Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

EquivDP3: A SIM(3)-Invariant Point-Cloud Encoder for Data-Efficient Humanoid Loco-Manipulation

Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose EquivDP3: a two-stage policy where a high-level diffusion planner with a SIM(3)-equivariant VNN encoder emits 6 Hz whole-body command chunks, executed at 50 Hz by a frozen, pre-trained RL locomotion policy and a differential inverse-kinematics module for the arms, trained end-to-end by behavior cloning. Across two simulated IsaacLab benchmarks and four non-equivariant baselines (5-100 demonstrations, in- and out-of-distribution), EquivDP3's advantage concentrates in the low-data regime: at 5-10 demonstrations it reaches 67.1% success versus 38.2-52.4% for the baselines, while by 50-100 all encoders converge (74.3-85.2%) and the ordering is no longer meaningful. A proprioception-only control confirms this gap is genuinely perceptual: with the point cloud removed, success drops to 31% vs. 60% (EquivDP3) at 5 demonstrations and 78% vs. 99% at 10, but vanishes by 50-100, showing the high-data plateau reflects a benchmark ceiling, not five encoders learning the same invariance. The encoder costs only 0.8 ms of extra latency per action chunk over the PointNet encoder it replaces. Baking geometric symmetry into a hierarchical diffusion policy's perception backbone is a practical, nearly free way to improve data efficiency for humanoid loco-manipulation when demonstrations are scarce.

cs.RO↗

Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI splits an attack objective into incomplete fragments, embeds them in external content retrieved through the agent's tools, and uses a reconstruction cue to prompt the agent to combine them. It then iteratively refines the fragments and cue with OpenEvolve using graded scoring and natural-language execution feedback from the target agent. Empirically, AdaLCPI achieves higher attack success than strong adaptive baselines, reaching 61.4\% macro-average ASR compared with 32.8\% for Trojan Hippo-style and 30.0\% for AgentVigil. Safety evaluations should therefore test whether agents remain robust when harmful objectives must be reconstructed from incomplete fragments.

cs.AI↗

Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code

cs.SD↗

Factorized Scheduling Principle: Learning Interpretable and Transferable Policies via Structured Additive Functions

Scheduling problems arise from repeatedly selecting one item from a set of candidates based on their states. These problems often reduce to assigning priority scores and choosing the highest-ranked item. In this work, we propose a factorized scheduling principle (FSP) framework to learn interpretable and transferable scheduling rules. The FSP framework represents system states as condition distributions and decomposes a global scheduling principle into additive univariate and pairwise components with identifiability constraints. The scheduling principle enables the framework to maintain a simple priority-based structure during deployment. This principle is learned by using a policy-based objective combined with a temporal-difference signal defined on the condition distribution. Experiments on synthetic and realistic scheduling tasks demonstrate the FSP framework's strong performance, interpretability, and zero-shot generalization across different system scales.

cs.LG↗

Root-system structure of sloppiness in passive Gaussian metrology

Passive transformations of an $n$-mode squeezed probe become locally unidentifiable when the quantum Fisher matrix loses rank. For pure zero-mean Gaussian probes, we show that the $C_{n}$ restricted-root system of $\mathrm{Sp}(2n,\mathbb{R})/\mathrm{U}(n)$ governs this exact sloppiness. For squeezing magnitudes $r_{1},\ldots,r_{n}$ in the canonical basis, the roots $2r_{j}$ and $r_{j}\pm r_{k}$ label the local phase rotations and two beam-splitter quadratures, and their vanishing identifies every Fisher-null direction. The associated Fisher weights quantify the approach to each singular wall. At fixed positive mean photon number, we solve the E-optimal probe-design problem of maximizing the smallest passive Fisher eigenvalue in two explicit generator normalizations. In the canonical phase and beam-splitter angle convention, the unique optimum in the fundamental Weyl chamber is an arithmetic progression approaching the consecutive-odd-integer dual-Weyl direction at large resource. With an invariant generator norm, the smallest eigenvalue is independent of the passive frame and the optimum lies exactly on that ray at every resource. Uniform squeezing instead minimizes the local-phase A-optimal cost while leaving one beam-splitter quadrature unidentifiable for every mode pair. Finally, the two quadratures of each identifiable mode pair saturate the quantum-geometric incompatibility bound, and the largest compatible passive submodel has $n(n+1)/2$ parameters at a regular spectrum. The resulting classification separates local identifiability, sensitivity optimization, and simultaneous attainability.

quant-ph↗

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.

cs.AI↗

MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms

Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.

cs.AI↗

Inferring Soil Friction Angle from Robot Foot-Ground Force Histories: A Bayesian Inverse Approach to Proprioceptive Soil Sensing

Foot-ground interaction signals recorded by quadruped robots may enable spatially distributed, in situ characterization of soil strength. As a first step, we test whether the internal friction angle $ϕ$ of cohesionless soil can be identified from the force history of a simplified rotating leg. A two-dimensional continuum model implemented with the material point method, benchmarked against measured rotating-leg force histories, generates the training data, and two Gaussian-process surrogates support Bayesian inversion of the full histories. In matched-model experiments, the framework recovers 14 off-grid friction angles with a median absolute error of approximately $0.1^\circ$ (maximum $\sim 0.7^\circ$); the reported credible intervals contain the true value in every case. These results establish that $ϕ$ is identifiable when the forward model is correctly specified, and support further development of proprioceptive soil sensing for spatially variable terrain, with applications from physics-grounded world models for robot training to post-wildfire slope assessment.

cs.RO↗

Attack-Resiliency Analytics for Wide-Area Control Systems in Smart Grids

Wide-area monitoring, protection, and control (WAMPAC) systems damp inter-area oscillations in large interconnected grids, but their reliance on synchronized PMU measurements carried over wide-area networks exposes them to false data injection (FDI) attacks. This paper presents an attack-resiliency analytics framework that formally models the coupled dynamics of the wide-area damping loop, automatic generation control, and the governor and excitation systems, and formulates the optimal stealthy FDI attack as a mixed-integer linear program (MILP). The anomaly detection model (ADM) enters as a replaceable constraint set: either a static bad-data detection (BDD) rule or a boundary learned from benign operation calibrated to a common false-positive rate. On the IEEE 39 and 118 bus systems, with reference dynamics validated on an OPAL-RT hardware-in-the-loop testbed, the optimal attack with wide-area access attains 3.3 and 8.1 times the benign objective and reaches a 0.5Hz frequency excursion more than twice as fast as an attack confined to automatic generation control. Learned boundaries reduce the attack objective by 15.5 55.0% and prevent over-frequency relay trips in all configurations, with feasible stealthy attacks remaining in every case, and residual risk tracks the width of the learned boundary rather than the detector family.

eess.SY↗

Making Analog Training Scale: Co-Designing Mapping, Optimizer, and Converters

Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains an open challenge due to severe hardware non-idealities, including physical weights with finite dynamic range and write granularity, analog-digital converters with finite resolution, and noisy and asymmetric updates. Guided by the insight that gradient accumulation is sensitive to precision and rounding errors, we adopt a mixed-precision training paradigm: executing forward and backward matrix multiplications in the analog domain while computing weight gradients in the digital domain. To enable scalable training, we present a holistic system-algorithm co-design that co-optimizes weight mapping to ensure well-conditioned physical and logical weight profiles, couples a preconditioned optimizer with threshold-triggered open-loop pulsing to stabilize training trajectories, and aligns converter dynamic ranges to suppress quantization errors. Evaluated via hardware-calibrated architectural simulations calibrated with electrochemical RAM measurements, our framework scales Transformer training up to $123\text{M}$ parameters with validation loss scaling as $L\propto N^{-0.231}$, where $N$ is the parameter count, comparable to $L\propto N^{-0.238}$ for digital training.

cs.LG↗

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this computation with all model weights frozen. Qwen3-8B improves from 15.5% to 99% exact accuracy on 24-line chains; a longer-trained LoRA reaches 50 lines. Ouro-1.4B reaches 60 lines after four loops and at least 160 after eight. The LoRA starts a relay: program lines pass on their chain identity through a short range of middle layers. Frozen heads read progressively further up the chain, and removing parent-line attention stops the relay. A frozen-model measurement locates the last useful intervention layer within tolerance in three of four held-out models. Task-specific LoRAs also improve MuSiQue. Default answers therefore understate the computation accessible through a tiny edit. Code and an interactive demo are available at https://lunamos.github.io/stop-thinking-too-early/

cs.AI↗

Surrogate-Enhanced Fractional Programming for MIMO Device-to-Device Interference Networks

Interference management in multi-stream multi-input multi-output (MIMO) device-to-device (D2D) networks often leads to weighted sum-rate maximization with sum-log-determinant involving matrix-valued signal-to-interference-plus-noise ratios. The state-of-the-art paradigms, including weighted minimum mean-square error (WMMSE) and fractional programming (FP), have achieved tremendous success in link scheduling, power control, and beamforming problems. Recently, an upgraded FP approach, named surrogate-enhanced FP (SEFP) in a scalar form, has attained improved performance in joint uplink scheduling and power control in coordinated multicell SISO networks. The proposed SEFP improves the surrogate construction of the classical Lagrangian dual transform plus quadratic transform for logarithmic fractional objectives with a novel reciprocal-inverse transform (RIT), yet its extension to the matrix form for MIMO settings does not seem straightforward because the matrix ratio and the auxiliary matrix generally do not commute. In this paper, by leveraging mathematical tools from Hermitian functional calculus, we extend RIT to matrix ratios by lifting the scalar RIT along the eigendirections of an auxiliary matrix and recasting the resulting directional construction in an operator form. Following, we develop a matrix SEFP framework for weighted sum-log-determinant maximization. Further, we establish a unified view of SEFP and the recently proposed XMMSE method, specify when the two algorithms attribute to identical minorization-maximization (MM) surrogates and variable update trajectories, and prove XMMSE can be considered as a special case of SEFP under the algorithmic family perspective.Inspired by the unified view, we develop SEFPLinQ for the joint scheduling and beamforming optimization in flexible-association MIMO D2D networks. Numerical results demonstrate consistent performance gains.

cs.IT↗

Learned Reporting Preferences in RLVR Can Conflict with the Current Request

Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions and independent human calibration. On GSM8K, boxed-format RLVR reduces the fraction of Qwen2.5-7B responses containing the requested hash-format payload by 35.33--74.37 percentage points relative to a 95.45% initial baseline in four of five training seeds; the fifth improves by 2.50 points. In the four deteriorating runs, almost every response that omits the requested payload instead retains the trained boxed convention, and the same four seeds deteriorate under two fixed paraphrases. Changing only the final-answer marker in supervised targets reverses which reporting convention the model prefers across three seeds, providing controlled evidence that this preference is learnable. Across three settings with independent human calibration, gains under a convention-sensitive scorer exceed the corresponding gains in committed-answer correctness, i.e., the correctness of the answer the model actually commits to. Together, these results separate three distinct post-training outcomes: learned reporting preference, current-request adherence, and committed-answer correctness. They show that convention-matched accuracy alone does not fully characterize post-training behavior and motivate evaluating current-request adherence alongside convention-matched task accuracy.

cs.LG↗

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to initialization shift with reduced human cost. Second, offline credit-filtered tuning assigns credit to individual agents and fine-tunes on per-agent trajectories with positive advantage rather than on entire joint rollouts. Third, we find existing online RL for VLAs are less effective for hard multi-agent tasks, which we attribute to noisy co-exploration and unstable updates. We instead use online latent-space fine tuning, which freeze the VLA and perform RL in its latent noise space. We evaluate our multi-agent VLA with both $π_0$ and $π_{0.5}$ backbones across 11 tasks in RoboTwin, RoboFactory and real-world manipulation with two Franka robots. Our multi-agent VLA improves the average success rate by $+23.1\%$, $+16.4\%$, and $+44\%$ on RoboTwin, RoboFactory, and real-world tasks, respectively. Code available at https://anonymous.4open.science/r/mavla_rft-2BC0/.

cs.RO↗

Non-Linear Pricing Restores Tractability for a Data Seller

We consider a data seller who designs pricing mechanisms over multiple datasets to maximize revenue from budget-constrained buyers. The seller offers multiple datasets and assigns each a pricing function that maps the quantity purchased to a total payment. The goal is to design these pricing functions to maximize revenue, anticipating that buyers---who trade off accuracy gains against cost---choose bundles optimally subject to their budget constraints. Prior work [Chaudhury et al., 2026] studies such optimal pricing under the restriction that each dataset is assigned a linear price, and shows that computing optimal linear prices is computationally intractable. In contrast, we allow each dataset to be priced via a general function and show that this additional flexibility can not only increase the revenue but also restore tractability, yielding a surprising simultaneous improvement in economic performance and computational efficiency. Even when pricing functions are only required to be monotone and lower-continuous, optimal pricing admits a highly structured and simple form: each pricing function is piecewise linear and convex (PLC), and the optimal solution can be computed in polynomial time. Moreover, the total number of kinks across all pricing functions is bounded by the number of buyers. Consequently, when datasets significantly outnumber buyers, most pricing functions are effectively linear. We further empirically study the structure of optimal pricing by analyzing the number of kinks and the revenue gap between optimal nonlinear pricing and optimal linear pricing on simulations generated from a real dataset.

cs.GT↗

SEED: Self-Speculative Decoding via Implicit Encoder-Decoder

Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-decoder (SEED), a self-speculative method that obtains high-quality drafts cheaply by reusing the deep contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder-decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder-decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the lightweight decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.7$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 28% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning. Code is available at https://github.com/lhk2004/SEED.

cs.CL↗

Mechanical Signature of a Chemically Driven Bath

Using optical tweezers, we demonstrate that the CuAAC click reaction increases the positional fluctuations of a trapped 2 micrometer colloid by at least 20% without altering the trap's corner frequency. This excess variance relaxes to the thermal baseline as reagents are depleted. The additional noise is Gaussian, with a flat power spectrum extending past the corner frequency up to 7 kHz. Notably, the noise amplitude decouples from the macroscale reaction rate, and independent molecular force dipoles are orders of magnitude too small to generate the observed variance. Although the chemical free energy released locally exceeds the energy absorbed by the bead by five orders of magnitude, this energy couples inefficiently. The mechanical action of the reacting bath cannot be represented as a sum of independent molecular events.

cond-mat.soft↗

Accurate recovery of the two linewidths hidden in laser beatnote statistics

Photons from ultrastable lasers can remain coherent over distances approaching the Earth-Sun separation, enabling optical clocks projected to lose less than one second over the age of the Universe. Yet characterizing these photons poses an identifiability problem: a two-laser heterodyne measurement produces a single beatnote linewidth that conflates contributions from both lasers. For more than half a century, the standard solution has been the three-cornered-hat (TCH) method, which requires three independent ultrastable lasers. Here we derive the mathematical form and elucidate the physical origin of asymmetric beatnote-linewidth distributions arising from finite photon wave trains, a long-observed feature not captured by canonical Gaussian or Lorentzian statistics. This finding accurately recovers two linewidths hidden in laser beatnote statistics without a third laser. The framework consistently captures the observed coherence-length and coherence-time statistics of photons, while comparison with TCH measurements confirms the quantitative validity of the extracted individual linewidths across five independent ultrastable-laser datasets spanning nearly two orders of magnitude. Most notably, the method resolves the 7.8-mHz linewidth of a cryogenic silicon-cavity laser beating with a broader laser. This transformative function of beatnote-linewidth distributions, together with the resulting method, could fundamentally advance optical clocks and precision metrology.

physics.optics↗