Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,153 records · Page 64Linked to original sources

MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.

cs.LG↗

ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context

Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current observation, and whether PRMs can estimate progress when context is needed remains underexplored. We therefore build ContextProgress-Bench, with 24 manipulation tasks for 120 episodes. The benchmark covers three settings: (i) State Recall, where information needed for progress appeared earlier but is not in the current frame; (ii) Sequence Tracking, where steps follow a fixed order, so progress requires knowing which steps are done and which comes next; and (iii) Recurrence Disambiguation, where look-alike frames sit at very different progress. We then run a paired diagnosis: each PRM keeps the same input format in both runs, and in one run its instruction integrates the right context. Even PRMs that read the entire history get lost in estimating progress, yet with the right context the same five models cut their progress error by 77-82%. Embodied PRMs are thus not incapable of progress estimation, but lost without the right context. We therefore propose ProgressCompass, an autonomous agentic loop that reorients an existing PRM and uses current general-purpose VLMs to supply the context the PRM needs. Wrapped in the loop, the same frozen PRM cuts its progress error by 63% and raises its rank agreement by 76%. With such a compass, PRMs estimate progress far better on longer, more complex tasks.

cs.CL↗

When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation

Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework for co-speech gesture generation. We first learn a discrete motion prior that represents continuous gestures in a compact and structured motion-code space. We then introduce a dual-branch semantic contribution estimation mechanism consisting of a full multimodal branch and an audio-only branch. Their distributional discrepancy is formulated as conditional information gain to quantify how much textual semantics changes the predicted motion. Based on this estimate, a controllable semantic-rhythm objective selectively strengthens semantic guidance in content-relevant segments while limiting unnecessary semantic intervention in rhythm-dominant segments. Furthermore, we treat background noise as an acoustic reliability condition and introduce noise-conditioned feature modulation together with beneficial latent perturbation to improve generation robustness under realistic acoustic environments. Experiments on benchmark datasets demonstrate that the proposed framework achieves a favorable balance among semantic expressiveness, rhythmic synchronization, motion diversity, and robustness, enabling reliable and controllable co-speech gesture generation.

cs.CV↗

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline's top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at https://github.com/cruiseresearchgroup/DecompRCA.

cs.LG↗

MyoCodec: A Streaming Neural Codec for Electromyography

Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.

eess.SP↗

GRP v0.1 Technical Report

Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.

cs.IR↗

CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference

Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergraphs into three stages, evidence weighting, evidence aggregation, and source fusion, and propose CHAIN, which designs stage-specific mechanisms to mitigate bias at each stage: (i) modulating the temporal decay function by causal topological distance, (ii) aggregating approximately independent causal chains via Noisy-OR after direction-aware deduplication, and (iii) driving adaptive fusion by causal coverage and directional balance. Experimental results on cross-domain forecasting benchmarks show CHAIN outperforms existing methods in expected calibration error, Brier score, and accuracy. Our project is available at https://github.com/QwenQKing/Chain.

cs.LG↗

Co-design of Silicon Microring Modulator beyond 200 Gb/s per Lane: Device Physics, Operating Point, and Compact Models for Scale-Up and Scale-Out Optical I/O

Optical input/output (I/O) supports high-bandwidth communication between processors in artificial intelligence (AI) systems. Depletion-mode silicon microring modulators offer compact footprints, wavelength multiplexing and femtojoule-scale junction switching energy per bit. However, as lane rates increase to 200~Gb/s and beyond, a digital signal processor (DSP) performing equalization and forward-error correction can consume up to half of the optical module power. We analyze the junction and cavity physics of silicon microring modulators, relating modulation efficiency, capacitance, optical loss and coupling to bandwidth and drive requirements to guide device optimization above 200~Gb/s per lane. Laser detuning and optimal operating points are examined by distinguishing maximum static slope, static and dynamic OMA, and gain--bandwidth product (GBW), including the influence of the junction RC response. The effects of optical self-heating, photocarriers and bias-network voltage droop on the resonance are discussed together with heater tuning and wavelength assignment in DWDM arrays. Cold-resonance design points are calculated from the required hot detuning, link-budget-derived optical power and assumed thermal parameters, with heater reserve and startup acquisition included as design constraints. A compact model of the coupled electrical, optical and thermal dynamics is used to evaluate PAM4 eye diagrams, and a Verilog-A implementation of the model core is provided for circuit-level simulation.

physics.optics↗

Video2Skill: From Streaming Experience to Reusable Embodied Skills

Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.

cs.CL↗

Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training

Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram representation, we observe that it jointly processes marginal-scale and interaction information. This opens an alternative way to organize geometric information hierarchically, motivating the Normalize-Then-Precondition framework. Specifically, it first uses diagonal-Gram information to construct a marginally normalized update, then applies spectral preconditioning to its directional interaction geometry. Building on this framework, we develop NormPre with NormPre-G and NormPre-L adopting global and localized spectral preconditioning, grounded in spectral-norm steepest descent and a regularized formulation followed by leading mode selection, respectively. To enable large-scale training, NormPre-G uses Newton-Schulz iterations and NormPre-L employs randomized sketching to approximate the leading interaction eigenspace. Theoretically, we establish $\mathcal{O}(T^{-1/2})$ convergence guarantees for simplified versions of NormPre. Across extensive pretraining experiments on GPT-2 Small, LLaMA and Qwen3, both variants consistently outperform AdamW, Muon and MANO under matched training budgets. Further efficiency and spectral analyses reveal the complementary strengths of two variants and characterize their performance-efficiency trade-off. We open-source our code through a GitHub repository at https://github.com/zx-gong/NormPre.

cs.LG↗

AESplat: Advancing Pose-Free Feed-Forward 3D Gaussian Splatting via Decoupled Appearance Modeling

Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a $0.8$ dB improvement in PSNR over the pose-free method NAS3R and a $1.1$ dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset. Project page: https://aesplat.github.io/.

cs.CV↗

Optimized Randomized Hamiltonian Simulation via Average-Error Analysis

Hamiltonian simulation is a central application of quantum computing. Randomized Hamiltonian simulation approximates the target dynamics by sampling quantum circuits and often allows simpler circuit implementations. We develop a framework for randomized Hamiltonian simulation in which Hamiltonian terms are sampled with arbitrary probabilities and the corresponding short-time evolutions are implemented sequentially. The leading channel error relative to ideal time evolution is governed by the variance of the sampled generators. Minimizing the variance-based upper bound on the worst-case error recovers qDRIFT, a leading randomized Hamiltonian simulation algorithm that samples terms in proportion to their operator norms. For typical input states, however, this sampling distribution need not be optimal. Minimizing Haar-averaged error bounds instead yields sampling probabilities proportional to the Hilbert-Schmidt norms of the terms. This choice can provide smaller average-error bounds than conventional qDRIFT. For Hamiltonians expressed as sums of Pauli strings, we combine this framework with commuting Pauli grouping to obtain smaller error bounds than conventional qDRIFT at fixed $R_z$ depth. Numerical benchmarks for the Sachdev-Ye-Kitaev model show error-bound improvement factors consistent with linear scaling in the number of qubits. For the 108-qubit FeMoco Hamiltonian, the error bounds are reduced by a factor of about 5.65. These results establish average-error optimization as a practical design principle for randomized Hamiltonian simulation.

quant-ph↗

Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation

Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow. InFlow first retrieves potentially informative sources using certainty-calibrated hidden-state trajectories, then measures their realized effect through the Jensen--Shannon divergence between the teacher's initial and retrieval-conditioned answer beliefs. Examples with larger belief shifts are selected for on-policy distillation. Our analysis formalizes the information optimized by retrieval and selection and relates the answer-level shift to the teacher--student distillation gap. Across four open-weight language models and three knowledge domains, InFlow achieves the strongest cross-model average among the compared selection methods, with ablations supporting both stages of the framework. Our code is available at https://github.com/1240148048/INFLOW.

cs.LG↗

Sensitivities of Long and Medium Baseline Experiments to Sterile Neutrinos and Non-Standard Interactions

In this work, we present for the first time, the simultaneous analysis of the non-standard neutrino interactions (NSI) and a light sterile neutrino in the 3+1 framework for the future long- and medium-baseline experiments DUNE and MOMENT, respectively. We show that treating the two sectors separately yields over optimistic sensitivities. The simultaneous presence of sterile-neutrino mixing and complex NSI weakens the $|\varepsilon_{eμ}|$ sensitivity by a factor of $\sim (2-3)$ with the dominant degradation associated with the additional parameter freedom in the flavor-changing NSI sector. DUNE alone fails to achieve $5σ$ CP-violation discovery, once NSI phases are marginalized over. The DUNE+MOMENT combination recovers $>5σ$ sensitivity over a wide range of $δ_{\rm CP}$ space, constrains $\sin^2θ_{24}\lesssim10^{-2}$ and $|\varepsilon_{eμ}|\lesssim0.025$ at $95\%$~C.L., and resolves degeneracies that are intractable for either experiment alone. Our results establish that a multi-baseline strategy combining matter-rich and near-vacuum baselines is necessary for robust parameter extraction in the simultaneous presence of two new physics scenarios. It is also important to perform simultaneous new physics analyses to reliably interpret future precision neutrino oscillation data.

hep-ph↗

Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms

To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are presented.

cs.LG↗

Electronic excitation spectra and recovery of excited states with neural network wave functions

Accurate electronic spectra require both a flexible description of electron correlation and a tractable treatment of the many states contributing to the response. We combine neural network wave functions with the Lorentz integral transform to calculate electronic spectra directly in continuous coordinates, without truncation error from a fixed one-electron basis and with polynomial computational cost per optimization step. Instead of constructing a prescribed set of excited states, the method solves an inhomogeneous Schrödinger equation at a chosen complex energy. This formulation gives access, in principle, to the entire spectrum coupled to a perturbation, including bound excitations and the ionization continuum, without explicitly determining all lower-lying eigenstates. A finite imaginary energy controls the resolution and keeps the response square integrable. Near an isolated bound excitation, the normalized response also recovers the corresponding eigenstate as the width tends to zero. A helium application illustrates the extraction of an excitation energy and oscillator strength. The formulation provides a route from neural descriptions of electronic correlation to spectra beyond a small manifold of low-lying states.

physics.chem-ph↗

Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG

When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answering (QA) benchmarks into underspecified conversations and evaluate ten LLM assistants with eight retrieval systems across 1.5 million simulated conversations. Our findings reveal that multi-turn interaction causes widespread performance degradation, incurring relative performance drops of up to 21% and increasing unreliability by 47%, making RAG systems simultaneously less accurate and less reliable. We identify two distinct failure modes behind this degradation. Systems are either lost in translation, where conversational rephrasing distorts the retrieval query, or lost in conversation, where retrieval succeeds but the LLM fails to synthesize evidence distributed across turns.

cs.CL↗

Interpolating Neural Operator (INO): A Data-Free and Efficient Approach for Learning PDE Solution Operators

Neural operators have become a popular approach to approximate the solution operators of parametric partial differential equations (PDEs). However, existing neural operators either require a large amount of simulation data or a long physics-informed training on GPUs, and they cannot tell how accurate an individual prediction is. In this paper, we propose the Interpolating Neural Operator (INO), a data-free interpolating neural network that is trained directly on the weak form of the PDE. In INO, the Karhunen-Loève coordinates of the input field are treated as additional inputs together with the spatial coordinates, and each input is approximated by a C-HiDeNN sub-network whose trainable parameters are nodal values. Since the network is multilinear in its parameters, training reduces to a sequence of one-dimensional linear solves by greedy alternating least squares. As a result, INO trains on one CPU core and predicts a new solution in microseconds. For coercive problems, the total error of every prediction is bounded by a computable residual bound that requires no reference solution, and the same bound applies to the predictions of other methods that satisfy the boundary conditions exactly. Before each prediction, INO checks whether the leading coordinates of the input lie within the range on which it is trained, and inputs outside this range can be passed to a conventional solver or to an INO trained on a wider range. INO is compared with physics-informed FNO and DeepONet on different benchmarks. INO is the most accurate model on most of these problems, by 15$\times$ on two-dimensional Helmholtz at $65^2$ and 53$\times$ on the diffusion-reaction benchmark, and on the one- and two-dimensional problems its training on one CPU core takes 3-80$\times$ less time than the physics-informed baselines on one GPU.

cs.CE↗