Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 541 records · Page 30Linked to original sources

Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models

Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the original data and gradients of off-the-shelf pre-training LLMs required by these methods are strictly unknown and highly diverse. To bridge this critical gap, we propose EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. Furthermore, we formulate a multi-objective optimization problem and introduce a first-order efficient Pareto optimizer that jointly optimizes LLM parameters and the soft prompt, rigorously enforcing orthogonality between new task updates and the estimated pre-training gradients. Extensive experiments across multiple LLMs demonstrate that EoupCT effectively preserves both task-specific proficiency and inherent general-purpose knowledge, successfully mitigating the catastrophic forgetting.

cs.CL↗

Self-Play Search Distillation for Large Language Model Reasoning

Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.

cs.AI↗

Sub-Horizon and Quasi-Static Approximations in Interacting Dark Sector Models Involving Scalar Field Dark Energy

We investigate the validity of the sub-horizon approximation (SHA) and quasi-static approximation (QSA) for linear matter perturbations in interacting scalar-field dark-energy models. We consider both quintessence and phantom scalar fields and derive their full relativistic perturbation evolution to compare with the corresponding SHA and QSA solutions. We find that, while the SHA reproduces the full evolution with percent-level accuracy for non-interacting and weakly interacting models on sufficiently sub-horizon scales, its accuracy progressively deteriorates with increasing interaction strength and towards larger scales. Moreover, unlike in $Λ$CDM and non-interacting scalar-field models, applying the SHA to interacting models does not fully recover the standard Newtonian fluid equations. The QSA generally provides a closer description of the full relativistic evolution than the SHA over the scales considered. By comparing the scalar-field model with $Λ$CDM using the full relativistic equations and, separately, using the SHA system, we evaluated whether the SHA is capable of capturing the full extent of the physical differences introduced by the alternative model. We find that the difference between these two comparisons becomes increasingly important for stronger interactions and smaller wavenumbers. Finally, we demonstrate that the resulting differences can reach several percent on scales accessible to large-volume hydrodynamical simulations, implying that the validity of the SHA cannot be assumed solely from the sub-horizon condition. Our results highlight the need to benchmark the approximation over the relevant interaction parameter and wavenumber ranges before employing interacting-dark-energy models in large-volume structure-formation simulations and precision comparisons with large-scale-structure observations.

gr-qc↗

Landscape Limits of Quantum-Inspired Evolutionary Optimization across 256 continuous functions

Quantum-inspired evolutionary optimization (QIEO) represents design variables as a set of qubits and searches a continuous, multi-dimensional landscape through rotation of the qubit's amplitude pair. Every generation rotates those amplitudes toward a single elite, which corresponds to that generation's best. The update is cheap, almost parameter-free, and well-suited for massive parallel implementation, which has encouraged its adoption in engineering, design, and planning applications. However, there are critical issues with this formulation, principally, the treatment of design variables as independent probability components which make it incapable of exploiting local curvature, anisotropy, or variable coupling. Despite this, QIEO is believed to hold promise, and has been used extensively to solve real-world problems, with significant qualitative and computational advantage over its classical counterpart, Genetic Algorithm (GA). A collection of 256 (actually 508; 256 unshifted + 252 shifted, 4 could not be shifted) continuous function are selected from the prior works, in such a way that they represent eleven landscape characteristics, namely continuity, differentiability, separability, scalability, modality, convexity, conditioning, symmetry, maximum dimensionality, dimension dependency, and the coupling pattern of the design variables. These functions are then solved by three QIEO variants, two GA encodings and Hansen's Covariance Matrix Adaptation Evolution Strategy (CMA-ES). The results are evaluated in terms of computational cost, solution precision, and specialization across landscape characteristics. They identify the conditions under which QIEO provides competitive performance, clarify where its independent-variable representation becomes limiting, and establish whether particular QIEO variants offer advantages for specific landscape characteristics.

cs.NE↗

MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory

Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, cognitive restructuring, behavioral experiments, and treatment monitoring) into five collaborative agents. CD Memory tracks cognitive-distortion type, frequency, severity, and restructuring efficacy across sessions to generate pre-session pathology reports and intervention-priority recommendations. We construct a Chinese CBT dialogue corpus via dual-role large language model simulation and train a Qwen3-14B backbone with supervised fine-tuning and direct preference optimization. Evaluation with GPT-4 judges shows MACBT outperforms MeChat, SoulChat, PsyChat, and CPsyCounX in professionalism (2.62) and clinical authenticity (2.25). The full memory-augmented system further improves session quality by 12.6% and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.

cs.AI↗

Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms

Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\% of baseline bank-run episodes and 83\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.

cs.AI↗

Spackle: Completing Large View Single Image NVS with Adaptive Gaussians

Single-image novel view synthesis (NVS) enables photorealistic rendering of un- observed viewpoints from a single input. Practical NVS systems require two key capabilities: robust reconstruction of occluded regions and high inference effi- ciency. While hybrid decoupled frameworks combining feedforward 3D Gaussian Splatting (3DGS) and diffusion models show promise for large-view-deviation NVS, they suffer from capacity competition: a fixed number of Gaussians forces resource shifts from visible to newly disoccluded areas, degrading original scene fidelity when the target view deviates significantly from the input. To address this, we propose Spackle, a lightweight residual learning framework that mit- igates capacity competition without sacrificing efficiency. Spackle operates in three stages: predicting base 3DGS attributes from given views, automatically identifying poorly reconstructed regions, and learning a residual 3DGS optimized exclusively for these areas. At inference, we combine the baseline and aug- mented Gaussians for NVS. We conduct comprehensive experiments and show that Spackle achieves state-of-the-art performance on large-view-deviation cases.

cs.CV↗

Non-factor quantum dynamics and relativity

We consider the effect of relativistic principles on the structure of the local evolutions of non-factor quantum systems such as those which naturally arise in quantum spacetimes and causal structures. Generalising a former result that had been given for factor systems, we show that any evolution which can be applied locally on a direct sum of factor systems and without producing superluminal signals must be linear, unitary, and respect the block decomposition of the system.

quant-ph↗

LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting

Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section generation or rely on manually crafted outlines, limiting scalable automation in realistic settings. In this work, we propose LogicTree-RAG, a logic tree-guided retrieval-augmented generation framework that induces a hierarchical logic tree as a global organizational backbone to organize and ground technical disclosures, without relying on expert-defined drafting priors. Each node in the logic tree represents a technical element and is constructed through evidence-guided recursive generation. A hybrid traversal mechanism then maps the logic tree into patent sections, enabling controllable and section-balanced generation. Extensive experiments show that LogicTree-RAG consistently improves content quality and language conformity over strong LLM-based baselines and achieves longer structured generation with high token efficiency, demonstrating the effectiveness of logic-centric generation for complex technical document drafting.

cs.AI↗

Representation-Aware Transport-Information Measure for Non-inclusive Discrete Supports

Information-theoretic measures for comparing probability distributions are widely used across physics and other fields. When two discrete distributions have non-inclusive supports, however, the Kullback-Leibler (KL) divergence is in general not directly applicable, and various alternative divergences and distances have been introduced. These measures compare the resulting distributions themselves, but do not generally retain information about the representation transformations by which the discrete distributions are generated from underlying continuous ones. Here we introduce a representation-aware transport-information measure for discrete distributions with non-inclusive supports, formulated based on the standard KL divergence. We consider two continuous reference distributions, each transformed into a discrete representation through its own discretization scheme. Rather than comparing only the resulting discrete distributions or their continuous references, we additionally retain local information associated with the representation-change schemes. The resulting measure can therefore distinguish discrete representations that may have identical discrete probability landscapes but originate from different continuous references or discretization schemes. The construction is based on the transport-information cost of continuous-to-discrete representation in the framework of unavoidable canonical nonlinearity (UCN). UCN provides a non-arbitrary correspondence between the transport cost of discretization as an extrinsic geometric operation and the information-theoretic indistinguishability of nearby continuous distributions, thereby allowing a discrete representation to be associated with a local family of underlying continuous distributions on the statistical manifold.

cond-mat.stat-mech↗

Coupled Meta-Adaptive Filtering for Active Noise Control Under Time-Varying Acoustic Paths

Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic paths remain a major challenge. In particular, the physics-informed optimizer features are constructed using a secondary path estimate and become mismatched when the physical path changes, leading to inaccurate filter updates and degraded noise reduction. To address this problem, this paper proposes a Coupled Meta-Adaptive Filtering Active Noise Control (CoMeta-AF-ANC) method, which applies meta-learning to jointly learn the control filter adaptation and acoustic path tracking within a closed-loop framework. The proposed Meta-Gated Joint Path Identifier (MG-JPI) simultaneously tracks the primary and secondary paths from the available ANC signals without auxiliary noise, while the updated secondary path estimate is fed back to reconstruct the Meta-AF controller features. A delayless dual-rate realization performs learned adaptation at the frame rate while generating the control signal at the sampling rate in the time domain. Evaluation using measured headrest acoustic paths shows that CoMeta-AF-ANC outperforms representative ANC algorithms in tracking time-varying acoustic paths, while maintaining higher stability across unseen head movement scenarios. It also generalizes well to real-world noises not encountered during training.

eess.AS↗

OneWorld: Learning Consistent Physics Across Actions in World Models

Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain a coherent understanding of the underlying world and limiting its reliability for planning and decision-making. To address these issues, we propose OneWorld, a shared-mechanism counterfactual generation framework that jointly models multiple action-conditioned futures under a common latent physical mechanism. A physical mechanism interpreter first infers a distribution over latent mechanisms from each action-outcome branch. These distributions are then aggregated into shared-world evidence, which captures whether the branches admit a common physical explanation while accounting for uncertainty in less informative branches. This evidence constrains flow training and guides sampling, encouraging consistency in the underlying physical mechanism while preserving the distinct outcomes induced by different actions. We further introduce a multi-intervention evaluation protocol in controlled environments, following the interaction settings of ACWM-Phys, to assess whether generated futures can be jointly explained by the same physical parameters, alongside standard measures of single-rollout prediction quality. Experiments in these environments show that OneWorld improves cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.

cs.CV↗

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

Visual odometry is essential for autonomous navigation in GPS-denied environments, yet RGB-based methods remain vulnerable to motion blur, challenging illumination, and dropped frames. Event cameras complement conventional cameras with high temporal resolution and dynamic range, but their asynchronous measurements complicate reliable correspondence estimation. We present DAPEVO, a learned visual odometry system that estimates image and event correspondences independently at shared patch locations and fuses their correlation evidence before motion refinement. Each tracked patch maintains image and event descriptors, and a learned scalar gate combines modality-specific correlation embeddings for each patch--frame edge before a shared recurrent refinement and bundle-adjustment update. DAPEVO also supports event-only observations, enabling continued tracking when RGB frames are sparse or unavailable, while modality-aware keyframe culling preserves scarce frame constraints. On UZH-FPV, when retaining only one in six RGB frames, DAPEVO's mean absolute trajectory error (ATE) increases by only 36%, from 1.00 to 1.36m, whereas the ATE of DPVO and RAMP-VO rises by factors of $3.7\times$ and $3.1\times$, respectively. On TartanEvent, DAPEVO similarly remains below 1m ATE at 3Hz RGB input, while DPVO and RAMP-VO exceed 9m. Under degraded RGB input on TartanEvent, DAPEVO achieves an ATE of 0.60m, compared with more than 4m for both DPVO and RAMP-VO, while also outperforming event-only DEVO at 0.87m.

cs.CV↗

PORL: Pretrained Offline Reinforcement Learning for the Job Shop Scheduling Problem

The Job Shop Scheduling Problem (JSSP) is a fundamental combinatorial optimization problem in industrial optimization. This work introduces Pretrained Offline Reinforcement Learning (PORL), a hybrid approach that combines simulation-based online pretraining with offline fine-tuning on production-specific data. Reinforcement learning through online interaction enables exploration of general scheduling strategies, but typically relies on simulation environments and may suffer from a simulation-to-reality gap. In contrast, offline RL avoids direct interaction with the environment by learning from historical data, but its performance is strongly influenced by dataset quality and coverage. PORL combines the strengths of both paradigms by first learning a general scheduling policy through online interaction and subsequently adapting it offline to a target distribution. A KL-divergence-based policy constraint is introduced to limit deviations from the pretrained policy during fine-tuning. The approach is evaluated on JSSP instances with distribution shift and datasets generated from heuristic, noisy-expert, and random behavioral policies. The results show that PORL consistently achieves lower optimality gaps than standalone offline RL and the considered general scheduling baselines. Furthermore, its advantage over standalone offline RL increases as dataset quality decreases, indicating reduced sensitivity to the quality and coverage of the available offline data. The results suggest that offline adaptation of pretrained policies is a promising approach for industrial scheduling environments where direct online exploration is impractical.

cs.LG↗

A unified approach to nonlinear Stein theorems

We study regularity for the variable exponent quasilinear $\mathfrak{p}\left(x\right)$-Laplace type equation on domains in Heisenberg groups and Euclidean spaces. We establish borderline continuity estimates for the appropriate first order derivatives of solutions. More precisely, we show that for any weak solution $u \in \mathbb{E}W^{1, \mathfrak{p}(\cdot)}\left(Ω\right)$ of \begin{align*} \operatorname{div}_{\mathbb{E}} \left( \mathfrak{a}(x) \lvert \nabla_{\mathbb{E}} u \rvert^{\mathfrak{p}\left(x\right)-2} \nabla_{\mathbb{E}} u \right) = f \qquad \text{ in } Ω, \end{align*} where $Ω\subset \mathbb{E}^{n}$, $\mathfrak{a}$ is a uniformly positive bounded scalar function and the exponent function $\mathfrak{p}$ is uniformly bounded away from $1$ and $\infty,$ $\nabla_{\mathbb{E}}u$ is continuous in $Ω$ as soon as $f \in L^{\left( Q_{\mathbb{E}}, 1\right)}\left(Ω\right)$ and $\mathfrak{a}, \mathfrak{p}$ satisfies some conditions regarding the summability of their mean-oscillations. Here $\mathbb{E}^{n}$ is either the Heisenberg group $\mathbb{H}_{n}$ or the Euclidean space $\mathbb{R}^{n}$ and $\operatorname{div}_{\mathbb{E}}$, $\nabla_{\mathbb{E}}$, $Q_{\mathbb{E}}$ stands for the corresponding divergence, gradient and homogeneous dimension, respectively. We treat the elliptic and subelliptic cases in a unified manner using Euclidean techniques and our conditions on $\mathfrak{a}$ and $\mathfrak{p}$ are new and weaker than all the known sufficient conditions even in the Euclidean case. However, all the known sufficient conditions imply our conditions, achieving yet another unification.

math.AP↗

Low-Bit Recurrent States in Hybrid Language Models

Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.

cs.LG↗

Bundled Contact Gradients: Stabilizing Differentiable Simulation for Deployable Dynamic Tasks

Differentiable simulation provides analytic gradients of robot dynamics, enabling fast and sample-efficient first-order policy optimization. However, obtaining smooth and informative gradients through rigid-body contact typically requires softened contact models, often at the expense of physical fidelity and thereby limiting learned policies largely to simulation. This trade-off becomes particularly consequential for dynamic humanoid motions, where accurate contact dynamics are critical for transferring policies to the real world. Increasing contact stiffness in rigid-body simulation improves the fidelity of interactions, but also makes the dynamics increasingly sensitive to small state perturbations, producing high-variance gradients that can destabilize first-order policy learning. To address this, we propose \emph{Bundled Contact Gradients (BCG)}, a contact-local randomized smoothing framework for differentiable policy learning. When stiff contact is detected, our method evaluates a local bundle of randomized perturbation rollouts around the stiff contact configuration and aggregates their gradient signal thereby reducing gradient variance. We demonstrate the effectiveness of our method by successfully training and transferring dynamic motions zero-shot onto a real-world Unitree G1 humanoid platform. Videos and supplementary information can be found at https://bundledcontactgradients.github.io/

cs.RO↗

MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Multi-view video understanding requires integrating spatial and temporal evidence across multiple, often non-overlapping camera streams: tracking entities as they transition between viewpoints, aligning events across time, and reasoning about latent 4D continuity rather than any single visible frame. We introduce MVVBench, a benchmark for multi-view video reasoning built from real world multi camera datasets. Questions are curated to be monocular-ambiguous along both the view and the temporal axis: each question is unanswerable from any single view in the designated input set, and the majority are further unanswerable from any single moment. Each question becomes uniquely solvable only by jointly reasoning across views and across time. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, with human-authored QA and rigorous verification. Beyond benchmarking, we provide an extensive analysis of when and why current vision language models succeed or fail, characterizing errors due to temporal mis-localization, cross-view identity breaks, and brittle multi-hop reasoning. We then study inference-time elicitation strategies that unlock latent multi-view competence---task-specific chain-of-thought scaffolds and structured cross-view evidence aggregation---yielding substantial gains without retraining. Finally, we present preliminary evidence that reinforcement learning with verifiable rewards can elicit some latent multi-view competence in the base model, pointing to training-time approaches as a promising direction for future work. Together, MVVBench offers a rigorous evaluation of 4D multi-view reasoning and a foundation for future progress toward reliable embodied perception.

cs.CV↗