SearcharxivSearch

arXiv subjects

Shuo Zhang

Publications and source records attributed to Shuo Zhang.

At least 19 recordsLinked to original sources

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.

cs.CL

FAST Observations of Filamentary and Compact H I Structure in a Magellanic Stream IV Field

The Magellanic Stream (MS) is believed to have formed from gas removed from the Large and Small Magellanic Clouds through tidal forces and hydrodynamic interactions with the Milky Way's gaseous halo. It provides an important laboratory for studying how stripped gas fragments, mixes, and evolves in a circumgalactic environment. In this work, we present HI observations of a $3.8^\circ\times2.2^\circ$ field in the MS IV region, using the data from the Commensal Radio Astronomy FasT Survey (CRAFTS). The total HI mass in the analyzed field is $\simeq 5.3 \times 10^{6} (d/120\,{\rm kpc})^{2}\,M_{\odot}$, where $d$ is the distance to the MS IV gas. The data resolve the emission into three coherent filamentary HI structures with related but distinct velocity trends. To characterize the HI structures and study potential multiphase gas, we adopt a Gaussian decomposition procedure to identify and reconstruct sources. Our results indicate that the field is dominated by one major filamentary HI complex, together with several smaller kinematic clump-like components. Among these identified sources, only one source likely shows multiple velocity components. Because several sources appear spatially overlapped in projection, we further examine their apparent overlap regions using position-velocity (P-V) diagrams. P-V diagrams across the apparent overlap region show no clear intermediate-velocity bridge or V-shaped structure, favoring line-of-sight projection over direct cloud-cloud collision. These results demonstrate the value of deep, high-angular-resolution HI observations for resolving faint emission, compact morphologies, and kinematic structure in the MS.

astro-ph.GA

Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG

Retrieval-Augmented Generation (RAG) enhances language models with external knowledge, but the lengthy retrieved context inflates the input and degrades inference efficiency. Soft context compression encodes each document into a substantially shorter embedding sequence. However, most existing approaches are trained by distilling outputs from uncompressed RAG systems, inherently limiting their performance relative to the original model. To address this limitation, we propose DEX-Comp, a two-stage training recipe: Pure Distillation warm-starts the compression model on the uncompressed RAG's correct responses only, and Hard Exploration then runs reinforcement learning solely on queries the uncompressed RAG fails, forcing the model to explore computation patterns better suited to compressed representations. On five open-domain QA benchmarks at retrieval depths from top-5 to top-30, DEX-Comp compresses retrieved contexts by $16\times$ and accelerates inference by $4\times$--$24\times$, while achieving performance comparable to or exceeding the uncompressed RAG baseline across retrieval depths. Ablations and evaluations across diverse datasets and backbones further confirm the contribution of each stage and the generalization of our approach.

cs.CL

Dense Process Supervision for Search Agents via Fact Utility Estimation

Reinforcement learning (RL) for search agents typically relies on outcome rewards. However, it often fails to achieve effective credit assignment, due to the unclear value of intermediate steps. It is hard to separate their contributions from the final result. In this paper, we propose a dense process supervision method based on fact utility estimation, which models the reasoning process as the accumulation of discrete evidence facts. We first extract structured facts from raw observations and organize them into an explicit fact store. To support credit assignment, we then cluster semantically equivalent facts and infer the posterior utility of each fact cluster using Bayesian estimation over group rollouts. Finally, we convert the estimated fact utilities into dense step-level rewards to guide RL training. Experiments on seven single-hop and multi-hop QA benchmarks show that our method consistently outperforms existing baselines. Ablation studies validate clear relative improvements on multi-hop QA compared to outcome reward-only training.

cs.CL

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues across multiple audio inputs, requiring models to identify types of degradation, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and comprehend low-level acoustic phenomena over multi-turn dialogues. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. Human evaluations further reveal a significant perception gap between LALMs and human listeners. MRMAD thus exposes a critical yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.

cs.SD

BTF-PINN: Enforcing Dirichlet Boundary Conditions Without Boundary Training

The homogeneous Dirichlet boundary value problem captures the core difficulty of solving Dirichlet problems with non-interpolatory methods. We propose BTF-PINN (Boundary-Training-Free Physics-Informed Neural Network), an interior-only strategy for solving homogeneous Dirichlet boundary value problems that requires no boundary training, boundary penalties, or boundary-conforming parametrizations. The key idea is to embed the essential boundary condition into a newly designed boundary-free loss function. We prove the equivalence between the proposed interior-only variational formulation and the original boundary value problem, establish a sharp threshold condition for the residual weight, and develop a convergence analysis based on the coercivity of the functional. Numerical experiments on high-dimensional problems, irregular geometries, and anisotropic elliptic equations demonstrate the effectiveness of BTF-PINN. Comparisons with standard boundary-penalty PINNs further show that BTF-PINN achieves superior boundary trace accuracy.

math.NA

Asymptotics of Titchmarsh--Weyl functions near the real axis andan application to the KdV hierarchy

We establish high-energy asymptotic expansions of Titchmarsh--Weyl functions for one-dimensional Schr\"odinger and Dirac operators in regions whose boundaries approach the spectrum at a prescribed polynomial rate. For bounded potentials with bounded derivatives, the expansions remain uniform up to these boundaries, with an explicit loss in the remainder determined by the rate of approach. We treat both self-adjoint Dirac operators and non-self-adjoint Dirac operators with skew-adjoint potential matrices, keeping track of the distinct half-planes in which the two scalar Weyl coordinates are naturally defined. As an application of the Schr\"odinger expansion, we verify the high-energy hypothesis in Kotani's construction of KdV flows: for every odd integer $p\geq3$, each real-valued $q\in W^{2p-1,\infty}(\mathbb{R})$ generates a global classical solution of the member of the KdV hierarchy indexed by $(p+1)/2$.

math.SP

Explicit Hamiltonian Classification in the $F_4(0)$ Toric Degeneration of $CP^2$

We give an explicit coordinate description of the Hamiltonian isotopy classes of the regular Lagrangian torus fibers of the smoothing \(\widehat{F}_4(0)\) of the \(F_4(0)\) toric degeneration, expressed in the explicit Oakley--Usher coordinates on \(\CP^2(\sqrt2)\). For the wall fibers, they are not Hamiltonian isotopic to standard toric fibers, and no two distinct wall fibers are Hamiltonian isotopic. For the off-wall fibers, we find the standard toric fibers they are Hamiltonian isotopic to.

math.SG

HCPG-Flow:Hierarchical Contact-Progress Guidance for Flow-Policy Robot Manipulation

Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. When several proposals are sampled, critic-based ranking makes data collection depend on value estimates over candidate actions that may be weakly represented in replay. We introduce HCPG-Flow, an analytic rollout-time selector that augments SAC-Flow with hierarchical, object-centric contact-progress guidance while preserving its actor and critic objectives. HCPG switches from end-effector approach to task progress after contact, scores each proposal by the first-order reduction of a task-relevant distance, standardizes scores within the candidate set, and executes a temperature-controlled action embedding. Across ten simulated tasks, HCPG improves mean success over SAC-Flow on both benchmarks, including a 9.5 percentage-point gain on Maniskill. Four physical tasks further show high success with a 17.4% reduction in successful completion steps.Project page: https://hitxraz.github.io/HCPG-Flow/

cs.RO

Convolution for Large Language Models

Large language models (LLMs) largely rely on Transformers, where self-attention provides global token interaction but does not explicitly encode the locality of natural language. We study whether lightweight depthwise convolutions can supply this local inductive bias without materially increasing model size. Our macro-level ablation compares convolution at 17 locations in a Qwen3 Transformer block and finds the best results when convolution is applied to the projected queries, keys, and values before attention. A subsequent micro-level study favors a residual depthwise convolution with kernel size $k=3$, without additional normalization or activation. Across Qwen3 models and several pre-training data budgets, this design improves the average accuracy on seven downstream benchmarks while adding less than $0.01\%$ parameters. A representation-level case study further suggests that the convolution makes repeated token IDs more sensitive to their immediate context. These results support depthwise convolution as a lightweight complement to self-attention for modeling short-range token interactions.

cs.CL

Events as Spacetime Anchors: Local Irreversibility at the Interface of Quantum Field Theory and Relativity

General relativity (GR) is naturally organized around spacetime events and their causal order, whereas quantum field theory (QFT) is formulated in terms of states, operators, and unitary evolution, without an intrinsic criterion for when a quantum process becomes a definite spacetime fact. We propose an event-centered framework in which events are locally irreversible records generated by quantum-environment interactions and serve as an interface between quantum dynamics and relativistic spacetime structure. Event anchoring is characterized operationally by three jointly sufficient conditions: local classicalization, redundant environmental recording, and irreversibility against recovery. Their joint satisfaction defines local generative freezing. We realize these criteria in an explicit repeated-collision open-system model. The model is illustrative rather than a derivation from relativistic QFT, but it remains globally unitary and yields the freezing time in closed form. For all admissible tolerances, local classicalization is certified no later than channel-level irrecoverability, and generically earlier. Full anchoring occurs only when both irrecoverability and the required record redundancy have been reached, producing irrecoverability-limited and redundancy-limited regimes. At the level of anchored records, physical history is therefore represented as a partially ordered causal skeleton of frozen events on which effective field-theoretic descriptions operate. The framework addresses event anchoring--when a candidate outcome becomes a stable spacetime fact--while remaining compatible with standard QFT and relativistic causality. It does not derive the Born rule or solve single-outcome selection, which belong to the separate problem of event generation.

quant-ph

Agentic Routing: The Harness-Native Data Flywheel

Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becoming structurally specialized: a model that is strong at code editing, long-context recovery, tool use, mathematical reasoning, or low-latency response may not dominate on the other axes. This makes model selection inside an agent a core systems problem rather than a per-query serving trick. Existing routing methods mostly optimize single-turn cost-quality trade-offs and therefore miss the execution state, intermediate failures, and feedback loops that make agents different from chat completion. We propose Harness-Native agentic routing, a step-level routing paradigm that selects either a single best-fit model for cost-effective execution or multiple complementary models for ensemble-style accuracy improvement, conditioned on the full harness state. The key insight is that every routing decision naturally produces a structured data record -- consisting of the query, harness state, model choice or model set, execution trace, outcome, and cost -- whose labels are supplied by the environment rather than by the router itself. These records form a harness-native data flywheel: execution traces train better routers and harness-native models, which improve cost-quality trade-offs and generate more traces under the same budget. We instantiate this idea in OpenSquilla with a four-layer routing stack, an open LightGBM cold-start ranker, and a staged router-model path that turns logged arena records into progressively stronger routing policies. The report studies singleton and multi-model routing on agentic benchmarks including DRACO and PinchBench, and argues that agentic routing is not merely cost control, but a data engine for agent-native training.

cs.CL

Global well-posedness of the Toda lattice on an exact spectral phase space

We identify an exact spectral phase space for the two-sided Toda lattice. Let $q=\{a_n,b_n\}_{n\in\mathbb Z}$ be coefficients of the right and left half-line Jacobi operators and denote their spectral measures by $\sigma_{\pm}^{q}$. Define a phase space \[ \mathcal Q=\left\{ \begin{array} [c]{c}% q=\{a_n,b_n\}_{n\in\mathbb Z}: a_n>0,\ b_{n} \in \mathbb{R} \text{ and} \int_{\mathbb R}e^{c|\lambda|}\sigma^q_\pm(d\lambda)<\infty \text{ for every }c>0 \end{array} \right\} . \] The integrability condition makes the representing measures unique. We prove that $q\in\mathcal Q$ if and only if the Toda lattice with initial datum $q$ admits a classical solution for all positive and negative times. Moreover, the solution remains in $\mathcal Q$, is unique, and depends continuously on the initial datum, uniformly on compact time intervals.

math.DS

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-level visual perception. We present MonkeyOCRv2, a visual-text pretrained model for document AI. First, we construct MonkeyDoc v2, to our knowledge the largest document-image pretraining corpus, comprising 113 million images spanning 17 languages. Second, we propose a pretraining strategy that jointly learns image-to-text generation and pixel-level document reconstruction: the former aligns visual representations with textual content, while the latter preserves character strokes and layout details. Extensive experiments are conducted on five representative document analysis tasks, including text recognition, formula recognition, text detection, document tampering detection, and overlapping text segmentation. Replacing the original encoders with MonkeyOCRv2 consistently improves performance across all five tasks. Finally, we validate its effectiveness as the vision encoder of multimodal large language models on the more challenging tasks of document parsing and document understanding. Kept frozen and paired with a lightweight language model, it yields a 0.7B document parsing model that sets a new open-source state-of-the-art on MDPBench, a recent benchmark spanning digital-born and photographed documents across 17 languages, surpassing the previous best 3B dots.mocr by 2.8% absolute with a vision encoder roughly 11$\times$ smaller. The frozen encoder also powers a document understanding model that outperforms counterparts built on CLIP, DINO, and SAM across eight benchmarks under identical training settings. These results suggest that document-oriented visual pretraining can serve as a foundation for document intelligence in its own right.

cs.CV

When and How to Ask: Dynamic Preference Elicitation Strategies for Conversational Recommendation

Conversational Recommender Systems (CRSs) are interactive systems that use multi-turn natural language dialogue to understand evolving user preferences and provide personalized recommendations. To achieve this goal, CRSs rely on preference elicitation strategies to actively gather informative preference cues from users; however, the timing and selection of these strategies during a conversation remain largely unexplored. While many existing studies emphasize eliciting explicit item attributes and tend to adopt relatively static elicitation strategies, the use of item-based preference elicitation and how it varies across different dialogue stages remains less explored. In this work, we conduct a systematic investigation of preference elicitation strategies from a stage-aware perspective. We provide empirical evidence that optimal preference elicitation strategies are stage-dependent and context-sensitive: attribute-based inquiries are effective in early stages, while item-based strategies become superior as preferences refine. To support this paradigm, we introduce InPE, a dataset enriched with fine-grained annotations for elicitation necessity and strategy selection. With this dataset, we propose COPE (COnversational Preference Elicitation via Mixture of Experts), a novel architecture for strategy modeling. Extensive offline evaluation on our dataset indicates that context-aware preference elicitation strategies are beneficial for conversational recommendation. In addition, the analysis of the predicted strategies uncovers consistent stage-wise tendencies in dialogue progression, providing empirical evidence of common interaction patterns in conversational recommendation systems. Our dataset is available at https://github.com/juanfacabian/InPE.

cs.IR

Penalty-Free Natural Deep Ritz Method Based on de Rham Complex for High-Dimensional Dirichlet Boundary Value Problems

Deep neural networks show great promise for high-dimensional PDEs, yet enforcing essential boundary conditions remains challenging, especially as penalty parameters require problem-specific retuning with increasing dimensionality. In this work, we extend the Natural Deep Ritz Method (NatDRM) [H. Yu and S. Zhang, J. Comput. Phys., 537 (2025)] to a unified framework for all dimensions $d \geq 2$ based on the de Rham complex and its penalty-free boundary decomposition: curl-type operators act on scalar potentials in 2D, vector potentials in 3D, and antisymmetric second-order tensor potentials in $d \geq 4$, respectively. This method converts Dirichlet constraints into three coupled natural (Neumann-type) subproblems with corresponding Ritz-type losses, eliminating the need for a boundary penalty parameter $\beta$. We derive dimension-unified discrete losses, lightweight boundary-based gauge-fixing regularizations to resolve curl-kernel non-uniqueness, and a joint training procedure; extensions to variable-coefficient elliptic and semilinear Poisson problems are formulated at the first subproblem level. Numerical experiments on smooth benchmarks up to 6D show that NatDRM, without any penalty tuning, matches or exceeds the accuracy of optimally tuned DRM and PINN in most cases. It converges stably in 6D where penalized DRM fails for most penalty values, and exhibits synchronous decay of interior and boundary errors, resolving the inherent imbalance of penalty-based methods.

math.NA

Primal finite element scheme of the Hodge-Laplace problem

In this paper, we construct nonconforming finite element spaces $\boldsymbol{V}^{\mathbf{d}\cap\mathring{\boldsymbol{\delta}}}_h\Lambda^k$ for the approximation of $H\Lambda^k\cap H^*_0\Lambda^k$ on simplicial meshes, for $n\ge 2$ and $1\le k\le n-1$, by enforcing adjoint continuity against piecewise Whitney spaces rather than trace matching. It holds, with $\mathbf{d}^k_h$ and $\boldsymbol{\delta}_{k,h}$ denoting respectively the piecewise action of differential and codifferential operators, and $\boldsymbol{\mathfrak{H}}_h\Lambda^k$ being the discrete harmonic forms in the FEEC sense, that $\boldsymbol{\mathfrak{H}}_h\Lambda^k=\{\boldsymbol{\mu}_h\in \boldsymbol{V}^{\mathbf{d}\cap\mathring{\boldsymbol{\delta}}}_h\Lambda^k:\mathbf{d}^k_h\boldsymbol{\mu}_h=0,\ \boldsymbol{\delta}_{k,h}\boldsymbol{\mu}_h=0\}$, which mirrors the continuous Hodge--Laplace kernel on domains with nontrivial topology. The space is not a classical Ciarlet-type finite element space; though, a uniform discrete Poincare inequality and locally supported basis functions (supported on at most two cells) are guaranteed. The resulting primal scheme yields an $O(h)$ error bound for smooth data and $O(h^s)$ on $s$-regular domains ($0<s\le 1$), nontrivial topology admitted. Two- and three-dimensional eigenvalue tests agree with the mixed method on perforated domains, which are given to verify the validity of the scheme.

math.NA

Precision-Aware Illumination-Disentangled Vision Transformer for Spacecraft 6D Pose Estimation

Vision sensors provide a lightweight solution for spacecraft proximity operations, but monocular spacecraft 6D pose estimation remains difficult under illumination variation, specular reflection, shadowing, weak texture, and background interference. These factors make local visual evidence spatially unreliable and can destabilize pose regression. This article proposes a Precision-Aware Illumination-Disentangled Vision Transformer (PAID-ViT) for robust spacecraft pose estimation.The proposed model separates pose-relevant structure tokens from illumination-sensitive appearance tokens, estimates patch reliability before pose aggregation, and uses foreground mask supervision to preserve silhouette cues. A parameter-free geometric recovery module converts normalized crop coordinates, log-depth, and a continuous 6D rotation representation into camera-frame rotation and translation. Experiments on SPEED+ V2, the SPEED+ validation/lightbox/sunlamp evaluation configuration used in this study, suggest that PAID-ViT reduces translation error and improves robustness in the challenging sunlamp domain, while ablation studies support the complementary roles of illumination disentanglement, reliability-aware token aggregation, mask supervision, and training-side regularization.

cs.CV