SearcharxivSearch

arXiv subjects

Jiawei He

Publications and source records attributed to Jiawei He.

At least 19 recordsLinked to original sources

Solutions to Three Conjectures and an Open Problem on Binary BCH Codes

BCH codes are among the most important classes of cyclic codes and have played a central role in coding theory and its applications. One of the fundamental problems in the study of BCH codes is to determine their exact minimum distances, which directly govern their error-correcting capability. Although the BCH bound provides a general lower bound, determining the exact minimum distance is often difficult, and many parameter families remain unresolved. In this paper, we investigate three conjectures and an open problem on binary BCH codes proposed by Chen, Xie, and Ding in \cite{Chen59}. We settle these conjectures on the exact minimum distances of three families of binary BCH codes by constructing codewords attaining the BCH bound. Beyond these conjectures, we further study the more difficult family codes and determine its minimum distance for some cases. We further study Open Problem 8.4: affirmative answers are obtained for the first two length families, while for the third family a sufficient condition is established and a counterexample shows that the unrestricted assertion does not hold in general.

cs.IT

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.

cs.RO

The Sharp Spectral Transition for Almost Mathieu Operators via Alternating Resonances

We prove that for any given frequency resonance exponent and phase resonance exponent, there exist a frequency and a phase exactly realizing these values such that the Almost Mathieu operator exhibits Anderson localization for $\ln|\lambda|>\max\{\beta(\alpha),\delta(\alpha,\theta)\}$. This resolves the conjecture in \cite{MR4686650}.

math-ph

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels

Coding agents are increasingly evaluated not only by whether they solve a task, but also by how they execute it. However, existing process-level evaluations often treat action prediction, task uncertainty, and step attribution as if they were the same problem, which makes it unclear what such evaluations actually measure. In this paper, we introduce a measurement framework for process evaluation in coding agents and instantiate step-level causal attribution with SCAE, a replay-based estimator derived from a structural causal model of agent execution. Our framework combines prefix-conditioned identification, replay/intervention-based estimation, and controlled judge-information manipulation to study process evaluation at the action, task, and step levels. Experiments on 499 file-localization episodes from 12 repositories show that next actions are driven primarily by execution provenance rather than code-graph transitions, execution uncertainty is structured at the task rather than step level, and full-trace judges exhibit systematic collider bias, suggesting that current process evaluation often measures semantic relevance rather than certified causal contribution.

cs.AI

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained $\pi_{0.5}$ instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.

cs.RO

DA-Nav: Direction-Aware City-Scale Vision-Language Navigation

City-scale outdoor navigation is currently hindered by the heavy reliance on dense maps or costly navigation supervision. In this work, we introduce a novel paradigm for leveraging directional instructions from commercial navigation tools (e.g., Google Maps). To bridge the gap between commercial instructions and executable navigation actions, while mitigating long-horizon error accumulation through robust trajectory recovery, we propose DA-Nav, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane. To achieve trajectory recovery, DA-Nav employs a Chain-of-Thought (CoT) reasoning process encompassing deviation assessment, action prediction, and target grid selection. We further introduce ReDA, a dataset that provides direction-aware instructions and recovery trajectories to enhance spatial grounding and support CoT recovery reasoning. Extensive experiments in CARLA demonstrate that DA-Nav achieves a high success rate of 56.16% in unseen urban environments, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability. Furthermore, without fine-tuning, DA-Nav seamlessly adapts to both quadruped and humanoid robots, enabling stable kilometer-scale closed-loop outdoor navigation in complex real world environments.

cs.RO

From Fragments to Paths: Task-Level Context Recovery for Large Industrial Codebases

Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often retrieve only local fragments and fail to recover the broader task-relevant context needed for complex repository-level tasks. We present DeepDiscovery, a task-level repository-understanding method for large industrial codebases. DeepDiscovery uses a two-stage \textit{Location--Inference} framework to localize high-confidence task anchors and recover broader task-relevant context over multi-relational repository structure under budget constraints. Across controlled method-level evaluation, organization-internal industrial repository-understanding scenarios, and end-to-end evaluation on SWE-bench Verified, DeepDiscovery consistently improves task-relevant file recovery and downstream SE performance. On 27 medium-scale tasks, DeepDiscovery achieves the best file recovery quality among five representative baselines without offline preprocessing. On organization-internal industrial tasks from a production-scale integrated codebase ecosystem, including 27 medium-scale tasks and 40 large-scale tasks, DeepDiscovery improves Full Recall Rate across multiple AI coding systems, with absolute gains ranging from 1.6 to 9.2 percentage points on large subprojects and from 2.5 to 7.4 percentage points on medium-scale subprojects. In a controlled end-to-end evaluation on SWE-bench Verified, a system equipped with DeepDiscovery achieves a 78.6\% Solve Rate, outperforming the corresponding baseline by 8.2 percentage points. These results suggest that stronger task-level repository understanding can improve coding-agent performance on complex SE tasks.

cs.SE

Cantor Spectrum via a Reducibility-Duality Bridge for the Mosaic Almost Mathieu Operator

We study the mosaic Almost Mathieu operator, a quasiperiodic model that naturally admits a singular strip-Jacobi representation. By establishing a duality framework and extending the correspondence between the integrated density of states and the fibered rotation number to this setting, we obtain an effective reduction to $SL(2,\mathbb{R})$ cocycles. As a consequence, combining Aubry duality, reducibility theory, and the Moser--P\"oschel argument, we prove that the spectrum is a Cantor set for all noncritical parameters.

math-ph

SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction

Joint Entity and Relation Extraction (JERE) is highly sensitive to training data quality, making data augmentation a natural way to improve generalization. However, existing augmentation methods often weaken entity relevance and disrupt semantic structure, limiting their effectiveness for JERE. In this paper, we propose \textbf{Structured Semantic Data Augmentation (SSDAU)}, a method designed to preserve triple-aware semantic structure during augmentation. SSDAU segments text by entity labels, captures semantic features through context-aware encoding, and restructures entity semantics to generate augmented data. To distinguish semantically similar entities, SSDAU combines contextualized embeddings with traditional similarity scores. To reduce topic inconsistency, we apply BERTopic-based filtering to remove irrelevant augmentations. We evaluate SSDAU on datasets with different annotation types and compare its performance on five representative JERE models against seven popular augmentation baselines. Experiments show that SSDAU generates semantically consistent data, is more robust to ambiguity than non-LLM methods (8.95\% vs. 23.58\% average relative F1 decrease), and significantly outperforms strong alternatives in most settings.

cs.CL

ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Existing benchmarks for LLM coding agents primarily evaluate final outcomes. While useful for measuring overall capability, these metrics provide limited visibility and often miss defects that arise during execution. We present ProcCtrlBench, a benchmark for execution-process evaluation in LLM coding agents. ProcCtrlBench organizes recurrent execution defects into a reusable ontology covering 11 defect types in 4 categories, and evaluates agent trajectories through standardized process evidence rather than final outcomes alone. To support comparison across heterogeneous agents, ProcCtrlBench standardizes raw logs into a unified trajectory representation and reports calibrated scorecards over process-level findings. In addition, ProcCtrlBench uses control preservation as a way to quantify execution-process quality, capturing whether execution remains interpretable, interruptible, correctable, reversible, and able to hand back authority when needed. We evaluate ProcCtrlBench on 200 cases sampled from three benchmarks: AndroidBench, TerminalBench, and SWE-bench-Verified. Results show that ProcCtrlBench can be instantiated with useful reliability, provides more stable semantics than direct thresholding, and reveals meaningful differences in execution quality that are often overlooked by conventional outcome-based evaluation.

cs.SE

ICT-NLP at SemEval-2026 Task 3: Less Is More -- Multilingual Encoder with Joint Training and Adaptive Ensemble for Dimensional Aspect Sentiment Regression

This paper describes our system to SemEval-2026 Task 3 Track A Subtask 1 on Dimensional Aspect Sentiment Regression (DimASR). We propose a lightweight and resource-efficient system built entirely on multilingual pre-trained encoders, without relying on LLMs or external corpora. We adopt joint multilingual and multi-domain training to facilitate cross-lingual transfer and alleviate data sparsity, introduce a bounded regression transformation that improves training stability while constraining predictions within the valid range, and employ an adaptive ensemble strategy via subset search to reduce prediction variance. Experimental results demonstrate that our system achieves strong and consistent performance, ranking 1st on zho-res, 2nd on zho-lap, and 3rd on jpn-hot, with all remaining datasets placed within the top half of participating teams.

cs.CL

Non-Hermitian Anomalous Scaling Engineering

Non-Hermitian systems exhibit anomalous scaling, a striking departure from conventional bulk laws, rooted in the non-Hermitian skin effect (NHSE). Here, we experimentally uncover this scaling and demonstrate its active control in a temporal photonic lattice. By tracking the real-time evolution of all eigenstates as system size varies, we directly observe scaling-driven spectral reshaping and eigenstate localization, revealing phenomena absent in Hermitian or NHSE-free lattices. In a Su-Schrieffer-Heeger lattice, scaling alone can trigger a non-Hermitian topological phase transition, with edge modes remaining protected. Crucially, Kerr interactions open the frontier of nonlinear non-Hermitian physics: weak nonlinearity accelerates or decelerates anomalous scaling, while strong nonlinearity suppresses it entirely. These results establish the first experimental platform for linear and nonlinear anomalous scaling engineering, paving the way for compact non-Hermitian devices and exploration of nonlinear and many-body non-Hermitian phenomena.

physics.optics

FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation

Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to reason about spatial structures. Although human-readable spatial schematics (e.g., floor plans) are ubiquitous in real-world buildings, current agents lack the cognitive ability to comprehend and utilize them. To bridge this gap, we introduce \textbf{FloorPlan-VLN}, a new paradigm that leverages structured semantic floor plans as global spatial priors to enable navigation with only concise instructions. We first construct the FloorPlan-VLN dataset, which comprises over 10k episodes across 72 scenes. It pairs more than 100 semantically annotated floor plans with Matterport3D-based navigation trajectories and concise instructions that omit step-by-step guidance. Then, we propose a simple yet effective method \textbf{FP-Nav} that uses a dual-view, spatio-temporally aligned video sequence, and auxiliary reasoning tasks to align observations, floor plans, and instructions. When evaluated under this new benchmark, our method significantly outperforms adapted state-of-the-art VLN baselines, achieving more than a 60\% relative improvement in navigation success rate. Furthermore, comprehensive noise modeling and real-world deployments demonstrate the feasibility and robustness of FP-Nav to actuation drift and floor plan distortions. These results validate the effectiveness of floor plan guided navigation and highlight FloorPlan-VLN as a promising step toward more spatially intelligent navigation.

cs.RO

Admissibility approach to nonuniform exponential dichotomies roughness with nonlocal perturbations

Nonuniform exponential dichotomy serves as an important characteristic of nonuniform hyperbolicity, while admissibility of function classes is often used to characterize nonuniform exponential dichotomy. In this paper, we investigate the preservation of nonuniform exponential dichotomy under certain nonlocal perturbations. By utilizing the concept of admissibility of a pair of function classes, we establish sufficient conditions to ensure that the dichotomy results are consistent with those in the homogeneous situation. These results need to satisfy a smallness integrability condition.

math.DS

Modular Expert Merging for Biomedical Retrieval

Adapting general-purpose LLMs into domain-specialized dense retrievers typically requires large-scale training on mixed-domain data. We show that merging independently trained domain-specialized experts consistently exceeds this approach across four decoder-only LLM families (0.6B-7B), four merging methods, and twelve medical and general retrieval tasks from MTEB, suggesting that parameter-space composition captures complementary domain strengths that large-scale mixed-domain training averages out. To further maximize expert quality, we introduce Synthesize-Train-Merge (STM), a modular framework that synthesizes hard negatives with a top-tier LLM and fine-tunes domain-specialized experts via LoRA before merging them, without continual pre-training. Synthesized hard negatives yield the largest gains for smaller models, and STM achieves strong performance on biomedical retrieval tasks while maintaining competitive general-domain results across all four backbone families.

cs.CL

Permutation groups and symmetric Hecke algebras

The endomorphism algebras of the permutation modules for transitive permutation groups, known as Hecke algebras, are fundamental objects in representation theory. While group algebras are known to be symmetric over any field, it is natural to ask whether this property extends to Hecke algebras. To study this, we introduce the new concepts of $p$-$S$-permutation groups (for a prime $p$) and $S$-permutation groups. A \emph{ $p$-$S$-permutation group} is a transitive permutation group whose associated Hecke algebra is symmetric over every field of characteristic $p$. An \emph{ $S$-permutation group} is a transitive permutation group that is a $p$-$S$-permutation group for all primes $p$. In this paper, we study Hecke algebras from a group-theoretical perspective and we show that several classes of permutation groups are $p$-$S$-permutation groups and $S$-permutation groups in our sense. This result represents a substantial extension of earlier work by Li and He. (Transform Groups, 30(4), 2025), and reframes the question of determining when the algebra \(\End_{KG}(K\Omega)\) is symmetric within a more general theoretical framework.

math.RT

Approximation for stochastic time-space fractional cable equations driven by rough noise

The time-space fractional cable equation arises from extending the generalized fractional Ohm's law to model anomalous diffusion processes. In this paper, we develop and analyze a numerical approximation for stochastic nonlinear time-space fractional cable equation driven by rough noise. The model involves both two nonlocal terms in time and one in space. By an operator theoretic approach, we establish the existence, uniqueness, and regularities of solutions. We also obtain a convergence result for the regularized equation via Wong-Zakai approximation to regularize the rough noise. The numerical scheme approximates the model in space by the standard spectral Galerkin method and in time by the backward Euler convolution quadrature method. After that, error estimates are established.

math.NA

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spatial referring and real-world metric measurement. However, existing methods struggle with this compositional task. To this end, we propose RoboTracer, a 3D-aware VLM that first achieves both 3D spatial referring and measuring via a universal spatial encoder and a regression-supervised decoder to enhance scale awareness during supervised fine-tuning (SFT). Moreover, RoboTracer advances multi-step metric-grounded reasoning via reinforcement fine-tuning (RFT) with metric-sensitive process rewards, supervising key intermediate perceptual cues to accurately generate spatial traces. To support SFT and RFT training, we introduce TraceSpatial, a large-scale dataset of 30M QA pairs, spanning outdoor/indoor/tabletop scenes and supporting complex reasoning processes (up to 9 steps). We further present TraceSpatial-Bench, a challenging benchmark filling the gap to evaluate spatial tracing. Experimental results show that RoboTracer surpasses baselines in spatial understanding, measuring, and referring, with an average success rate of 79.1%, and also achieves SOTA performance on TraceSpatial-Bench by a large margin, exceeding Gemini-2.5-Pro by 36% accuracy. Notably, RoboTracer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (UR5, G1 humanoid) in cluttered real-world scenes. Please see the project page at https://zhoues.github.io/RoboTracer.

cs.RO