SearcharxivSearch

arXiv subjects

Yanyan Zhang

Publications and source records attributed to Yanyan Zhang.

At least 19 recordsLinked to original sources

Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goal, scene configuration, or robot state. Existing recovery methods often require failure data, policy retraining, or external corrective agents, introducing additional data requirements and execution risks. We propose Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data. Upon detecting a deviation, CoRe imagines how the policy would continue toward the current goal from a recent viable state, using synthesized observations in place of physical execution, and then minimally realigns the robot and scene to rejoin this imagined continuation before returning control to the policy. Recovery is therefore planned without physical trial-and-error, preserves completed task progress, and handles both mid-episode instruction changes and physical perturbations in a unified manner. Extensive experiments across multiple simulators, VLA backbones, and real-world settings show that CoRe improves success rates by up to 85.0 percentage points to near-nominal levels while reducing physical restorations by 42.2%, without policy fine-tuning or failure-specific recovery training.

cs.RO

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

Large language models can solve harder reasoning problems with more inference-time compute. The term "test-time scaling," however, covers several inference algorithms: extending deliberation along one trajectory, sampling completed candidates and aggregating them by voting or verification, and searching over partial states. These algorithms differ in statistical structure, compute requirements, and failure modes. Treating them as interchangeable under a scalar "budget," or reporting accuracy without specifying the inference protocol, makes results difficult to compare across studies. We study test-time scaling along three axes. First, we formalize it as budgeted inference over the implicit prefix tree of an autoregressive model and distinguish single-trajectory sequential scaling, leaf-level scaling with terminal reduction, and prefix-level scaling. Second, we treat the full inference system as the evaluated object and separate end-to-end performance from candidate-bank diagnostics. We introduce an evaluation profile whose coordinates and simple functionals recover or bound common repeated-sampling metrics, and require compute accounting and uncertainty estimates that match the protocol. Third, we distinguish exact replay from distributional reproducibility and state the requirements for each. We also organize open-weight reasoning models by model-side and interface mechanisms. Our empirical study covers broad knowledge, symbolic reasoning, and competition mathematics, and we publicly release 1,403,520 sampled model attempts. The project website is available at https://mohsenhariri.github.io/scorio/tts. The released datasets are Trace (https://huggingface.co/datasets/harimo/scorio-trace), Lite (https://huggingface.co/datasets/harimo/scorio-lite), Math (https://huggingface.co/buckets/harimo/scorio-math), and SuperGPQA (https://huggingface.co/buckets/harimo/scorio-gpqa).

cs.LG

Towards Trustworthy Physical AI: From Theory to Practice Across Life Cycle

Physical AI refers to AI systems that understand, reason about, and act in accordance with the physical world and its underlying laws, dynamics, and constraints. Unlike conventional AI systems, physical AI interacts continuously with uncertain physical environments, and its actions produce consequences that are physically irreversible. As existing trustworthy AI frameworks have been developed primarily for digital AI systems, they do not fully capture the distinctive challenges of physical AI, such as physical safety, cyber-physical security, and physical manufacturing process. To address this gap, we present a survey of trustworthy physical AI principles. First, we characterize the core capabilities and challenges of physical AI. Second, we examine the role of physics in AI. Third, we trace the end-to-end physical AI life cycle across five core stages and introduce Trustworthy Physical AI Operationalization (T-PAIO). Fourth, we develop the Trustworthy Physical AI (T-PAI) framework, a theoretical framework that organizes key trustworthiness principles and provides a foundation for governing trustworthy physical AI systems.

cs.AI

STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models. However, existing RLVR methods typically rely on final-answer correctness to assign trajectory-level rewards, providing sparse supervision and treating all tokens uniformly regardless of their actual contribution to reasoning. Although recent studies introduce intermediate signals such as process rewards, high-entropy tokens, and semantic uncertainty, these signals are often not inherently verifiable and may fail to distinguish beneficial strategic patterns from harmful ones. To address this limitation, we propose STRIDE (Strategic Trajectory Reasoning with Discriminative Estimation), a fine-grained RLVR framework that derives strategic reasoning supervision from verifiable outcomes. STRIDE contrasts successful and failed trajectories within each response group to estimate the outcome-discriminative preference of each $n$-gram strategic pattern, and further combines this signal with reasoning saliency entropy to identify decision-relevant strategic patterns. These patterns are assigned differentiated advantage values during RL optimization, enabling more precise credit assignment while preserving the verifiability of RLVR. Extensive experiments demonstrate that STRIDE consistently improves reasoning performance across diverse models, tasks, and extended settings, including VLMs and agent-based systems.

cs.AI

CausalGuard: Conformal Inference under Graph Uncertainty

Estimating treatment effects from observational data requires choosing an adjustment set, but valid adjustment depends on an unknown causal graph. Graph misspecification can cause under-coverage, while graph-agnostic conformal wrappers may regain nominal coverage only through large padding. We introduce CausalGuard, a structure-weighted conformal framework that calibrates after aggregating graph-conditional doubly robust pseudo-outcomes. Candidate DAGs are proposed from an LLM-derived edge prior, pruned by conditional-independence tests, and reweighted by Bayesian Information Criterion. A composite nonconformity score then calibrates the posterior-weighted pseudo-outcome. CausalGuard provides distribution-free finite-sample marginal coverage for this aggregated pseudo-outcome; under causal identification, overlap, conditional-mean nuisance stability, and concentration on target-aligned valid adjustment strategies, its conditional mean converges to the true Conditional Average Treatment Effect. Across five benchmarks, CausalGuard attains mean coverage above the nominal 90% level for the directly evaluable target and reduces width when graph-agnostic conformal baselines require large padding. Stress tests show that CausalGuard suppresses invalid collider adjustment and remains stable under misspecified priors when the retained candidate set is data-supported.

cs.LG

Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models

Vision-Language-Action (VLA) models achieve remarkable flexibility and generalization beyond classical control paradigms. However, most prevailing VLAs are trained under a single-frame observation paradigm, which leaves them structurally blind to temporal dynamics. Consequently, these models degrade severely in non-stationary scenarios, even when trained or finetuned on dynamic datasets. Existing approaches either require expensive retraining or suffer from latency bottlenecks and poor temporal consistency across action chunks. We propose Pace-and-Path Correction, a training-free, closed-form inference-time operator that wraps any chunked-action VLA. From a single quadratic cost, joint minimization yields a unified solution that decomposes orthogonally into two distinct channels. The pace channel compresses execution along the planned direction, while the path channel applies an orthogonal spatial offset, jointly absorbing the perceived dynamics within the chunk window. We evaluate our approach on a comprehensive diagnostic benchmark MoveBench designed to isolate motion as the sole controlled variable. Empirical results demonstrate that our framework consistently outperforms state-of-the-art training-free wrappers and dynamic-adaptive methods and improves success rates by up to 28.8% and 25.9% in absolute terms over foundational VLA models in dynamic-only and static-dynamic mixed environments, respectively.

cs.RO

Well-posedness and asymptotic behavior of solutions to a second order nonlocal parabolic MEMS equation

We consider a second-order nonlocal parabolic MEMS equation with Dirichlet boundary conditions: \[ u_t-\Delta u=\frac{\lambda}{(1-u)^2\bigl(1+\int_\Omega\frac{1}{1-u}\,dx\bigr)^2},\quad x\in\Omega,\ t>0, \] where \(\Omega\subset\mathbb{R}^N\) \((1\le N\le3)\) is a bounded smooth domain and \(\lambda>0\). Using operator semigroups and the contraction mapping principle, we prove local existence and give a quenching criterion. Under suitable smallness conditions on \(\lambda\) and the initial data, global existence and exponential convergence to the minimal steady state are obtained. Assuming the global solution stays uniformly away from the singularity \(u=1\), we show that the system forms a gradient system. By establishing analyticity of the energy and a Lojasiewicz--Simon inequality, we prove that the solution converges to a steady state with either exponential or algebraic rate depending on the Lojasiewicz exponent. Numerical experiments in 1D and 2D illustrate the results and support conjectures on the \(\lambda\)-dichotomy.

math.AP

Asymptotic Behaviors of Global Solutions to Fourth-order Parabolic and Hyperbolic Equations with Dirichlet Boundary Conditions

This paper investigates the asymptotic behaviors of global solutions to fourth-order parabolic and hyperbolic equations with Dirichlet boundary conditions. The equations model Micro-Electro-Mechanical Systems (MEMS) and are depending on a positive voltage parameter $\lambda$. We establish the convergence of global solutions to an equilibrium, along with the convergence rate estimates. Supporting numerical simulations are presented.

math.AP

Low-Energy, Octave-Spanning Supercontinuum Generation in Ta_2O_5 Waveguides: Towards Optical Coherence Metrology

Supercontinuum generation (SCG) on integrated photonic platforms is a pivotal technology for developing next-generation chip-scale systems for precision spectroscopy and metrology. While significant progress has been made with silicon (Si) and silicon nitride ($Si_3N_4$) platforms, they are often constrained by two-photon absorption (TPA) or moderate nonlinear coefficients, necessitating a trade-off between energy efficiency and bandwidth. Tantalum pentoxide ($Ta_2O_5$), possessing both high nonlinearity and a wide bandgap, emerges as a promising candidate; however, current implementations remain challenged by high pump energy consumption. Here, we report a low-loss $Ta_2O_5$ integrated waveguide fabricated via the Damascene process. It enables the generation of a two-octave-spanning spectrum with a low pulse energy of only 92.9 pJ (60 fs, 1550 nm). Notably, the corresponding peak power is a mere 1.36 kW, which is nearly an order of magnitude lower than that of state-of-the-art comparable broadband sources. Furthermore, at the maximum pump energy, our spectrum exhibits an ultrabroad coverage from 450 nm to 3400 nm, spanning nearly three octaves. Supported by numerical simulations, we analyze the dynamics of soliton fission. Furthermore, a Michelson interferometry system developed using this source exhibits superior performance, achieving not only micrometer-scale axial resolution but also a 6 dB sensitivity roll-off length of 3.1 mm. This exceptional roll-off performance, combined with a displacement measurement sensitivity of 346 nm, underscores the immense potential of the $Ta_2O_5$ platform for applications in biomedical imaging and precision metrology.

physics.optics

Analytic description of the moving moisture front in soils

The fact that moisture propagates in soils at a finite speed is confirmed by natural everyday experience as well as by controlled laboratory tests. In this text, we rigorously derive analytical upper bounds for the speed of moisture front propagation under gravity for the solution to the Richards equation with compactly supported initial data. The main result is an explicit criterion describing a competition between gravity and capillarity, where the dominant effect is determined by the characteristics of the soil. If capillarity prevails, the initially wet regions remain wet for all times, while if gravity is dominant, moisture travels downward at a speed that is asymptotically bounded from below and above. As a by-product, we prove the existence and uniqueness of a solution to an initial value problem for the degenerate Richards equation on the whole space. Numerical simulations based on the proposed model confirm the theoretical predictions, with results that closely match experimental observations.

math.AP

On-chip quadratically nonlinear photodetector

Involving deterministically nonlinear photoresponse in on-chip photodetector is intriguing to develop sophisticated functions in photonic integrated circuits, such as in-sensor computing and optoelectronic mixing, though the corresponding devices are still lack of sufficient investigation. Here, we demonstrate an on-chip quadratically nonlinear photodetector (QNPD) by configuring an InSe p-i-n homojunction on a silicon waveguide. Telecom-band light guiding in the waveguide couples with the InSe evanescently and is frequency up-converted into visible light via InSe's second-harmonic generation (SHG), which is subsequently absorbed by InSe and finally generates photocurrent under the built-in electric field of the p-i-n homojunction. Governed by these sequential processes, the on-chip QNPD presents a quadratic function between photocurrent and optical power. Thanks to the efficient SHG and well-established homojunction in InSe, the QNPD reaches a high normalized responsivity of 37.1 A/W2 and low dark current of 1 pA, representing greatly improved performances among reported nonlinear photodetectors. Benefiting from the extra SHG process, the on-chip QNPD intrinsically incorporates light-light interactions, enabling straightforwardly monitoring all-optically mixing signals electrically. As an example, an array of 16-pixel QNPDs was designed to implement a fully single-shot on-chip autocorrelator without requirement of bulky optics and external cameras, which precisely measures picosecond pulses with high sensitivity of 6.1*10-10 W2.

physics.optics

Keyword search is all you need: Achieving RAG-Level Performance without vector databases using agentic tool use

While Retrieval-Augmented Generation (RAG) has proven effective for generating accurate, context-based responses based on existing knowledge bases, it presents several challenges including retrieval quality dependencies, integration complexity and cost. Recent advances in agentic-RAG and tool-augmented LLM architectures have introduced alternative approaches to information retrieval and processing. We question how much additional value vector databases and semantic search bring to RAG over simple, agentic keyword search in documents for question-answering. In this study, we conducted a systematic comparison between RAG-based systems and tool-augmented LLM agents, specifically evaluating their retrieval mechanisms and response quality when the agent only has access to basic keyword search tools. Our empirical analysis demonstrates that tool-based keyword search implementations within an agentic framework can attain over $90\%$ of the performance metrics compared to traditional RAG systems without using a standing vector database. Our approach is simple to implement, cost effective, and is particularly useful in scenarios requiring frequent updates to knowledge bases.

cs.IR

Asymptotic Behavior of Rupture Solutions for the Elliptic MEMS Equation with H\'enon-Type and External Pressure Terms

This paper investigates an elliptic MEMS-Type equation with Henon and external pressure terms: Delta u = lambda|x|^alpha / u^p + F for x in R^N \ {0}, with u(0)=0 and u>0 for x in R^N \ {0}, where N >= 1, lambda > 0, p > 0, alpha > -2 and F in R are constants. We study positive rupture solutions with rupture point at the origin (u(0)=0). Our main emphasis is on asymptotic radial rupture solutions: we prove the existence of both radial and non-radial solutions, characterize their asymptotic behavior near the origin, and obtain a full asymptotic expansion of arbitrary order.

math.AP

NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?

The evaluation of Vision-Language-Action (VLA) agents is hindered by the coarse, end-task success metric that fails to provide precise skill diagnosis or measure robustness to real-world perturbations. This challenge is exacerbated by a fragmented data landscape that impedes reproducible research and the development of generalist models. To address these limitations, we introduce NEBULA, a unified ecosystem for single-arm manipulation that enables diagnostic and reproducible evaluation. NEBULA features a novel dual-axis evaluation protocol that combines fine-grained capability tests for precise skill diagnosis with systematic stress tests that measure robustness. A standardized API and a large-scale, aggregated dataset are provided to reduce fragmentation and support cross-dataset training and fair comparison. Using NEBULA, we demonstrate that top-performing VLAs struggle with key capabilities such as spatial reasoning and dynamic adaptation, which are consistently obscured by conventional end-task success metrics. By measuring both what an agent can do and when it does so reliably, NEBULA provides a practical foundation for robust, general-purpose embodied agents.

cs.RO

New measurement method for weak magnetic fields using magnetically induced deformation of chemical bonds in Co-CsPbBr3 quantum dots

The research on weak magnetic field detection is of great significance in advancing the development of bioscience, aerospace, chip manufacturing and other fields. However, the weak magnetic detecting still face some problems, including the large size of the detectors and the limited detection scale. To contribute to the detection of weak magnetic fields, the Co-CsPbBr3 colloidal quantum dots (QDs) composite magnetic material was synthesised on the basis of the theory of room temperature ferromagnetism, molecular polarisation and vibration level of chemical bond. The synthesis involved the mixing of Co2+ into CsPbBr3, an all-inorganic perovskite with activated ions. Subsequently, a weak magnetic field measurement system was devised, comprising working medium samples and a vibration level detection optical path. Following the acquisition, comparison, processing and analysis of multiple data sets, a Stokes displacement function model was established under different magnetic field sizes and the weak magnetic field intensity range of Pitsla (pT) was measured. The Pitsla weak magnetic field measurement system proposed in this paper provides a reference for the development of non-contact weak magnetic measurement methods and for the advancement of intelligent and low-dimensional weak signal measurement applications.

physics.optics

Multi-watt 1 GHz single-cycle frequency combs

Single-cycle optical pulses offer a strong carrier-envelope-offset (CEO) dependent electric field and the highest peak intensity for a given pulse energy. Absence of demonstrated GHz single-cycle lasers constrains exploration of single/sub-cycle dynamics at this repetition rate. By leveraging fiber soliton effects and suppressing higher-order dispersion, we achieve single-cycle pulse generation at a 1 GHz repetition rate in an all-fiber format. The laser produces 7.1 fs (1.1-cycle) pulses with 1.8 W average power, centered around 1970 nm. Temporal characterization shows 60% of the pulse energy is concentrated in the pulse center, yielding a peak power of 110 kW. The seed laser demonstrates a 43 dB signal-to-noise ratio for the CEO frequency, facilitating comb stabilization and CEO control. Our model, which matches experimental observations, identifies conditions for achieving single-cycle duration and predicts scalability to a 2 GHz repetition rate. This work presents the first GHz single-cycle source. We envision these advances will drive studies in single/sub-cycle light-matter interaction, spectroscopy, microscopy, and CEO-sensitive nonlinear optics.

physics.optics

Two-octave frequency combs from all-silica-fiber implementation

Mid-infrared frequency comb spectroscopy enables measurement of molecular at megahertz spectral resolution, sub-hertz frequency accuracy and microsecond acquisition speed. However, the widespread adoption of this technique has been hindered by the complexity and alignment sensitivity of mid-infrared frequency comb sources. Leveraging the underexplored mid-infrared window of silica fibers presents a promising approach to address these challenges. In this study, we present the first experimental demonstration and quantitative numerical description of mid-infrared frequency comb generation in silica fibers. Our all-silica-fiber frequency comb spans over two octaves (0.8 $\mu$m to 3.5 $\mu$m) with a power output of 100 mW in the mid-infrared region. The amplified quantum noise is suppressed using four-cycle (25 fs) driving pulses, with the carrier-envelope offset frequency exhibiting a signal-to-noise ratio of 40 dB and a free-running bandwidth of 90 kHz. Our developed model provides quantitative guidelines for mid-infrared frequency comb generation in silica fibers, enabling all-fiber frequency comb spectroscopy in diverse fields such as organic synthesis, pharmacokinetics processes, and environmental monitoring.

physics.optics

Microcavity induced by few-layer GaSe crystal on silicon photonic crystal waveguide for efficient optical frequency conversion

We demonstrate the post-induction of high-quality microcavity on silicon photonic crystal (PC) waveguide by integrating few-layer GaSe crystal, which promises highly efficient on-chip optical frequency conversions. The integration of GaSe shifts the dispersion bands of the PC waveguide mode into the bandgap, resulting in localized modes confined by the bare PC waveguides. Thanks to the small contrast of refractive index at the boundaries of microcavity, it is reliably to obtain quality (Q) factors exceeding 10^4. With the enhanced light-GaSe interaction by the microcavity modes and high second-order nonlinearity of GaSe, remarkable second-harmonic generation (SHG) and sum-frequency generation (SFG) are achieved. A record-high on-chip SHG conversion efficiency of 131100% W^-1 is obtained, enabling the clear SHG imaging of the resonant modes with the pump of sub-milliwatts continuous-wave (CW) laser. Driven by a pump of on-resonance CW laser, strong SFGs are successfully carried out with the other pump of a CW laser spanning over the broad telecom-band. Broadband frequency conversion of an incoherent superluminescent light-emitting diode with low spectral power density is also realized in the integrated GaSe-PC waveguide. Our results are expected to provide new strategies for high-efficiency light-matter interactions, nonlinear photonics and light source generation in silicon photonic integrated circuits.

physics.optics