SearcharxivSearch

arXiv subjects

Yao Wang

Publications and source records attributed to Yao Wang.

At least 19 recordsLinked to original sources

An Eccentric Massive Protobinary Assembled via a Core-merger Parabolic Encounter

Most massive stars form in binary systems, which profoundly influence their subsequent evolution. However, how such systems form remains poorly understood, with several competing scenarios proposed, including disk fragmentation, core fragmentation and capture. Determining the orbital architectures of massive binaries, particularly during their earliest embedded phases, is therefore crucial for distinguishing among these formation pathways, but direct measurements of their three-dimensional motions have remained exceptionally challenging. Here we present high-resolution, multi-epoch sub-millimeter-to-centimeter ALMA and JVLA observations of the massive protobinary IRAS 07299$-$1651, complemented by JWST and VLT infrared imaging. We detect orbital proper motion of the binary components, enabling a full three-dimensional orbital reconstruction. Combining orbital fitting, multi-wavelength continuum modelling, hydrogen recombination line kinematics and jet observations, we find that the preferred orbital solutions are highly eccentric and close to parabolic, while both circumstellar disks are strongly misaligned with the orbital plane. These properties are naturally explained by a ``core-merger'' scenario in which the two protostars originated independently from initially unbound cores that recently underwent a near-parabolic encounter, producing an eccentric binary with a current separation of about 200 au. These findings suggest that the core-merger process may represent an important pathway for forming eccentric massive binaries.

astro-ph.SR

FoRIS: Progressive Foreground Refinement for Training-Free In-Context Segmentation

In-Context Segmentation (ICS) aims to precisely segment arbitrary semantic concepts, such as objects or parts, given one or a few annotated visual exemplars. In this paper, we revisit ICS from a more classical segmentation perspective, viewing it as a coarse-to-fine progressive refinement process. Rather than directly predicting the final mask through reference-query matching, we progressively refine the segmentation from coarse and ambiguous foreground responses to precise and complete foreground structures. Building upon this perspective, we propose a training-free in-context segmentation framework, termed FoRIS. Specifically, FoRIS consists of three key stages: Foreground Purification, Foreground Localization, and Foreground Consolidation, which progressively suppress background distractions, localize discriminative target regions, and recover complete foreground structures through semantic aggregation. Experimental results demonstrate that FoRIS achieves SOTA performance across semantic and part segmentation tasks, with average improvements of 4.5 and 4.8 mIoU points over existing approaches in the 1-shot and 5-shot settings, respectively. Code: https://github.com/Xi-Mu-Yu/FoRIS.

cs.CV

Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs

Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit formats such as MXFP4 and HiF4, has accelerated research into efficient MLLM training and deployment. In this work, we present a systematic study of these quantization schemes in representative MLLMs that span both video generation and reasoning tasks. Our analysis shows that MXFP8 achieves near-lossless performance, whereas aggressive 4-bit quantization leads to significant degradation. Through extensive ablations, we identify activation quantization as the primary source of this performance loss, contributing substantially more than weight quantization. Motivated by this observation, we propose Residual Fallback Quantization (RFQ), a lightweight activation reconstruction framework that supplements the primary ulta-low-bit activation representation with an auxiliary quantized residual pathway. By explicitly modeling and compensating for quantization errors, RFQ improves activation fidelity while preserving the efficiency advantages of ultra-low-bit computation. RFQ requires no architectural modifications and incurs negligible computational overhead. Extensive experiments on Wan2.2 and Qwen3-VL demonstrate that RFQ consistently recovers a substantial portion of the performance lost under the quantization of MXFP4 and HiF4, significantly narrowing the gap to BF16 baselines across both generation and 4 reasoning benchmarks. Our findings establish activation quantization as the dominant bottleneck in ultra-low-bit MLLMs and highlight residual-based activation reconstruction as an effective and practical strategy for robust 4-bit deployment.

cs.LG

Error Propagation Theory for Variational Non-Markovian Open Quantum Dynamics

Variational approaches based on neural quantum states and physics-informed neural networks provide powerful paradigms for simulating non-Markovian open quantum dynamics. However, extending these methods into the strongly non-Markovian regime reveals a critical bottleneck: even minute errors in the time evolution can translate into substantial deviations in physical observables. The fundamental origin of this stringent precision requirement, as well as how non-Markovianity governs it, remains an open question. Here, we develop a theoretical framework that systematically characterizes error propagation in variational non-Markovian dynamics. By combining analytical derivations with numerical verification, we present the first quantitative description of variational error evolution over time. Our analysis uncovers an intrinsic error-backflow mechanism driven by long-lived environmental memory. This mechanism establishes a fundamental precision barrier and provides concrete guidance for designing robust variational algorithms for strongly non-Markovian quantum systems.

quant-ph

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured task and step representations; trust-state updates, projec- tions, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.

cs.AI

Coupled-cluster molecular properties across the main group that extrapolate beyond training size

Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.

physics.chem-ph

Transport-Noise Witnesses of Electronic Multipartite Entanglement

Entanglement among particles is a defining feature of strongly correlated quantum materials, distinguishing them from conventional metals and semiconductors. The ability to certify intrinsic entanglement among interacting electrons in solid-state materials is important not only for classifying quantum states of matter, but also for developing material-based quantum technologies. Here, we introduce a transport-based protocol for witnessing multipartite entangled electronic states, based on the equilibrium noise spectrum as an experimentally accessible observable. The appropriately integrated, symmetrized, and projected current noise obeys an upper bound that can be derived from microscopic model parameters and is invariant with respect to the choice of electronic basis. We benchmark this framework in several paradigmatic systems, including twisted bilayer graphene, twisted bilayer MoTe$_2$, and Hubbard models, certifying entanglement in the fractional Chern insulating state. The method extends recently developed scattering-based entanglement witnesses to ultralow-temperature materials, where conventional spectroscopic probes are inaccessible but candidate entangled states are expected to arise.

cond-mat.str-el

CrossTracer: Cross-Embodiment Navigation via VLA Model Reasoning and Trace Residuals Adapting

Vision-language-action (VLA) models provide strong semantic priors for robot navigation, but they often ignore embodiment-specific mobility constraints. A path that is semantically plausible for one robot may be physically infeasible for another. We propose CrossTracer, a hierarchical framework for cross-embodiment navigation through adaptive trace residuals. CrossTracer represents navigation plans as normalized image-plane waypoints, forming a unified pixel-space interface between semantic reasoning and physical grounding. First, Vision-Language Trace Proposer (VL-Tracer) adapts a pretrained VLA model to predict an initial navigation trace from egocentric observations and flexible goal specifications. Second, CE-Adapter refines this trace by predicting embodiment-conditioned residual corrections from visual traversability cues, robot identity, and the initial trace. To train the refinement module without costly manual annotation, Cross-Embodiment RRT* (CE-RRT*) converts panoptic segmentation into robot-conditioned traversability cost maps and generates cost-minimizing pixel-space traces. We evaluate CrossTracer on the NaviTrace benchmark, which tests whether a model can generate embodiment-consistent navigation traces from egocentric observations, language instructions, and robot embodiment types. CrossTracer achieves a total score of 45.68, outperforming the strongest evaluated general-purpose baseline, Gemini-2.5-Pro, by 10.01 points, corresponding to a 28.1% relative improvement. Real-world deployment on wheeled and legged robots further shows improved navigation success and execution efficiency.

cs.RO

Sublattice-resolved coherent phonon dynamics in charge density waves

Phonons govern fundamental material properties and play a central role in various electronic phase transitions. Coherent driving of specific phonon modes enables on-demand phase control, motivating sublattice-resolved identification of real-space phonon motions. Yet experimentally resolving these motions remains challenging, limiting precise phonon-based control. Here, we introduce a dynamical protocol to track element-resolved phonon dynamics in the charge density wave material EuTe4, in which the dominant Te-sublattice charge order is accompanied by a previously unreported Eu-sublattice component. We leverage the elemental selectivity of time-resolved resonant X-ray scattering to reveal three coherent phonon modes with distinct sublattice character, thereby disentangling Eu- and Te-dominated lattice dynamics, in good agreement with theoretical calculations of the phonon eigenvectors. This time-domain approach, which surpasses the energy-resolution limits of conventional frequency-domain inelastic scattering, provides a broadly applicable framework for decomposing coherent phonons in multi-element materials, which is crucial for the targeted control of phases of matter.

cond-mat.mtrl-sci

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout-training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul -- a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout's compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format -- whose three-level hierarchical scaling preserves resolution under FP4's tight 4-bit budget -- closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to-end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.

cs.LG

Stable FP4 Training via Transposition-Invariant Block Quantization

Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-bit oating point (FP4) remains challenging due to instability during optimization. We identify a fundamental source of this instability in existing microscaling approaches: scale inconsistency induced by tensor transposition. In conventional 1D block quantization, forward and backward passes assign di erent scaling factors to the same values after transposition, leading to biased and unstable gradient updates. To address this issue, we propose a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations. We further combine this with truncation-free scaling and stochastic rounding to control quantization error and maintain unbiased gradients. To handle the sensitivity of attention mechanisms, we adopt MXFP8 quantization for query and key projections, yielding a practical mixed-precision design. We evaluate our method on dense LLMs up to 7B parameters and a 30B Mixture-of-Experts model, trained on up to 100B tokens. Across all settings, our approach achieves stable end-to-end FP4 training and closely matches BF16 performance, with less than 1.3% degradation in perplexity and downstream accuracy. These results demonstrate that enforcing forwardbackward scaling consistency is su cient to enable practical FP4 training at scale, providing a simple and e ective pathway toward more e cient LLM training.

cs.LG

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.

cs.CV

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing one-step editing methods primarily rely on text conditioning for semantic transformation, lacking explicit spatial control over \textit{where} to edit. More importantly, even when spatial constraints are introduced, these methods often struggle to achieve strong and stable semantic modifications within the target regions. In this work, we revisit one-step image editing from a spatially controlled perspective and identify two key challenges: discovering editable regions and achieving effective localized semantic transformation. We reveal that existing methods perform global semantic transport, which limits high-intensity local editing under the one-step setting. To address this issue, we propose \textbf{WhereEdit}, a framework that reformulates one-step editing as localized adaptive editing. WhereEdit automatically identifies semantically relevant regions from internal model features and applies adaptive local modulation to enhance target-region editing while preserving non-target areas and structural consistency. Experiments on the PIE-Bench benchmark demonstrate that WhereEdit consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficiency of one-step generation. Additional experiments with region-level supervision further highlight the importance of explicit spatial reasoning for high-quality one-step image editing.

cs.CV

A reduction scheme for general-order Ising-like Hamiltonians in quantum heuristic solvers

The Ising model is ubiquitous in various optimization problems but notoriously difficult to solve due to combinatorial explosion. In view of this, Hamiltonian reduction is a useful preprocessing technique for reducing the effective problem size before applying heuristic solvers. However, existing reduction techniques mainly target second-order Ising models, whereas many pseudo-Boolean formulations naturally contain higher-order interactions. In this work, we generalize the concept of non-separable groups to arbitrary-order Ising-like models and develop a Hamiltonian reduction framework that iteratively detects and merges constrained spin groups into single variables. We benchmark the reduction on synthetic hypergraphs and higher-order network datasets, and evaluate its integration with downstream order-reduction and solver workflows. Our results establish a foundation for Hamiltonian reduction in higher-order Ising-like optimization problems.

quant-ph

Variational non-gaussian approach to interacting spin-boson models

We apply a hybrid variational framework to interacting spin-boson Hamiltonians, targeting regimes where simulations are limited by the unbounded bosonic Hilbert space and strong many-body correlations. The bosonic sector and spin-boson correlations are captured within a compact non-Gaussian variational manifold, while the minimized spin sector is obtained as the solution to an effective spin Hamiltonian. Minimization is carried out inside a self-consistent energy-minimization loop, where variational parameters are minimized and the effective Hamiltonian is solved via DMRG. The results are obtained without eliminating or truncating the photonic field. We benchmark the method on the Dicke and Dicke-Ising models by comparison to converged spin-boson DMRG, finding accurate ground-state solutions with reduced bond dimension.

quant-ph

UMCP: A Unified Multi-Task Collaborative Perception Network for Luggage Trolley Pose Estimation

In robotic autonomous luggage trolley collection, robots must continuously localize scattered luggage trolleys in cluttered and dynamic environments. This requires the vision system to achieve both high accuracy and real-time performance. However, existing visual perception approaches for luggage trolleys often rely on cascaded multi-model inference, leading to increased inference latency and high deployment costs. To address these limitations, this article presents a unified multi-task collaborative perception network (UMCP) that simultaneously performs luggage trolley detection, keypoint detection and orientation estimation. Based on the YOLOv12 architecture, keypoint features are fused with orientation features and then fed into an orientation feature enhancement module (OFEM), thereby improving orientation estimation accuracy. In addition, circular probability distribution modeling with a Kullback-Leibler (KL) divergence loss is adopted to enhance orientation estimation accuracy further. Experimental results demonstrate that the proposed method achieves competitive overall accuracy while substantially reducing model complexity and computational cost compared with existing methods. A website about this work is available at https://sites.google.com/view/robot-umcp.

cs.RO

The Wigner function for Integer quantum Hall effect

Wigner's quasi-probability distribution function in phase space is a specialized representation of the density matrix, possessing significant physical importance. In this article, we first review the wave function describing electronic motion in an electromagnetic field under the Landau gauge. Next, based on an introduction to the properties of the Wigner function, we calculate the Wigner function for the integer quantum Hall effect using the integral method.

cond-mat.mes-hall

Hessian sparsity-constrained self-supervised network for near-infrared single-photon single-pixel imaging

Near-infrared (NIR) imaging has emerged as an important technology for night vision, remote sensing, and biological imaging, yet conventional array-detector-based systems are often limited by insufficient sensitivity, high cost, and substantial dark noise. Single-pixel imaging (SPI) offers an attractive alternative, enabling single-photon-level NIR imaging by using a cost-effective single-element detector. Nevertheless, SPI remains restricted by photon noise, leading to degraded imaging quality and limited frame rate under extremely low photon flux conditions. Here, we present a Hessian sparsity-constrained self-supervised network (HS3N) for single-photon NIR SPI, which can suppress noise and enable high-fidelity and real-time imaging under ultra-low illumination conditions. The HS3N integrates the physical forward model of SPI with an untrained neural network regularized by both sparsity priors and Hessian-based structural constraints, enabling effective noise suppression while preserving structural fidelity and continuity. Both simulated and experimental results demonstrate that HS3N enables high-fidelity reconstructions under ultra-low NIR photon levels down to ~0.01 photons per pixel. Furthermore, we demonstrate its dynamic capability by monitoring the dynamic evolution and detachment of infrared-absorbing droplets, at a frame rate of ~20 Hz under ~0.19 photons per pixel, highlighting its potential for high-sensitivity infrared inspection. The proposed reconstruction framework paves the way for practical NIR imaging in extreme low light conditions, which can be extended to visible, mid-infrared or terahertz imaging, offering broad potential for photon-efficient sensing across a wide spectral range.

physics.optics