SearcharxivSearch

arXiv subjects

Shuai Zhao

Publications and source records attributed to Shuai Zhao.

At least 19 recordsLinked to original sources

NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face structure-blind backbones, backbone-bound generalization, and flat single-objective optimization. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework. It introduces a Neuro-Symbolic Encoder for plan-constrained primitive inference, a Neuro-Symbolic Solver that conditions a backbone-agnostic policy on the active primitive, and Hierarchical Joint Policy Optimization with reward-granularity matching. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability and expanded exploration space. Our code is publicly available.

cs.RO

Experimental demonstration of asynchronous measurement-device-independent quantum cryptographic conferencing

Quantum network enable a variety of quantum information processing tasks, where multi-user quantum communication is one of the important objectives. Quantum cryptographic conferencing (QCC) serves as an essential solution to establish secure keys to realize secure multi-user communications. However, existing QCC implementations have been fundamentally limited by the low probability of multi-user coincidence detection to measure or construct the Greenberger-Horne-Zeilinger (GHZ) entangled state. In this work, we report the experimental realization of QCC eliminating the need for coincidence detection, where the GHZ state is constructed by correlating detection events occurring within the coherence time, thereby greatly enhancing the success probability of GHZ-state measurement. Meanwhile, to establish and maintain high-visibility GHZ measurement among three independent users, we developed a three-party phase compensation scheme combined with precise temporal and polarization alignment within a time-bin-phase encoding framework. Furthermore, we designed an efficient pairing strategy to simplify subsequent data processing and enhance processing efficiency. Based on these techniques, we successfully performed QCC experiments over total loss of 66.3 dB, with the channel loss of 51.24 dB, achieving secure key rates of 5.4 bit/s, whereas previous QCC experiments have been limited to 100 km. The results establish a new regime of scalable, multi-user quantum communication and paving the way for metropolitan quantum networks.

quant-ph

MRI-Guided Reslice-Refined Cross-Slice SDF Reconstruction of the Left Ventricle from Cardiac MRI with Sparse Axial Supervision

Reconstructing a three-dimensional left-ventricular (LV) endocardial surface from cardiac magnetic resonance (CMR) data is challenging when supervision is available on only a small number of axial slices. Through-plane geometry is weakly constrained, and automatically generated two-dimensional masks can propagate segmentation errors into the recovered shape. We present MR-RS-SDFR, a per-case implicit signed distance field (SDF) framework that reconstructs a continuous LV surface from a CMR volume and sparse axial weak masks. The method first builds a cross-slice SDF initialization from axial and longitudinal geometric cues and then refines the field using two complementary signals: MRI edge-field normal alignment, which provides an image-derived boundary cue independent of the weak masks, and differentiable reslice Dice and contour consistency, which preserve agreement with the observed planes. We evaluate three weak-mask generators -- LOO TransUNet, LOO nnU-Net, and an off-the-shelf Medical SAM3 model used without MM-WHS-specific training or fine-tuning -- and five sparsity levels from 4 to 64 axial planes. In the sparse-16 setting, final MR-RS-SDFR reconstruction reaches 0.928 Dice and 3.80mm HD95 with Medical SAM3 masks. The upstream generators do not exhibit a single common ranking across 2D and dense 3D segmentation, and nnU-Net- and Medical-SAM3-driven sparse reconstruction achieve the same mean final Dice despite different upstream error profiles. Across all three sparse-16 mask sources, MR-RS-SDFR is numerically better than protocol-matched full GHD+DVS in both Dice and HD95. Final Dice improves markedly from sparse-4 to sparse-16 and then saturates at the reported precision through sparse-64. These results support MRI-guided per-case SDF refinement as a reconstruction strategy that remains effective across weak-mask generators and supervision densities.

cs.CV

HarmReduction: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

Millions of individuals' well-being are challenged by the harms of substance use. Harm reduction as a public health strategy provides non-judgemental, evidence-based information intended to improve health outcomes and reduce associated safety risks. Some large language models (LLMs) have demonstrated a high level of medical reasoning, promising to address the information needs of people who use drugs (PWUD). However, their performance in relevant tasks remains largely unexplored. We introduce HarmReduction, a benchmark designed to evaluate LLMs' accuracy and safety risks in harm reduction information provision. The benchmark dataset (HR-Basic) has 2,160 question-answer-evidence pairs. The scope covers three tasks: checking safety boundaries, providing quantitative values, and inferring polysubstance use risks. We build the Instruction and RAG schemes to evaluate model behaviours based on their inherent knowledge and the integration of domain knowledge. Our results indicate that state-of-the-art LLMs still struggle to provide accurate harm reduction information, and sometimes, present severe safety risks to PWUD. This work contributes an evaluation framework for LLMs to deliver harm reduction information to avoid introducing negative health outcomes through the use of LLMs.

cs.CL

Uncertainty-Aware Parameter Estimation for Condition Monitoring of Power Converters

Parameter estimation is widely used for condition monitoring of power converters, but most existing methods provide only point estimates and therefore cannot quantify whether an observed parameter change is statistically significant. This paper proposes an uncertainty-aware parameter estimation framework based on Bayesian maximum a posteriori optimization and a differentiable converter model. A Laplace approximation is used to obtain a local Gaussian posterior, enabling uncertainty quantification, consistency testing, estimator-resolution analysis, and precision-weighted pooling across data windows. The method is validated on synthetic and hardware Buck converter. It demonstrates accurate estimation of well-identified parameters and reveal the weak practical identifiability of parameters such as MOSFET on-resistance under the available sensing configuration.

eess.SY

Vision Guided Target Conditioned Control for Autonomous Excavation

Autonomous excavation requires an intelligent control system that can convert spatial work intent into coordinated bucket motion under contact-rich soil interaction. This paper presents a target-conditioned intelligent control framework for autonomous excavation in a physics-based deformable-soil simulation workflow. An image-aligned target mask serves as a visual spatial command for the desired digging region, while a mask-conditioned Action Chunking Transformer maps multi-view RGB observations, proprioception, and the target mask to temporally extended joystick commands. To reduce target-ignoring behavior, demonstrations are organized with paired-condition supervision, where the same or closely matched scene is demonstrated with different target masks and corresponding action chunks. The framework is evaluated through both a diagnostic manipulation task and an excavation simulation benchmark with single-scoop and sequential pile-clearing protocols. In manipulation, target success is 4\% for no-condition ACT, 63\% for non-paired mask-conditioned ACT, and 96\% for paired-condition mask-conditioned ACT. In sequential pile clearing, paired-condition mask-conditioned ACT removes 76.8\% of the pile versus 27.4\% and 15.7\% for the two baselines, with 91.0\% human-normalized efficiency. The results show that visual target conditioning, paired demonstration structure, and action-chunk control form a practical cyber-physical simulation pipeline for excavator automation.

cs.RO

Certifying optimal device-independent quantum randomness in quantum networks

Bell nonlocality provides a device-independent (DI) way to certify quantum randomness, based on which true random numbers can be extracted from the observed correlations without detail characterizations on devices for quantum state preparation and measurement. However, the efficiency of current strategies for DI randomness certification is still heavily constrained when it comes to non-maximal Bell values, especially for multiple parties. Here, we present a family of multipartite Bell inequalities that allows to certify optimal quantum randomness and self-test GHZ (Greenberger-Horne-Zeilinger) states, which are inspired from the stabilizer group of the GHZ state. Due to the simple representation of stabilizer group for GHZ states, this family of Bell inequalities is of simple structure and can be easily expanded to more parties. Compared with the Mermin-type inequalities, this family of Bell inequality is more efficient in certifying quantum randomness when non-maximal Bell values achieved. Meanwhile, the general analytical upper bound for the Holevo quantity is presented, and achieves better performance compared with the MABK (Mermin-Ardehali-Belinskii-Klyshko) inequality, Parity-CHSH (Clauser-Horne-Shimony-Holt) inequality and Holz inequality at $N=3$, which is of particular interests for experimental researches on DI quantum cryptography in quantum networks.

quant-ph

Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding

Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.

cs.CL

CP Polarimetry with Linearly Polarized Photon Fusion and Double-Tagged Protons

Double forward-proton tagging turns the forward detectors into event-by-event photon polarimeters because each measured proton recoil fixes the transverse momentum, and hence the linear-polarization axis, of the emitted photon. We show that this production-side polarimetry gives a decay-analyzer-independent measurement of the CP phase of a photon-coupled spin-zero resonance. We derive the leading-power photon-density contraction for a CP-mixed hard amplitude with scalar and pseudoscalar couplings. The CP phase appears as a translation of the second harmonic in the signed proton--proton azimuthal angle. This provides a compact production-side CP measurement for axionlike particles and more general spin-zero resonances.

hep-ph

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and semantically inconsistent text-image pairs that make cross-modal evidence unreliable. We propose Expert-Guided Mutual Distillation (EGMD), which learns what evidence to trust across the prediction pipeline. At the input level, input-level calibration encodes pair-level coherence as a shared gain before fusion. At the representation level, an expert-guided teacher aligns domain statistics and encourages domain-specific patterns to concentrate in specialized experts. At the decision level, prototype-anchored domain-specific students use mutual learning and dual-channel distillation to inherit the teacher's feature geometry and calibrated predictions while discouraging local domain priors. We further construct Weibo_Balanced, a domain-balanced benchmark that isolates the effect of imbalance on generalization. Across four datasets in two languages, EGMD achieves state-of-the-art accuracy while reducing domain bias by up to 57.3%.

cs.CL

Photon-Tagged Energy Flow in Inclusive Endpoint $B$ Decays

We introduce a $B$-decay tagged energy correlator (BTEC) to resolve the angular structure of energy flow in inclusive $B$ decays in the endpoint region, focusing on the direct-photon contribution to $\overline B\to X_sγ$. At leading power and at the natural collinear angular scale $τ\sim s\sim QΛ_{\rm QCD}$, we derive a factorization relation involving the standard hard coefficient and $B$-meson shape function together with a new measured quark jet function; no new leading-power nonperturbative function is introduced. We calculate the measured jet function at one-loop accuracy and verify that its angular integral reproduces the standard inclusive quark jet function. For the central Bosch--Lange--Neubert--Paz parameter set, an illustrative benchmark gives $3 - 16\%$ migration outside the fixed angular cuts $τ_c=2 - 4~{\rm GeV}^2$. With an independently constrained shape function, the angular cumulatives provide a closure test of leading-power endpoint factorization and are sensitive to direct-$O_7$ power corrections and resolved-photon effects. The BTEC thereby adds information on the angular structure of the inclusive recoil jet beyond the ordinary photon spectrum and may improve signal--background discrimination when their energy-flow profiles differ.

hep-ph

Heavy quark mass dependence of the $Λ_Q$ light-cone distribution amplitude in QCD

We study the heavy quark mass dependence of the leading-twist light-cone distribution amplitude (LCDA) of the $Λ_Q$ baryon in QCD. Starting from the factorization formula that relates the QCD LCDA to the boosted heavy-quark effective theory (bHQET) LCDA, we derive a first-order partial differential equation governing this mass dependence in the peak region. The equation is solved analytically, and the explicit factor connecting LCDAs at different heavy quark masses is presented. We further incorporate Borel-resummed perturbative corrections from a renormalon model into the factorization. The impact of these renormalon corrections on the mass dependence is studied, and a numerical analysis using a simple LCDA model is performed to illustrate the behavior and to assess the uncertainties arising from the corrections, thereby providing a numerical estimate of the associated power corrections to the mass dependence. Our results provide an essential tool for extrapolating lattice QCD calculations of heavy-baryon LCDAs from smaller simulated masses to the physical bottom quark mass.

hep-ph

Accessing the HQET B-Meson Shape Function from a LaMET Quasi-Shape Function

The shape function and the light-cone distribution amplitude of heavy meson jointly characterize the nonperturbative structure of the heavy meson on the light-cone, with the former being essential for theoretical predictions of inclusive decays and the latter for exclusive decays. While first-principles lattice QCD results for the heavy meson LCDA have become available in recent years, lattice results for the shape function remain absent. In this work, we establish a two-step factorization scheme -- known as the HQLaMET framework -- for computing the $B$-meson shape function on the lattice, which fully disentangles the effects of the disparate scales $P_B^z$, $m_b$, and $Λ_{\textrm{QCD}}$. For illustration, starting from a phenomenological model for the shape function in HQET, we provide a graphical presentation of the entire procedure of this framework. The results of the current work lay the foundation for nonperturbative lattice QCD determinations of the shape function in the near future.

hep-ph

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has emerged as a promising approach for alignment, acting as an RL-free alternative to Reinforcement Learning from Human Feedback (RLHF). Despite DPO's various advancements and inherent limitations, an in-depth review of these aspects is currently lacking in the literature. In this work, we present a comprehensive review of the challenges and opportunities in DPO, covering theoretical analyses, variants, relevant preference datasets, and applications. Specifically, we categorize recent studies on DPO based on key research questions to provide a thorough understanding of DPO's current landscape. Additionally, we propose several future research directions to offer insights on model alignment for the research community. An updated collection of relevant papers can be found on https://github.com/Mr-Loevan/DPO-Survey.

cs.AI

ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. However, a critical ambiguity persists: it remains unclear whether these models genuinely synthesize cross-modal information to construct physically grounded reasoning chains, or if they merely exploit strong language priors to mask single-modality reliance, thereby hallucinating advanced multimodal capabilities. Motivated by this, and to rigorously mitigate language modality bias and shortcuts, we propose a novel multimodal Chrono}logical Physical Dynamics Reasoning Benchmark ChronoPhyBench, which unifies next state prediction with Visual Question Answering (VQA) paradigms by conditioning on historical video context and textual captions to enforce models to deduce subsequent physical states through both single image selection and the inherently more complex task of multiple frame chronological sorting. Concurrently, we construct a large-scale multimodal reasoning dataset curated using the ChronoPhyBench criteria, comprising over 10,000 long-form videos paired with meticulously annotated captions, totaling 5M tokens. Our experimental evaluations reveal a stark contrast to conclusions drawn by previous benchmarks. The capacity of current open-source models to perform physically grounded multimodal reasoning remains in its infancy. Ultimately, this work seeks to systematically stress-test the reasoning capabilities of multimodal models, quantify hallucination rates, and advance the development of Physical AI, thereby providing the community with a robust and transparent evaluation framework toward Artificial General Intelligence (AGI).

cs.CV

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization

Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning ability of large language models. However, widely used RLVR algorithms, such as GRPO, often suffer from entropy collapse, leading to premature determinism and unstable optimization. Existing remedies, including entropy regularization and ratio-based clipping heuristics, either control entropy in a coarse-grained manner or rely on approximate on-policy training. In this paper, we revisit entropy collapse from a token-level entropy flow perspective. Our analysis reveals that entropy-decreasing tokens consistently outweigh entropy-increasing ones, resulting in a severely imbalanced entropy flow. This perspective provides a unified explanation of entropy collapse in existing RLVR algorithms and highlights the importance of balancing entropy dynamics. Motivated by this analysis, we propose On-Policy Entropy Flow Optimization (OPEFO), an adaptive entropy flow balancing mechanism that rescales entropy-increasing and entropy-decreasing updates according to their contributions to entropy change, while remaining strict on-policy. Experiments on six mathematical reasoning benchmarks demonstrate that OPEFO improves training stability and final performance. We will release the code and models upon publication.

cs.LG

EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation

Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedagogical narratives across multi-shot instructional videos, especially in STEM domains. To address these challenges, we propose EduStory, a unified framework for reliable instructional video generation. EduStory integrates pedagogical state modeling to track persistent knowledge states, script-guided structured control to organize multi-shot narratives, and learning-oriented evaluation metrics to assess knowledge fidelity and constraint satisfaction. To support rigorous evaluation, we further introduce EduVideoBench, a diagnostic benchmark with multi-granularity annotations, including pedagogical storyboards, shot-level semantics, and knowledge state transitions, together with baseline tasks for controllable instructional video generation. Extensive experiments demonstrate that domain-aware state modeling and structured control substantially reduce narrative breakdown and improve alignment with instructional intent. These results highlight the significance of domain-specific structural constraints and tailored benchmarks for advancing reliable, controllable, and also trustworthy long-horizon video generation.

cs.CV

Self-Debias: Self-correcting for Debiasing Large Language Models

Although Large Language Models (LLMs) demonstrate remarkable reasoning capabilities, inherent social biases often cascade throughout the Chain-of-Thought (CoT) process, leading to continuous "Bias Propagation". Existing debiasing methods primarily focus on static constraints or external interventions, failing to identify and interrupt this propagation once triggered. To address this limitation, we introduce Self-Debias, a progressive framework designed to instill intrinsic self-correction capabilities. Specifically, we reformulate the debiasing process as a strategic resource redistribution problem, treating the model's output probability mass as a limited resource to be reallocated from biased heuristics to unbiased reasoning paths. Unlike standard preference optimization which applies broad penalties, Self-Debias employs a fine-grained trajectory-level objective subject to dynamic debiasing constraints. This enables the model to selectively revise biased reasoning suffixes while preserving valid contextual prefixes. Furthermore, we integrate an online self-improvement mechanism utilizing consistency filtering to autonomously synthesize supervision signals. With merely 20k annotated samples, Self-Debias activates efficient self-correction, achieving superior debiasing performance while preserving general reasoning capabilities without continuous external oversight.

cs.CL