SearcharxivSearch

arXiv subjects

Xinhao Hu

Publications and source records attributed to Xinhao Hu.

11 recordsLinked to original sources

Learning from the Self-future: On-policy Self-distillation for dLLMs

On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitraryorder generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from "self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms RLVR and SFT baselines with superior sample efficiency, requiring only around 10% of the optimization steps by RLVR and opening a promising pathway for dLLM posttraining. The code is available at https://github.com/xingzhejun/d-OPSD.

cs.CL

Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

Zero-shot vision-and-language navigation (VLN) has gained significant attention due to its minimal data collection costs and inherent generalization. This paradigm is typically driven by the integration of pre-trained Vision-Language Models (VLMs) and Large Language Models (LLMs), where VLMs construct 3D scene graphs while LLMs handle high-level reasoning and decision-making. However, a critical bottleneck exists in this system: current 3D perception models prioritize pixel-level accuracy, directly conflicting with the strict computational limits and real-time efficiency demanded by embodied navigation. To address this gap, this paper quantifies the actual impact of 3D scene understanding capability on VLN performance. Based on typical VLM-LLM frameworks, we propose statistical success rate (SR) upper bounds for two core subsystems: 1) the slow LLM planner, which relies on topological mapping semantics, and 2) the fast reactive navigator, which utilizes spatial coordinates and bounding boxes to execute LLM decisions. Evaluations using state-of-the-art 3D scene understanding models validate our proposed bounds and reveal a perception saturation phenomenon, indicating that improvements in perception accuracy beyond a certain threshold yield diminishing returns in navigation success. Our findings suggest that 3D scene understanding for VLN should pivot away from strict pixel-level precision, prioritizing instead navigation-relevant core vocabularies and accurate bounding box proportions.

cs.RO

Enhancing Domain Generalization in 3D Human Pose Estimation through Controllable Generative Augmentation

Pedestrian motion, due to its causal nature, is strongly influenced by domain gaps arising from discrepancies between training and testing data distributions. Focusing on 3D human pose estimation, this work presents a controllable human pose generation framework that synthesizes diverse video data by systematically varying poses, backgrounds, and camera viewpoints. This generative augmentation enriches training datasets, enhances model generalization, and alleviates the limitations of existing methods in handling domain discrepancies. By leveraging both indoor/real-world and outdoor/virtual datasets, we perform cross-domain data fusion and controllable video generation to construct enriched training data, tailored to realistic deployment settings. Extensive experiments show that the augmented datasets significantly improve model performance on unseen scenarios and datasets, validating the effectiveness of the proposed approach.

cs.CV

All-optical intracellular thermal profiling using nanodiamond-based "thermal radar"

The local thermal conductivity (\k{appa}) is a pivotal biophysical parameter, governing intracellular heat flux and underlying functional processes like metabolic regulation and stress response. However, label-free mapping with sub-micron resolution in living cells remains challenge. Here, we present frequency-domain fluorescence thermometry (FD-FTM), an all-optical method based on a hybrid nanodiamond-on-gold-membrane platform, which enables quantitative mapping of \k{appa} in biological systems. Fluorescence nanodiamonds (FNDs) are deposited on substrates coated with a 50 nm gold membrane, where FNDs function as nanoscale thermometers, and the gold membrane serves as a photothermal heat source. We validate FD-FTM across reference materials and biological media, with fitting uncertainties of ~10%. By varying the modulation frequency, we tune the thermal penetration depths, enabling controlled heat propagation from the substrate to the cell nucleus. The method delivers sensitivity sufficient to resolve changes in biofluid thermal conductivity on the order of 16% relative to water. Using these capabilities, we demonstrate non-invasive thermal profiling across scales: at the cellular level, nuclear chromatin packing yields \k{appa} higher by ~10% relative to the cytoplasm; at the organelle level, we resolve \k{appa} variations associated with protein aggregates formed during liquid-liquid phase separation in an amyotrophic lateral sclerosis disease model. Temporal measurements in living cells over 30 minutes further reveal spatially resolved intracellular responses to osmotic stress, linking nanoscale thermal dynamics to biomolecular condensates. These results establish FD-FTM as a label-free, robust, and quantitative platform for thermally decoding intracellular processes, opening avenues for studying metabolic heterogeneity, disease mechanisms, and therapeutic responses.

physics.bio-ph

SFCo-Nav: Efficient Zero-Shot Visual Language Navigation via Collaboration of Slow LLM and Fast Attributed Graph Alignment

Recent advances in large vision-language models (VLMs) and large language models (LLMs) have enabled zero-shot approaches to visual language navigation (VLN), where an agent follows natural language instructions using only ego perception and reasoning. However, existing zero-shot methods typically construct a naive observation graph and perform per-step VLM-LLM inference on it, resulting in high latency and computation costs that limit real-time deployment. To address this, we present SFCo-Nav, an efficient zero-shot VLN framework inspired by the principle of slow-fast cognitive collaboration. SFCo-Nav integrates three key modules: 1) a slow LLM-based planner that produces a strategic chain of subgoals, each linked to an imagined object graph; 2) a fast reactive navigator for real-time object graph construction and subgoal execution; and 3) a lightweight asynchronous slow-fast bridge aligns advanced structured, attributed imagined and perceived graphs to estimate navigation confidence, triggering the slow LLM planner only when necessary. To the best of our knowledge, SFCo-Nav is the first slow-fast collaboration zero-shot VLN system supporting asynchronous LLM triggering according to the internal confidence. Evaluated on the public R2R and REVERIE benchmarks, SFCo-Nav matches or exceeds prior state-of-the-art zero-shot VLN success rates while cutting total token consumption per trajectory by over 50% and running more than 3.5 times faster. Finally, we demonstrate SFCo-Nav on a legged robot in a hotel suite, showcasing its efficiency and practicality in indoor environments.

cs.RO

Swelling-Induced Stress-Assisted Transfer of Nanodiamond Arrays with a PVA Carrier Tape for Conformal Bio-Integrated Sensing and Labelling

The conformal integration of nitrogen-vacancy (NV) center nanodiamond arrays onto soft, hydrated, curvilinear biological interfaces remain a fundamental challenge for in vivo quantum sensing and imaging. Conventional transfer techniques often fail due to reliance on high temperature, corrosive chemicals, or mechanical peeling, leading to pattern damage, low fidelity, or poor biocompatibility. Here, we report a transfer strategy utilizing polyvinyl alcohol (PVA) carrier soluble tape, enabling rapid, residue-free, high-fidelity transfer of nanodiamond patterns onto diverse biointerfaces. The success of this method is rooted in a unique "hydrate-soften-expand-self-peel" mechanism of the soluble tape with PVA backing. In situ mechanical tracking reveals non-uniform PVA swelling upon hydration generates transient local normal and shear stresses at the interface. These stresses delaminate the tape within 3 minutes at room temperature while promoting adhesion of the nanodiamond array to the substrate. In contrast, conventional water-soluble tapes with composite structures undergo passive dissolution and collapse, causing residue contamination and reduced efficiency. Leveraging this mechanism, we achieve conformal patterning on ultra-soft hydrogels (~0.6 kPa) and highly curved bio-surfaces (hair, 100 {\mu}m^-1). Additionally, we demonstrate a dual-identity verification system integrating data storage and physical unclonable functions on a hydrogel contact lens. This work provides a versatile tool for bio-interface engineering and a general framework for gentle, efficient transfer of functional nanomaterials.

physics.bio-ph

Nanodiamond-Enabled Torsion Microscopy Uncovers Multidimensional Cell-Matrix Mechanical Interactions

Traditional cellular force-sensing techniques, such as traction force microscopy (TFM), are predominantly limited to measuring linear tractions, overlooking and technically unable to capture the nanoscale torsional forces that are critical in cell-matrix interactions. Here, we introduce a nanodiamond-enabled torsion microscopy (DTM) that integrates nitrogen-vacancy (NV) centers as orientation markers with micropillar arrays to decouple and quantify nanoscale rotational and translational motions induced by cells. This approach achieves high precision (~1.47 degree rotational accuracy and ~3.13*10-15 Nm torque sensitivity), enabling reconstruction of cellular torsional force fields and twisting energy distributions previously underestimated. Our findings reveal the widespread presence of torsional forces in cell-matrix interactions, introducing "cellular mechanical modes" where different adhesion patterns dictate the balance between traction- and torque- mediated mechanical energy transferred to the substrate. Notably, in immune cells like macrophages that generally exert low linear tractions, torque overwhelmingly dominates traction, highlighting a unique mechanical output for specific cellular functions. By uncovering these differential modes, DTM provides a versatile tool to advance biomechanical investigations, with potential applications in disease diagnostics and therapeutics.

physics.bio-ph

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive steps into a coherent 'chunk' and shifting the policy optimization paradigm from GRPO's step level to the chunk level can effectively mitigate the negative impact of this issue. Building on this insight, we propose Group Chunking Policy Optimization (GCPO), the first chunk-level reinforcement learning approach for post-training flow matching. Extensive experiments demonstrate that GCPO achieves superior performance on both standard T2I benchmarks and preference alignment, with up to 43% relative gains over GRPO, highlighting the promise of chunk-level policy optimization. The code is available on https://github.com/xingzhejun/GCPO.

cs.CV

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO, the first method to incorporate Group Relative Policy Optimization (GRPO)-based RL into this overlooked paradigm. Our core insight is to redefine the transition probability, which is different from current approaches, and formulate the unmasking process as a multi-step decision-making problem. To further enhance our method, we explore several useful strategies, including removing the KL constraint, applying the reduction strategy, and filtering out low-quality samples. Using Mask-GRPO, we improve a base model, Show-o, with substantial improvements on standard T2I benchmarks and preference alignment, outperforming existing state-of-the-art approaches. The code is available on https://github.com/xingzhejun/Mask-GRPO

cs.CV

Examining Different Placement Strategies for Indoor Environmental Quality Sensors in Office Environments

Collecting Indoor Environmental Quality (IEQ) data from an occupant's immediate surroundings can provide personalized insights for healthy environmental conditions aligned with occupant preferences, but effective sensor placement for data accuracy and reliability has not been thoroughly explored. This paper explores various positioning of IEQ multi-sensing devices at individual workstations in typical office settings, aiming to identify sensor placements that most accurately reflect the environmental conditions experienced by occupants. We examined five unique positions close to an occupant (above and below the monitor, right side of the desk, ceiling, and chair backrest), two orientations, and three desk locations characterized by different lighting levels, thermal and airflow conditions. Data on temperature, humidity, carbon dioxide (CO2), particulate matters (PM1, PM2.5, PM10), illuminance, and sound were collected over a 2-week longitudinal experiment, followed by short-term experiments simulating common pollution events such as coughing and sneezing. Principal Component Analysis, Spearman's rank correlation, R2, and Mean Absolute Error were applied to identify the position and orientation that best captures the most information and matches breathing zone measurements. It was found that above the monitor position, facing the occupant, best captures the IEQ conditions experienced by the occupant.

eess.SP

Discriminative Addressing of Versatile Nanodiamonds via Physically-Enabled Classifier in Complex Bio-Systems

Nitrogen-vacancy (NV) centers show great potentials for nanoscale bio-sensing and bio-imaging. Nevertheless, their envisioned bio-applications suffer from intrinsic background noise due to unavoidable light scattering and autofluorescence in cells and tissues. Herein, we develop a novel all-optical modulated imaging method via physically-enabled classifier, for on-demand and direct access to NV fluorescence at pixel resolution while effectively filtering out background noise. Specifically, NV fluorescence can be modulated optically to exhibit sinusoid-like variations, providing basis for classification. We validate our method in various complex biological scenarios with fluorescence interference, ranging from cells to organisms. Notably, our classification-based approach achieves almost 10^6 times enhancement of signal-to-background ratio (SBR) for fluorescent nanodiamonds (FNDs) in neural protein imaging. We also demonstrate 4-fold contrast improvement in optically-detected magnetic resonance measurements (ODMR) of FNDs inside stained cells. Our technique offers a generic, explainable and robust solution, applicable for realistic high-fidelity imaging and sensing in challenging noise-laden scenarios.

physics.optics