SearcharxivSearch

arXiv subjects

Haoyu Xie

Publications and source records attributed to Haoyu Xie.

At least 19 recordsLinked to original sources

Dynamic Water-Wave Tweezers

Following a recent demonstration of stable trapping of floating particles by stationary (monochromatic) structured water waves [Nature 638, 394 (2025)], we report dynamic water-wave tweezers that enable controllable transport of the trapped particle along an arbitrary trajectory on the water surface. Furthermore, we demonstrate simultaneous transport of two trapped particles along different trajectories. We employ a triangular lattice formed by the interference of three plane waves, which can trap particles, depending on the wave frequency and particle parameters, either at intensity maxima or at intensity zeros (vortices). By introducing small frequency detunings of the interfering waves, we control 2D motion of the lattice and trapped particles. This approach is robust and effective over a relatively broad range of particle sizes and wave frequencies, offering remarkable new possibilities for noncontact manipulation of floating (e.g., biological and soft-matter) objects in fluidic environments.

physics.flu-dyn

mmSimPrior: Learning Simulation Priors for Data-Efficient and Generalizable Real-World Radar-based Human Motion Reconstruction

Millimeter-wave (mmWave) radar enables privacy-preserving and illumination-robust human motion reconstruction, but training generalizable models typically requires costly paired radar-motion recordings. Simulation can scale such supervision, yet even physics-based simulators cannot fully reproduce real-world multipath, clutter, hardware-specific response statistics, or distance-dependent resolution degradation, leaving a sim-to-real gap. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. To learn transferable signal and motion priors, we pretrain a multimodal radar encoder with a physics-informed domain-randomization curriculum designed to mitigate the sim-to-real gap by approximating real-world propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A dual-mode mapping module predicts either motion-code distributions for structurally constrained zero-shot reconstruction or continuous motion parameters for flexible adaptation from limited real data. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that prevents any exact subject-environment-location-motion tuple from appearing in both the adaptation and test sets. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7-39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without fine-tuning.

cs.CV

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current guardrail models lack agentic risk awareness and transparency in risk diagnosis. To introduce an agentic guardrail that covers complex and numerous risky behaviors, we first propose a unified three-dimensional taxonomy that orthogonally categorizes agentic risks by their source (where), failure mode (how), and consequence (what). Guided by this structured and hierarchical taxonomy, we introduce a new fine-grained agentic safety benchmark (ATBench) and a Diagnostic Guardrail framework for agent safety and security (AgentDoG). AgentDoG provides fine-grained and contextual monitoring across agent trajectories. More Crucially, AgentDoG can diagnose the root causes of unsafe actions and seemingly safe but unreasonable actions, offering provenance and transparency beyond binary labels to facilitate effective agent alignment. AgentDoG variants are available in three sizes (4B, 7B, and 8B parameters) across Qwen and Llama model families. Extensive experimental results demonstrate that AgentDoG achieves state-of-the-art performance in agentic safety moderation in diverse and complex interactive scenarios. All models and datasets are openly released.

cs.AI

Monocular Models are Strong Learners for Multi-View Human Mesh Recovery

Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision.

cs.CV

A wafer-scale ultrasensitive programmable chiroptical sensor

Chiroptical enantioselective sensing is gaining traction across various applications. However, intrinsic molecular chiroptical responses are weak, and existing amplification approaches add synthesis, manufacturing, or operational complexity that limits sensitivity, scalability, and dynamic control. Here, we present a fundamentally new sensing paradigm merging adsorption-driven chirality induction with wafer-scale optical transduction in a programmable heterostructure containing twisted aligned carbon nanotubes (CNTs) and phase change materials (PCMs). Chiral molecules adsorb onto CNTs to form chiroptically active composites that are macroscopically assembled by alignment and rotational stacking, yielding large ultraviolet circular dichroism (CD). We resolve molecule concentration and handedness in a single device without lithography, hotspot delivery, or differential protocols, achieving sub-$μ$M sensitivity for CD-silent glucose and chiral amino acids enabled by $>10^5\,\mathrm{M^{-1}}$ adsorption constants. We validate adsorption using molecular dynamics simulations, reproduce experimental results using chiral transfer matrix simulations, and realize sensor programmability by tuning the PCM layer. This platform enables cost-effective in-situ enantiomer monitoring in aqueous environments.

physics.optics

TopoFR: A Closer Look at Topology Alignment on Face Recognition

The field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks has demonstrated the effectiveness of data structure information. Considering that the FR task can leverage large-scale training data, which intrinsically contains significant structure information, we aim to investigate how to encode such critical structure information into the latent space. As revealed from our observations, directly aligning the structure information between the input and latent spaces inevitably suffers from an overfitting problem, leading to a structure collapse phenomenon in the latent space. To address this problem, we propose TopoFR, a novel FR model that leverages a topological structure alignment strategy called PTSA and a hard sample mining strategy named SDE. Concretely, PTSA uses persistent homology to align the topological structures of the input and latent spaces, effectively preserving the structure information and improving the generalization performance of FR model. To mitigate the impact of hard samples on the latent space structure, SDE accurately identifies hard samples by automatically computing structure damage score (SDS) for each sample, and directs the model to prioritize optimizing these samples. Experimental results on popular face benchmarks demonstrate the superiority of our TopoFR over the state-of-the-art methods. Code and models are available at: https://github.com/modelscope/facechain/tree/main/face_module/TopoFR.

cs.CV

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.

cs.SD

Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition

Wearable Human Activity Recognition (WHAR) is a prominent research area within ubiquitous computing. Multi-sensor synchronous measurement has proven to be more effective for WHAR than using a single sensor. However, existing WHAR methods use shared convolutional kernels for indiscriminate temporal feature extraction across each sensor variable, which fails to effectively capture spatio-temporal relationships of intra-sensor and inter-sensor variables. We propose the DecomposeWHAR model consisting of a decomposition phase and a fusion phase to better model the relationships between modality variables. The decomposition creates high-dimensional representations of each intra-sensor variable through the improved Depth Separable Convolution to capture local temporal features while preserving their unique characteristics. The fusion phase begins by capturing relationships between intra-sensor variables and fusing their features at both the channel and variable levels. Long-range temporal dependencies are modeled using the State Space Model (SSM), and later cross-sensor interactions are dynamically captured through a self-attention mechanism, highlighting inter-sensor spatial correlations. Our model demonstrates superior performance on three widely used WHAR datasets, significantly outperforming state-of-the-art models while maintaining acceptable computational efficiency.

cs.CV

Carbon-Nanotube/$β$-Ga$_2$O$_3$ Heterojunction PIN Diodes

$β$-Ga$_2$O$_3$ is gaining attention as a promising semiconductor for next-generation high-power, high-efficiency, and high-temperature electronic devices, thanks to its exceptional material properties. However, challenges such as the lack of viable p-type doping have hindered its full potential, particularly in the development of ambipolar devices. This work introduces a novel heterojunction diode (HD) that combines p-type carbon nanotubes (CNTs) with i/n-type $β$-Ga$_2$O$_3$ to overcome these limitations. For the first time, a CNT/$β$-Ga$_2$O$_3$ hetero-p-n-junction diode is fabricated. Compared to a traditional Schottky barrier diode (SBD) with the same $β$-Ga$_2$O$_3$ epilayer, the CNT/$β$-Ga$_2$O$_3$ HD demonstrates significant improvements, including a higher rectifying ratio ($1.2 \times 10^{11}$), a larger turn-on voltage (1.96 V), a drastically reduced leakage current at temperatures up to 300 °C, and a 26.7% increase in breakdown voltage. Notably, the CNT/$β$-Ga$_2$O$_3$ HD exhibits a low ideality factor of 1.02, signifying an ideal interface between the materials. These results underline the potential of CNT/$β$-Ga$_2$O$_3$ heterojunctions for electronic applications, offering a promising solution to current limitations in $β$-Ga$_2$O$_3$-based devices.

physics.app-ph

FaceChain-FACT: Face Adapter with Decoupled Training for Identity-preserved Personalization

In the field of human-centric personalized image generation, the adapter-based method obtains the ability to customize and generate portraits by text-to-image training on facial data. This allows for identity-preserved personalization without additional fine-tuning in inference. Although there are improvements in efficiency and fidelity, there is often a significant performance decrease in test following ability, controllability, and diversity of generated faces compared to the base model. In this paper, we analyze that the performance degradation is attributed to the failure to decouple identity features from other attributes during extraction, as well as the failure to decouple the portrait generation training from the overall generation task. To address these issues, we propose the Face Adapter with deCoupled Training (FACT) framework, focusing on both model architecture and training strategy. To decouple identity features from others, we leverage a transformer-based face-export encoder and harness fine-grained identity features. To decouple the portrait generation training, we propose Face Adapting Increment Regularization~(FAIR), which effectively constrains the effect of face adapters on the facial region, preserving the generative ability of the base model. Additionally, we incorporate a face condition drop and shuffle mechanism, combined with curriculum learning, to enhance facial controllability and diversity. As a result, FACT solely learns identity preservation from training data, thereby minimizing the impact on the original text-to-image capabilities of the base model. Extensive experiments show that FACT has both controllability and fidelity in both text-to-image generation and inpainting solutions for portrait generation.

cs.CV

Optical modeling, solver, and design of wafer-scale single-enantiomer carbon nanotube film and reconfigurable chiral photonic device

The interaction of circularly polarized light with chiral matter and functional devices enables novel phenomena and applications. Recently, wafer-scale solid-state single-enantiomer carbon nanotube (CNT) films have become feasible and are emerging as a chiral photonic material platform thanks to their quantum-confinement-induced optical properties and facile scalable assembly. However, optical modeling, solver, and device design tools for such materials are non-existent. Here, we prepare wafer-scale single-enantiomer (6,5) and (11,-5) randomly oriented CNT films and create an optical material model based on measured experimental optical spectra. We also implement a highly-parallel graphic-processing-unit accelerated transfer matrix solver for general bi-anisotropic materials and layered structures. Further, we demonstrate reconfigurable chiral photonic devices in a heterostructure with phase change materials through machine learning-enabled efficient gradient-based inverse design and optimization. Our developed full stack of a chiral photonic material and device hardware platform and a corresponding high-performance differential-programming-enabled solver opens the door for future chiral photonic devices and applications based on single-enantiomer CNT films.

physics.optics

LiDAR-based 4D Occupancy Completion and Forecasting

Scene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In this paper, we introduce a novel LiDAR perception task of Occupancy Completion and Forecasting (OCF) in the context of autonomous driving to unify these aspects into a cohesive framework. This task requires new algorithms to address three challenges altogether: (1) sparse-to-dense reconstruction, (2) partial-to-complete hallucination, and (3) 3D-to-4D prediction. To enable supervision and evaluation, we curate a large-scale dataset termed OCFBench from public autonomous driving datasets. We analyze the performance of closely related existing baseline models and our own ones on our dataset. We envision that this research will inspire and call for further investigation in this evolving and crucial area of 4D perception. Our code for data curation and baseline implementation is available at https://github.com/ai4ce/Occ4cast.

cs.CV

A programmable wafer-scale chiroptical heterostructure of twisted aligned carbon nanotubes and phase change materials

The ability to design and dynamically control chiroptical responses in solid-state matter at wafer scale enables new opportunities in various areas. Here we present a full stack of computer-aided designs and experimental implementations of a dynamically programmable, unified, scalable chiroptical heterostructure containing twisted aligned one-dimensional (1D) carbon nanotubes (CNTs) and non-volatile phase change materials (PCMs). We develop a software infrastructure based on high-performance machine learning frameworks, including differentiable programming and derivative-free optimization, to efficiently optimize the tunability of both excitonic reciprocal and linear-anisotropy-induced nonreciprocal circular dichroism (CD) responses. We experimentally implement designed heterostructures with wafer-scale self-assembled aligned CNTs and deposited PCMs. We dynamically program reciprocal and nonreciprocal CD responses by inducing phase transitions of PCMs, and nonreciprocal responses display polarity reversal of CD upon sample flipping in broadband spectral ranges. All experimental results agree with simulations. Further, we demonstrate that the vertical dimension of heterostructure is scalable with the number of stacking layers and aligned CNTs play dual roles - the layer to produce CD responses and the Joule heating electrode to electrically program PCMs. This heterostructure platform is versatile and expandable to a library of 1D nanomaterials and electro-optic materials for exploring novel chiral phenomena and photonic and optoelectronic devices.

physics.optics

Large-scale Outdoor Cell-free mMIMO Channel Measurement in an Urban Scenario at 3.5 GHz

The design of cell-free massive MIMO (CF-mMIMO) systems requires accurate, measurement-based channel models. This paper provides the first results from the by far most extensive outdoor measurement campaign for CF-mMIMO channels in an urban environment. We measured impulse responses between over 20,000 potential access point (AP) locations and 80 user equipments (UEs) at 3.5 GHz with 350 MHz bandwidth (BW). Measurements use a "virtual array" approach at the AP and a hybrid switched/virtual approach at the UE. This paper describes the sounder design, measurement environment, data processing, and sample results, particularly the evolution of the power-delay profiles (PDPs) as a function of the AP locations, and its relation to the propagation environment.

eess.SP

PRCL: Probabilistic Representation Contrastive Learning for Semi-Supervised Semantic Segmentation

Tremendous breakthroughs have been developed in Semi-Supervised Semantic Segmentation (S4) through contrastive learning. However, due to limited annotations, the guidance on unlabeled images is generated by the model itself, which inevitably exists noise and disturbs the unsupervised training process. To address this issue, we propose a robust contrastive-based S4 framework, termed the Probabilistic Representation Contrastive Learning (PRCL) framework to enhance the robustness of the unsupervised training process. We model the pixel-wise representation as Probabilistic Representations (PR) via multivariate Gaussian distribution and tune the contribution of the ambiguous representations to tolerate the risk of inaccurate guidance in contrastive learning. Furthermore, we introduce Global Distribution Prototypes (GDP) by gathering all PRs throughout the whole training process. Since the GDP contains the information of all representations with the same class, it is robust from the instant noise in representations and bears the intra-class variance of representations. In addition, we generate Virtual Negatives (VNs) based on GDP to involve the contrastive learning process. Extensive experiments on two public benchmarks demonstrate the superiority of our PRCL framework.

cs.CV

FaceChain: A Playground for Human-centric Artificial Intelligence Generated Content

Recent advancement in personalized image generation have unveiled the intriguing capability of pre-trained text-to-image models on learning identity information from a collection of portrait images. However, existing solutions are vulnerable in producing truthful details, and usually suffer from several defects such as (i) The generated face exhibit its own unique characteristics, \ie facial shape and facial feature positioning may not resemble key characteristics of the input, and (ii) The synthesized face may contain warped, blurred or corrupted regions. In this paper, we present FaceChain, a personalized portrait generation framework that combines a series of customized image-generation model and a rich set of face-related perceptual understanding models (\eg, face detection, deep face embedding extraction, and facial attribute recognition), to tackle aforementioned challenges and to generate truthful personalized portraits, with only a handful of portrait images as input. Concretely, we inject several SOTA face models into the generation procedure, achieving a more efficient label-tagging, data-processing, and model post-processing compared to previous solutions, such as DreamBooth ~\cite{ruiz2023dreambooth} , InstantBooth ~\cite{shi2023instantbooth} , or other LoRA-only approaches ~\cite{hu2021lora} . Besides, based on FaceChain, we further develop several applications to build a broader playground for better showing its value, including virtual try on and 2D talking head. We hope it can grow to serve the burgeoning needs from the communities. Note that this is an ongoing work that will be consistently refined and improved upon. FaceChain is open-sourced under Apache-2.0 license at \url{https://github.com/modelscope/facechain}.

cs.CV

Space Engage: Collaborative Space Supervision for Contrastive-based Semi-Supervised Semantic Segmentation

Semi-Supervised Semantic Segmentation (S4) aims to train a segmentation model with limited labeled images and a substantial volume of unlabeled images. To improve the robustness of representations, powerful methods introduce a pixel-wise contrastive learning approach in latent space (i.e., representation space) that aggregates the representations to their prototypes in a fully supervised manner. However, previous contrastive-based S4 methods merely rely on the supervision from the model's output (logits) in logit space during unlabeled training. In contrast, we utilize the outputs in both logit space and representation space to obtain supervision in a collaborative way. The supervision from two spaces plays two roles: 1) reduces the risk of over-fitting to incorrect semantic information in logits with the help of representations; 2) enhances the knowledge exchange between the two spaces. Furthermore, unlike previous approaches, we use the similarity between representations and prototypes as a new indicator to tilt training those under-performing representations and achieve a more efficient contrastive learning process. Results on two public benchmarks demonstrate the competitive performance of our method compared with state-of-the-art methods.

cs.CV

Boosting Semi-Supervised Semantic Segmentation with Probabilistic Representations

Recent breakthroughs in semi-supervised semantic segmentation have been developed through contrastive learning. In prevalent pixel-wise contrastive learning solutions, the model maps pixels to deterministic representations and regularizes them in the latent space. However, there exist inaccurate pseudo-labels which map the ambiguous representations of pixels to the wrong classes due to the limited cognitive ability of the model. In this paper, we define pixel-wise representations from a new perspective of probability theory and propose a Probabilistic Representation Contrastive Learning (PRCL) framework that improves representation quality by taking its probability into consideration. Through modelling the mapping from pixels to representations as the probability via multivariate Gaussian distributions, we can tune the contribution of the ambiguous representations to tolerate the risk of inaccurate pseudo-labels. Furthermore, we define prototypes in the form of distributions, which indicates the confidence of a class, while the point prototype cannot. Moreover, we propose to regularize the distribution variance to enhance the reliability of representations. Taking advantage of these benefits, high-quality feature representations can be derived in the latent space, thereby the performance of semantic segmentation can be further improved. We conduct sufficient experiment to evaluate PRCL on Pascal VOC and CityScapes to demonstrate its superiority. The code is available at https://github.com/Haoyu-Xie/PRCL.

cs.CV