SearcharxivSearch

arXiv subjects

Xinmiao Wang

Publications and source records attributed to Xinmiao Wang.

9 recordsLinked to original sources

X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation

While recent advances have demonstrated strong performance in individual humanoid skills such as upright locomotion, fall recovery and whole-body coordination, learning a single policy that masters all these skills remains challenging due to the diverse dynamics and conflicting control objectives involved. To address this, we introduce X-Loco, a framework for training a vision-based generalist humanoid locomotion policy. X-Loco trains multiple oracle specialist policies and adopts a synergetic policy distillation with a case-adaptive specialist selection mechanism, which dynamically leverages multiple specialist policies to guide a vision-based student policy. This design enables the student to acquire a broad spectrum of locomotion skills, ranging from fall recovery to terrain traversal and whole-body coordination skills. To the best of our knowledge, X-Loco is the first framework to demonstrate vision-based humanoid locomotion that jointly integrates upright locomotion, whole-body coordination and fall recovery, while operating solely under velocity commands without relying on reference motions. Experimental results show that X-Loco achieves superior performance, demonstrated by tasks such as fall recovery and terrain traversal. Ablation studies further highlight that our framework effectively leverages specialist expertise and enhances learning efficiency.

cs.RO

VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training

Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging. We identify two critical mismatches: wrist-mounted fisheye views, with severe radial distortion and local gripper-centric perspectives, are out-of-distribution for pretrained VLMs; and human-collected trajectories frequently violate kinematic limits, incur collisions, or exceed controller bandwidth, teaching VLA policies physically infeasible actions. To address the challenges, we present VISTA, a framework that bridges this dual gap through three synergistic components. (i)~UMI-VQA, the first large-scale VQA dataset tailored to wrist-mounted fisheye observations, aligns VLM representations to the distorted visual regime via auxiliary vision-language supervision. (ii)~A systematic physical-validation pipeline performs a data-completeness pre-check and scores each valid trajectory for trajectory continuity, self-collision risk, and execution fidelity before it enters training. (iii)~A two-stage co-training recipe jointly learns vision-language grounding on UMI-VQA and action prediction on validated trajectories. Our experiments empirically show that incorporating UMI-VQA consistently improves downstream policy performance, and that physical-validation scores are strongly predictive of deployment success. On diverse simulation and real-world manipulation tasks, VISTA significantly outperforms strong baselines including $π_{0.5}$, LingBot-VLA, and Wall-X. We release the physical-validation pipeline, UMI-VQA, validated trajectory data, and the pre-trained model for the community.

cs.RO

Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning

Humanoid robots are promising to learn a diverse set of human-like locomotion behaviors, including standing up, walking, running, and jumping. However, existing methods predominantly require training independent policies for each skill, yielding behavior-specific controllers that exhibit limited generalization and brittle performance when deployed on irregular terrains and in diverse situations. To address this challenge, we propose Adaptive Humanoid Control (AHC) that adopts a two-stage framework to learn an adaptive humanoid locomotion controller across different skills and terrains. Specifically, we first train several primary locomotion policies and perform a multi-behavior distillation process to obtain a basic multi-behavior controller, facilitating adaptive behavior switching based on the environment. Then, we perform reinforced fine-tuning by collecting online feedback in performing adaptive behaviors on more diverse terrains, enhancing terrain adaptability for the controller. We conduct experiments in both simulation and real-world experiments in Unitree G1 robots. The results show that our method exhibits strong adaptability across various situations and terrains. Project website: https://ahc-humanoid.github.io.

cs.RO

PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning

Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and perceptive locomotion, they are constrained by fixed weighting strategies that produce a single suboptimal policy, rather than providing a diverse set of solutions for sophisticated multi-objective control. In this paper, we propose a novel framework leveraging Multi-Objective Reinforcement Learning (MORL) to achieve Preference-Conditioned Humanoid Control (PCHC). Unlike conventional methods that require training a series of policies to approximate the Pareto front, our framework enables a single, preference-conditioned policy to exhibit a wide spectrum of diverse behaviors. To effectively integrate these requirements, we introduce a Beta distribution-based alignment mechanism based on preference vectors modulating a Mixture-of-Experts (MoE) module. We validated our approach on two representative humanoid tasks. Extensive simulations and real-world experiments demonstrate that the proposed framework allows the robot to adaptively shift its objective priorities in real-time based on the input preference condition.

cs.RO

Hairless Black Hole by Superradiance

We investigate the interplaying effects of black hole scalarization and superradiance in the context of the Einstein-Maxwell-scalar model, with the scalar field possessing electric charge. Restricted to spherical symmetry, our linear analysis about a Reissner-Nordström background confirms the persistence of tachyonic scalar modes upon inclusion of electric charge. However, fully nonlinear numerical simulations reveal that the system no longer evolves into a scalarized, hairy black hole state. Instead, we find that the superradiance phenomenon (specifically the electromagnetic version of the effect) causes the scalar condensate to become fully depleted through accretion into the black hole and radiation to spatial infinity. The combination of the two effects, which we refer to as "tachyonic superradiance", may thus be seen as a particularly efficient mechanism for the extraction of energy from a black hole, exploiting both the tachyonic growth and superradiant emission. We accurately compute the amounts of energy transfer in different channels by deriving formulae for the energy fluxes of matter fields in the presence of a dynamical horizon, which are amenable for evaluation in numerical relativity calculations.

gr-qc

MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains

Humanoid robots have demonstrated robust locomotion capabilities using Reinforcement Learning (RL)-based approaches. Further, to obtain human-like behaviors, existing methods integrate human motion-tracking or motion prior in the RL framework. However, these methods are limited in flat terrains with proprioception only, restricting their abilities to traverse challenging terrains with human-like gaits. In this work, we propose a novel framework using a mixture of latent residual experts with multi-discriminators to train an RL policy, which is capable of traversing complex terrains in controllable lifelike gaits with exteroception. Our two-stage training pipeline first teaches the policy to traverse complex terrains using a depth camera, and then enables gait-commanded switching between human-like gait patterns. We also design gait rewards to adjust human-like behaviors like robot base height. Simulation and real-world experiments demonstrate that our framework exhibits exceptional performance in traversing complex terrains, and achieves seamless transitions between multiple human-like gait patterns.

cs.RO

Stable long-term evolution in numerical relativity

We report on the potential occurrence of a numerical instability in the long-time simulation of black holes using the Baumgarte-Shapiro-Shibata-Nakamura formulation of numerical relativity, even in the simple set-up of a Schwarzschild black hole. Through extensive numerical experiments, we identify that this "late-time instability" arises from accumulated violations of the momentum constraint. To address this issue, we propose two modified versions of the so-called conformal covariant Z4 scheme, designed to propagate momentum constraint violations without damping. Our results demonstrate that these alternative formulations, which we refer to as CCZ4' and CCZ3, effectively resolve the late-time numerical instability not only in Schwarzschild spacetimes but also in black hole spacetimes with matter fields. Notably, by preventing damping of the momentum constraint violation, the Hamiltonian constraint damping can be significantly increased, which plays a crucial role in stabilizing long-term evolution in our proposed schemes.

gr-qc

Black hole accretion of scalar clouds with spontaneous symmetry breaking

Spontaneous scalarization of black holes typically occurs through the condensation of a scalar field, with the field evolving from a $U(1)$-symmetric phase into a symmetry-breaking one with lower energy. We show that there exist symmetry-breaking phases which are themselves unstable to the formation of an additional scalar condensate, or `cloud', which is partly accreted into the black hole. By studying the fully nonlinear dynamical evolution of the process, we find that symmetry breaking causes the accretion channels of scalar clouds to be non-degenerate, favoring a dominant channel for evolution. Additionally, the final states form a characteristic energy band due to varying amounts of radiation emitted by clouds in different channels.

gr-qc

To Half--Be or Not To Be?

It has recently been argued that half degrees of freedom could emerge in Lorentz and parity invariant field theories, using a non-linear Proca field theory dubbed Proca-Nuevo as a specific example. We provide two proofs, using the Lagrangian and Hamiltonian pictures, that the theory possesses a pair of second class constraints, leaving $D-1$ degrees of freedom in $D$ spacetime dimensions, as befits a consistent Proca model. Our proofs are explicit and straightforward in two dimensions and we discuss how they generalize to an arbitrary number of dimensions. We also clarify why local Lorentz and parity invariant field theories cannot hold half degrees of freedom.

hep-th