SearcharxivSearch

arXiv subjects

Xinhua Wang

Publications and source records attributed to Xinhua Wang.

At least 19 recordsLinked to original sources

WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors

Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.

cs.RO

HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL

Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) foundation models are not readily applicable to humanoid whole-body loco-manipulation. The high dimensionality and interdependence of humanoid motions make it challenging for conventional single-stage VLA architectures to coordinate locomotion, waist posture, and dual-arm manipulation effectively. Moreover, policies trained through offline behavior cloning can remain suboptimal during real-world deployment. Although online reinforcement learning can refine policies through real-world interaction, directly tuning large VLA backbones demands excessive computation and may introduce safety risks during real-robot exploration. To address these bottlenecks, we introduce HAF (Humanoid Adaptation Framework), a two-part framework consisting of HAF-VLA and HAF-Steer that transfers off-the-shelf generalist VLA foundation models to humanoid whole-body loco-manipulation. HAF-VLA is a hierarchical action-flow generator built on a pretrained flow-matching VLA. It splits full-body action denoising into three sequential stages with stage embeddings and cross-stage KV caches that retain kinematic dependencies, avoiding incoherent whole-body actions from one-shot generation. On top of the frozen HAF-VLA, HAF-Steer is a latent offline-to-online RL pipeline that leverages flow-matching invertibility and DCT-based dimensionality reduction to restrict RL optimization to a compact noise subspace and train a regularized SAC policy. This avoids updating the large VLA backbone and enables efficient real-world policy refinement. Evaluated on seven real-world humanoid loco-manipulation tasks, HAF surpasses vanilla single-stage VLA baselines and improves whole-body coordination and task performance. Project website: https://grange007.github.io/HAF .

cs.RO

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still face two fundamental challenges: (i) producing precise low-level actions from high-dimensional observations, (ii) bridging domain gaps across heterogeneous data sources, including diverse robot embodiments and human demonstrations. Existing methods often encode latent variables from either visual dynamics or robotic actions to guide policy learning, but they fail to fully exploit the complementary multi-modal knowledge present in large-scale, heterogeneous datasets. In this work, we present X Robotic Model 1 (XR-1), a novel framework for versatile and scalable VLA learning across diverse robots, tasks, and environments. XR-1 introduces the \emph{Unified Vision-Motion Codes (UVMC)}, a discrete latent representation learned via a dual-branch VQ-VAE that jointly encodes visual dynamics and robotic motion. UVMC addresses these challenges by (i) serving as an intermediate representation between the observations and actions, and (ii) aligning multimodal dynamic information from heterogeneous data sources to capture complementary knowledge. To effectively exploit UVMC, we propose a three-stage training paradigm: (i) self-supervised UVMC learning, (ii) UVMC-guided pretraining on large-scale cross-embodiment robotic datasets, and (iii) task-specific post-training. We validate XR-1 through extensive real-world experiments with more than 14,000 rollouts on six different robot embodiments, spanning over 120 diverse manipulation tasks. XR-1 consistently outperforms state-of-the-art baselines such as $π_{0.5}$, $π_0$, RDT, UniVLA, and GR00T-N1.5 while demonstrating strong generalization to novel objects, background variations, distractors, and illumination changes. Our project is at https://xr-1-vla.github.io/.

cs.RO

HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state representation for scalable learning across heterogeneous embodiments, and incorporates a Mixture-of-Experts Unified Proprioceptive Predictor to model whole-body coordination and temporal motion dynamics from large-scale multi-embodiment trajectory data. To efficiently capture temporal visual context, HEX uses lightweight history tokens to summarize past observations, avoiding repeated encoding of historical images during inference. It further employs a residual-gated fusion mechanism with a flow-matching action head to adaptively integrate visual-language cues with proprioceptive dynamics for action generation. Experiments on real-world humanoid manipulation tasks show that HEX achieves state-of-the-art performance in task success rate and generalization, particularly in fast-reaction and long-horizon scenarios.

cs.RO

Depletion-mode N-polar AlN-based high electron mobility transistors with improved on/off ratios

We report N-polar AlN-based high-electron mobility transistors (HEMTs) with a GaN channel thickness of 5.2 nm on N-polar AlN on sapphire. The threshold voltage is around -2.4 to -3.0 V with saturation currents over 240 mA/mm and on/off ratios as high as 10,000, much higher than previously reported N-polar AlN-based HEMTs. The high on/off ratio is attributed to the use of an abrupt AlN/GaN heterostructure with a dedicated AlN transition layer, together with improved gate leakage. The high frequency properties as well as the on-resistance of ~20 Ohm mm are all limited by the 2000 Ohm/square sheet resistance of the channel layer.

cond-mat.mtrl-sci

RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to generalize across long-horizon bimanual tasks and mobile manipulation in unstructured environments remains limited. To bridge this gap, we present RoboMIND 2.0, a comprehensive real-world dataset comprising over 310K dual-arm manipulation trajectories collected across six distinct robot embodiments and 739 complex tasks. Crucially, to support research in contact-rich and spatially extended tasks, the dataset incorporates 12K tactile-enhanced episodes and 20K mobile manipulation trajectories. Complementing this physical data, we construct high-fidelity digital twins of our real-world environments, releasing an additional 20K-trajectory simulated dataset to facilitate robust sim-to-real transfer. To fully exploit the potential of RoboMIND 2.0, we propose MIND-2 system, a hierarchical dual-system frame-work optimized via offline reinforcement learning. MIND-2 integrates a high-level semantic planner (MIND-2-VLM) to decompose abstract natural language instructions into grounded subgoals, coupled with a low-level Vision-Language-Action executor (MIND-2-VLA), which generates precise, proprioception-aware motor actions.

cs.RO

RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation

The pursuit of general-purpose robotic manipulation is hindered by the scarcity of diverse, real-world interaction data. Unlike data collection from web in vision or language, robotic data collection is an active process incurring prohibitive physical costs. Consequently, automated task curation to maximize data value remains a critical yet under-explored challenge. Existing manual methods are unscalable and biased toward common tasks, while off-the-shelf foundation models often hallucinate physically infeasible instructions. To address this, we introduce RoboGene, an agentic framework designed to automate the generation of diverse, physically plausible manipulation tasks across single-arm, dual-arm, and mobile robots. RoboGene integrates three core components: diversity-driven sampling for broad task coverage, self-reflection mechanisms to enforce physical constraints, and human-in-the-loop refinement for continuous improvement. We conduct extensive quantitative analysis and large-scale real-world experiments, collecting datasets of 18k trajectories and introducing novel metrics to assess task quality, feasibility, and diversity. Results demonstrate that RoboGene significantly outperforms state-of-the-art foundation models (e.g., GPT-4o, Gemini 2.5 Pro). Furthermore, real-world experiments show that VLA models pre-trained with RoboGene achieve higher success rates and superior generalization, underscoring the importance of high-quality task generation. Our project is available at https://robogene-boost-vla.github.io.

cs.RO

RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation

Enhancing the generalization capability of robotic learning to enable robots to operate effectively in diverse, unseen scenes is a fundamental and challenging problem. Existing approaches often depend on pretraining with large-scale data collection, which is labor-intensive and time-consuming, or on semantic data augmentation techniques that necessitate an impractical assumption of flawless upstream object detection in real-world scenarios. In this work, we propose RoboAug, a novel generative data augmentation framework that significantly minimizes the reliance on large-scale pretraining and the perfect visual recognition assumption by requiring only the bounding box annotation of a single image during training. Leveraging this minimal information, RoboAug employs pre-trained generative models for precise semantic data augmentation and integrates a plug-and-play region-contrastive loss to help models focus on task-relevant regions, thereby improving generalization and boosting task success rates. We conduct extensive real-world experiments on three robots, namely UR-5e, AgileX, and Tien Kung 2.0, spanning over 35k rollouts. Empirical results demonstrate that RoboAug significantly outperforms state-of-the-art data augmentation baselines. Specifically, when evaluating generalization capabilities in unseen scenes featuring diverse combinations of backgrounds, distractors, and lighting conditions, our method achieves substantial gains over the baseline without augmentation. The success rates increase from 0.09 to 0.47 on UR-5e, from 0.16 to 0.60 on AgileX, and from 0.19 to 0.67 on Tien Kung 2.0. These results highlight the superior generalization and effectiveness of RoboAug in real-world manipulation tasks. Our project is available at https://x-roboaug.github.io/.

cs.RO

Larger Hausdorff Dimension in Scanning Pattern Facilitates Mamba-Based Methods in Low-Light Image Enhancement

We propose an innovative enhancement to the Mamba framework by increasing the Hausdorff dimension of its scanning pattern through a novel Hilbert Selective Scan mechanism. This mechanism explores the feature space more effectively, capturing intricate fine-scale details and improving overall coverage. As a result, it mitigates information inconsistencies while refining spatial locality to better capture subtle local interactions without sacrificing the model's ability to handle long-range dependencies. Extensive experiments on publicly available benchmarks demonstrate that our approach significantly improves both the quantitative metrics and qualitative visual fidelity of existing Mamba-based low-light image enhancement methods, all while reducing computational resource consumption and shortening inference time. We believe that this refined strategy not only advances the state-of-the-art in low-light image enhancement but also holds promise for broader applications in fields that leverage Mamba-based techniques.

cs.CV

Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning

Harmful fine-tuning issues present significant safety challenges for fine-tuning-as-a-service in large language models. Existing alignment-stage defenses, e.g., Vaccine, Repnoise, Booster, and T-Vaccine, mitigate harmful fine-tuning issues by enhancing the model's robustness during the alignment phase. While these methods have been proposed to mitigate the issue, they often overlook a critical upstream factor: the role of the original safety-alignment data. We observe that their defense performance and computational efficiency remain constrained by the quality and composition of the alignment dataset. To address this limitation, we propose Pharmacist, a safety alignment data curation solution that enhances defense against harmful fine-tuning by selecting a high-quality and safety-critical core subset from the original alignment data. The core idea of Pharmacist is to train an alignment data selector to rank alignment data. Specifically, up-ranking high-quality and safety-critical alignment data, down-ranking low-quality and non-safety-critical data. Empirical results indicate that models trained on datasets selected by Pharmacist outperform those trained on datasets selected by existing selection methods in both defense and inference performance. In addition, Pharmacist can be effectively integrated with mainstream alignment-stage defense methods. For example, when applied to RepNoise and T-Vaccine, using the dataset selected by Pharmacist instead of the full dataset leads to improvements in defense performance by 2.60\% and 3.30\%, respectively, and enhances inference performance by 3.50\% and 1.10\%. Notably, it reduces training time by 56.83\% and 57.63\%, respectively. Our code is available at https://github.com/Lslland/Pharmacist.

cs.CR

Modelling and Control of Subsonic Missile for Air-to-Air Interception

Subsonic missiles play an important role in modern air-to-air combat scenarios - utilized by the F-35 Lightning II - but require complex Guidance, Navigation and Control systems to manoeuvre with 30G's of acceleration to intercept successfully. Challenges with mathematically modelling and controlling such a dynamic system must be addressed, high frequency noise rejected, and actuator delay compensated for. This paper aims to investigate the control systems necessary for interception. It also proposes a subsonic design utilizing literature and prior research, suggests aerodynamic derivatives, and analyses a designed 2D reduced pitch autopilot control system response against performances. The pitch autopilot model contains an optimized PID controller, 2nd order actuator, lead compensator and Kalman Filter, that rejects time varying disturbances and high frequency noise expected during flight. Simulation results confirm the effectiveness of the proposed method through reduction in rise time (21%), settle time (10%), and highlighted its high frequency deficiency with respect to the compensator integration. The actuator delay of 100ms has been negated by the augmented compensator autopilot controller so that it exceeds system performance requirements (1) & (3). However, (2) is not satisfied as 370% overshoot exists. This research confirms the importance of a lead compensator in missile GNC systems and furthers control design application through a specific configuration. Future research should build upon methods and models presented to construct and test an interception scenario.

eess.SY

RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation

In this paper, we introduce RoboMIND (Multi-embodiment Intelligence Normative Data for Robot Manipulation), a dataset containing 107k demonstration trajectories across 479 diverse tasks involving 96 object classes. RoboMIND is collected through human teleoperation and encompasses comprehensive robotic-related information, including multi-view observations, proprioceptive robot state information, and linguistic task descriptions. To ensure data consistency and reliability for imitation learning, RoboMIND is built on a unified data collection platform and a standardized protocol, covering four distinct robotic embodiments: the Franka Emika Panda, the UR5e, the AgileX dual-arm robot, and a humanoid robot with dual dexterous hands. Our dataset also includes 5k real-world failure demonstrations, each accompanied by detailed causes, enabling failure reflection and correction during policy learning. Additionally, we created a digital twin environment in the Isaac Sim simulator, replicating the real-world tasks and assets, which facilitates the low-cost collection of additional training data and enables efficient evaluation. To demonstrate the quality and diversity of our dataset, we conducted extensive experiments using various imitation learning methods for single-task settings and state-of-the-art Vision-Language-Action (VLA) models for multi-task scenarios. By leveraging RoboMIND, the VLA models achieved high manipulation success rates and demonstrated strong generalization capabilities. To the best of our knowledge, RoboMIND is the largest multi-embodiment teleoperation dataset collected on a unified platform, providing large-scale and high-quality robotic training data. Our project is at https://x-humanoid-robomind.github.io/.

cs.RO

Unmanned F/A-18 Aircraft Landing Control on Aircraft Carrier in Adverse Conditions

Carrier landing of aircrafts is a challenge for control due to the existence of nonlinear wind disturbances and the requirements of changing reference trajectories. In this paper, a robust landing control system is presented for carrier landing of unmanned F/A-18 aircraft. In the control system, an augmented observer is applied to estimate the combined disturbances in the pitch dynamics of F/A-18 aircraft during carrier landing. Therefore, the control performance is improved through the control compensations from these estimations. Additionally, the controllers are designed to regulate the velocity, rate of descent and vertical position. A full model, including the nonlinear flight dynamics, controller, carrier deck motion, wind and measurement noise, is constructed numerically and implemented in software. Combining the observer with a proportional-derivative (PD) control, the proposed pitch control shows the better transient characteristics and stronger robustness than a proportional-integral-derivative (PID) controller. The simulations verify that the designed control system can make the aircraft quickly track a time-varying reference despite the existence of nonlinear disturbances and noise.

eess.SY

Modelling, design and control of middle-size tilt-rotor quadrotor

This paper explores the mathematical modelling and 3D design of a tilt-rotor quadrotor aircraft. The aircraft is a VTOL design and has capacity for one pilot. The design incorporates a part manual part automatic computerised flight control system and hybrid powertrain providing energy to eight ducted contrarotating propellers. Analysis of controllability was performed using the parameters derived from the developed model for the take-off phase of flight. The aircraft is of a lightweight design and is intended to fill a niche in the aviation market with potential for civilian and military applications. The aircraft has multiple redundancies in the instance of propeller failure. The viability of a hybrid powerplant is explored by combining industry standard gas turbine technology with electrical motor and battery systems. The combination of these systems results in a safe, versatile aircraft that can operate at different levels of automation depending on environmental factors or phase of flight. Simulation confirms the design of the aircraft.

eess.SY

Future state prediction based on observer for missile system

Guided missile accuracy and precision is negatively impacted by seeker delay, more specifically by the delay introduced by a mechanical seeker gimbal and the computational time taken to process the raw data. To meet the demands and expectations of modern missiles systems, the impact of this hardware limitation must be reduced. This paper presents a new observer design that predicts the future state of a seeker signal, augmenting the guidance system to mitigate the effects of this delay. The design is based on a novel two-step differentiator, which produces the estimated future time derivatives of the signal. The input signal can be nonlinear and provides for simple integration into existing systems. A bespoke numerical guided missile simulation is used to demonstrate the performance of the observer within a missile guidance system. Both non-manoeuvring and randomly manoeuvring target engagement scenarios are considered.

eess.SY

Longitudinal dynamic modelling and control for a quad-tilt rotor UAV

Tilt rotor aircraft combine the benefits of both helicopters and fixed wing aircraft, this makes them popular for a variety of applications, including Search and Rescue and VVIP transport. However, due to the multiple flight modes, significant challenges with regards to the control system design are experienced. The main challenges with VTOL aircraft, comes during the dynamic phase (mode transition), where the aircraft transitions from a hover state to full forwards flight. In this transition phase the aerodynamic lift and torque generated by the wing/control surfaces increases and as such, the rotor thrust, and the tilt rate must be carefully considered, such that the height and attitude remain invariant during the mode transition. In this paper, a digital PID controller with the applicable digital filter and data hold functions is designed so that a successful mode transition between hover and forwards flight can be ascertained. Finally, the presented control system for the tilt-rotor UAV is demonstrated through simulations by using the MATLAB software suite. The performance obtained from the simulations confirm the success of the implemented methods, with full stability in all three degrees of freedom being demonstrated.

eess.SY

Robust control for uncertain air-to-air missile systems

Air-to-air missiles are used on many modern military combat aircraft for self-defence. It is imperative for the pilots using the weapons that the missiles hit their target first time. The important goals for a missile control system to achieve are minimising the time constant, overshoot, and settling time of the missile dynamics. The combination of high angles of attack, time-varying mass, thrust, and centre of gravity, actuator delay, and signal noise create a highly non-linear dynamic system with many uncertainties that is extremely challenging to control. A robust control system based on saturated sliding mode control is proposed to overcome the time-varying parameters and non-linearities. A lag compensator is designed to overcome actuator delay. A second-order filter is selected to reduce high-frequency measurement noise. When combined, the proposed solutions can make the system stable despite the existence of changing mass, centre of gravity, thrust, and sensor noise. The system was evaluated for desired pitch angles of 0° to 90°. The time constant for the system stayed below 0.27s for all conditions, with satisfactory performance for both settling time and overshoot.

eess.SY

Mode transition control of large-size tiltrotor aircraft

Tiltrotors are an aircraft concept with the ability to rotate their rotors freely, achieving vertical take-off and fast forward flight. The combination of helicopter and fixed-wing flight into one aircraft provides versatility in mission selection, yet challenges persist in their construction and control. Tiltrotor aircraft can operate in three primary modes: helicopter, fixed-wing, and transition, with the transition mode facilitating the shift between helicopter and fixed-wing flight. However, control within this transition region is inherently challenging due to its non-linear nature, hence tiltrotors have been predominantly limited to military applications. Thus, this paper aims to explore transition mode control for a large-size tiltrotor aircraft, tailored to civil applications. A novel, large-sized, tiltrotor concept is presented, accompanied by a derived mathematical model describing the aircrafts behaviours. A PID control method has been used to control the height, pitch, and velocity variations within the transition mode with secondary control loop developed to control the tilt angle during transition. The derived model and control are then implemented within a MATLAB simulation, where the control method was iterated to improve performance. The results show a full transition was achieved in under 14 seconds, where altitude variations were kept below 10 metres. Though the transition mode control was successful, a collective look at the data showcases issues with assumptions as well as thrust discontinuities. The implications of these results are discussed, with suggested improvements proposed for future work.

eess.SY