SearcharxivSearch

arXiv subjects

Wanting Xu

Publications and source records attributed to Wanting Xu.

At least 19 recordsLinked to original sources

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models

The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or batch-workload submission, typically allocate an exclusive set of GPU and CPU resources to a single tenant. While this paradigm maximizes client flexibility, it burdens users with infrastructure adaptation, and the fixed card-hour accounting model renders short or bursty workloads both expensive for tenants and inefficient for the service provider. To address these challenges, we present JoyNexus, a unified service for multi-tenant VLA supervised fine-tuning, reinforcement learning, and evaluation. JoyNexus decouples the Training Model Service, Inference Model Service, and Environment Service, each accessed through APIs and backed by resident shared base models with tenant-specific slots. Tenants can directly invoke high-level semantic APIs for training, rollout, and evaluation, or compose custom algorithms using lower-level APIs and their assigned endpoints. Multiple tenants submit workloads concurrently; their action modules, optimizers, rollout records, and policy versions remain isolated, and the service is scheduled by the global Training Queue and Inference Queue. To further improve multi-tenant training efficiency, JoyNexus introduces group batching for heterogeneous VLA data schemas that share a compatible model-facing prefix, enabling a single shared backbone forward pass over grouped samples. Finally, we evaluate JoyNexus through workload simulation and a group-batching pipeline in a realistic embodied scenario. Results show that, compared with isolated single-tenant execution, JoyNexus reduces aggregate GPU time and improves service utilization via cross-tenant scheduling on shared resources.

cs.DC

Tailoring pure valley-Zeeman spin-orbit coupling in WSe$_2$-encapsulated monolayer graphene

Engineering proximity effects in twisted van der Waals heterostructures offers a powerful platform for designing electronic properties. While theoretical predictions of quantum interference in transition metal dichalcogenide-encapsulated graphene can selectively control the spin-orbit coupling component, experimental realizations have remained elusive. Here, we report pure valley-Zeeman spin-orbit coupling in monolayer graphene, achieved by encapsulation between two parallel twisted WSe$_2$ monolayers. We observed a symmetry-enforced reordering of Landau levels, which is driven by the competition between the fixed valley-Zeeman energy and the magnetic-field-dependent cyclotron energy. This reordering is characterized by a transition from symmetry-broken states in the quantum Hall effect to a restored fourfold degeneracy with integer or half-integer quantum Hall sequences. We also demonstrate the ability to completely quench the proximity spin-orbit coupling by tuning the encapsulated geometry.

cond-mat.mes-hall

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training

In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.

cs.DC

JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy

Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large differences across robot embodiments impede effective behavior knowledge transfer. To address these challenges, we propose JoyAI-RA, a vision-language-action (VLA) embodied foundation model tailored for generalizable robotic manipulation. JoyAI-RA presents a multi-source multi-level pretraining framework that integrates web data, large-scale egocentric human manipulation videos, simulation-generated trajectories, and real-robot data. Through training on heterogeneous multi-source data with explicit action-space unification, JoyAI-RA effectively bridges embodiment gaps, particularly between human manipulation and robotic control, thereby enhancing cross-embodiment behavior learning. JoyAI-RA outperforms state-of-the-art methods in both simulation and real-world benchmarks, especially on diverse tasks with generalization demands.

cs.RO

Thousand-GPU Large-Scale Training and Optimization Recipe for AI-Native Cloud Embodied Intelligence Infrastructure

Embodied intelligence is a key step towards Artificial General Intelligence (AGI), yet its development faces multiple challenges including data, frameworks, infrastructure, and evaluation systems. To address these issues, we have, for the first time in the industry, launched a cloud-based, thousand-GPU distributed training platform for embodied intelligence, built upon the widely adopted LeRobot framework, and have systematically overcome bottlenecks across the entire pipeline. At the data layer, we have restructured the data pipeline to optimize the flow of embodied training data. In terms of training, for the GR00T-N1.5 model, utilizing thousand-GPU clusters and data at the scale of hundreds of millions, the single-round training time has been reduced from 15 hours to just 22 minutes, achieving a 40-fold speedup. At the model layer, by combining variable-length FlashAttention and Data Packing, we have moved from sample redundancy to sequence integration, resulting in a 188% speed increase; {\pi}-0.5 attention optimization has accelerated training by 165%; and FP8 quantization has delivered a 140% speedup. On the infrastructure side, relying on high-performance storage, a 3.2T RDMA network, and a Ray-driven elastic AI data lake, we have achieved deep synergy among data, storage, communication, and computation. We have also built an end-to-end evaluation system, creating a closed loop from training to simulation to assessment. This framework has already been fully validated on thousand-GPU clusters, laying a crucial technical foundation for the development and application of next-generation autonomous intelligent robots, and is expected to accelerate the arrival of the era of human-machine integration.

cs.RO

Metric, inertially aligned monocular state estimation via kinetodynamic priors

Accurate state estimation for flexible robotic systems poses significant challenges, particularly for platforms with dynamically deforming structures that invalidate rigid-body assumptions. This paper addresses this problem and enables the extension of existing rigid-body pose estimation methods to non-rigid systems. Our approach integrates two core components: first, we capture elastic properties using a deformation-force model, efficiently learned via a Multi-Layer Perceptron; second, we resolve the platform's inherently smooth motion using continuous-time B-spline kinematic models. By continuously applying Newton's Second Law, our method formulates the relationship between visually-derived trajectory acceleration and predicted deformation-induced acceleration. We demonstrate that our approach not only enables robust and accurate pose estimation on non-rigid platforms, but also shows that the properly modeled platform physics allow for the recovery of inertial sensing properties. We validate this feasibility on a simple spring-camera system, showing how it robustly resolves the typically ill-posed problem of metric scale and gravity recovery in monocular visual odometry.

cs.RO

Chemical vapor deposition growth of continuous monolayer antiferromagnetic CrOCl films

The discovery of two-dimensional magnetic materials has provided an ideal platform for exploring physical phenomena in the two-dimensional limit. However, intrinsic two-dimensional antiferromagnetic materials have been rarely reported, limiting systematic studies of their electronic properties. The discovery of novel intrinsic two-dimensional antiferromagnets and the development of robust synthesis strategies, therefore, remain significant challenges. Here, we report the chemical vapor deposition synthesis of CrOCl monolayer films and nanosheets that exhibit excellent air stability. The CrOCl morphology is tunable, ranging from two-dimensional nanosheets to three-dimensional flower-like structures, with lateral sizes ranging from several microns to continuous monolayer films. Structural characterization confirms the materials composition and high crystalline quality. Furthermore, magnetic measurements, supported by theoretical calculations, reveal a N\'eel temperature for CrOCl of ~14 K. This work provides a reliable route for preparing two-dimensional antiferromagnetic materials.

cond-mat.mes-hall

The interplay of ferroelectricity and magneto-transport in non-magnetic moir\'{e} superlattices

The coupling of ferroelectricity and magnetic order provides rich tunability for engineering material properties and demonstrates great potential for uncovering novel quantum phenomena and multifunctional devices. Here, we report interfacial ferroelectricity in moir\'{e} superlattices constructed from graphene and hexagonal boron nitride. We observe ferroelectric polarization in an across-layer moir\'{e} superlattice with an intercalated layer, demonstrating a remnant polarization comparable to its non-intercalated counterpart. Remarkably, we reveal a magnetic-field enhancement of ferroelectric polarization that persists up to room temperature, showcasing an unconventional amplification of ferroelectricity in materials lacking magnetic elements. This phenomenon, consistent across devices with varying layer configurations, arises purely from electronic rather than ionic contributions. Furthermore, the ferroelectric polarization in turn modulates quantum transport characteristics, suppressing Shubnikov-de Haas oscillations and altering quantum Hall states in polarized phases. This interplay between ferroelectricity and magneto-transport in non-magnetic materials is crucial for exploring magnetoelectric effects and advancing two-dimensional memory and logic applications.

cond-mat.mtrl-sci

Anomalous Meets Topological Hall Effect in Cr2Ge2Te6 Heterostructures

Introducing topologically protected skyrmions in graphene holds significant importance for developing high-speed, low-energy spintronic devices. Here, we present a centrosymmetric ferromagnetic graphene/trilayer Cr2Ge2Te6/graphene heterostructure, demonstrating the anomalous and topological Hall effect due to the magnetic proximity effect. Through gate voltage control, we effectively tune the emergence and size of skyrmions. Micromagnetic simulations reveal the formation of skyrmions and antiskyrmions, which respond differently to external magnetic fields, leading to oscillations in the topological Hall signal. Our findings provide a novel pathway for the formation and manipulation of skyrmions in centrosymmetric two-dimensional magnetic systems, offering significant insights for developing topological spintronics.

cond-mat.mes-hall

Simulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric Vision

While automatically generated polynomial elimination templates have sparked great progress in the field of 3D computer vision, there remain many problems for which the degree of the constraints or the number of unknowns leads to intractability. In recent years, homotopy continuation has been introduced as a plausible alternative. However, the method currently depends on expensive parallel tracking of all possible solutions in the complex domain, or a classification network for starting problem-solution pairs trained over a limited set of real-world examples. Our innovation lies in a novel approach to finding solution-problem pairs, where we only need to predict a rough initial solution, with the corresponding problem generated by an online simulator. Subsequently, homotopy continuation is applied to track that single solution back to the original problem. We apply this elegant combination to generalized camera resectioning, and also introduce a new solution to the challenging generalized relative pose and scale problem. As demonstrated, the proposed method successfully compensates the raw error committed by the regressor alone, and leads to state-of-the-art efficiency and success rates.

cs.CV

Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model

We introduce Xmodel-VLM, a cutting-edge multimodal vision language model. It is designed for efficient deployment on consumer GPU servers. Our work directly confronts a pivotal industry issue by grappling with the prohibitive service costs that hinder the broad adoption of large-scale multimodal systems. Through rigorous training, we have developed a 1B-scale language model from the ground up, employing the LLaVA paradigm for modal alignment. The result, which we call Xmodel-VLM, is a lightweight yet powerful multimodal vision language model. Extensive testing across numerous classic multimodal benchmarks has revealed that despite its smaller size and faster execution, Xmodel-VLM delivers performance comparable to that of larger models. Our model checkpoints and code are publicly available on GitHub at https://github.com/XiaoduoAILab/XmodelVLM.

cs.CV

Tight Fusion of Events and Inertial Measurements for Direct Velocity Estimation

Traditional visual-inertial state estimation targets absolute camera poses and spatial landmark locations while first-order kinematics are typically resolved as an implicitly estimated sub-state. However, this poses a risk in velocity-based control scenarios, as the quality of the estimation of kinematics depends on the stability of absolute camera and landmark coordinates estimation. To address this issue, we propose a novel solution to tight visual-inertial fusion directly at the level of first-order kinematics by employing a dynamic vision sensor instead of a normal camera. More specifically, we leverage trifocal tensor geometry to establish an incidence relation that directly depends on events and camera velocity, and demonstrate how velocity estimates in highly dynamic situations can be obtained over short time intervals. Noise and outliers are dealt with using a nested two-layer RANSAC scheme. Additionally, smooth velocity signals are obtained from a tight fusion with pre-integrated inertial signals using a sliding window optimizer. Experiments on both simulated and real data demonstrate that the proposed tight event-inertial fusion leads to continuous and reliable velocity estimation in highly dynamic scenarios independently of absolute coordinates. Furthermore, in extreme cases, it achieves more stable and more accurate estimation of kinematics than traditional, point-position-based visual-inertial odometry.

cs.CV

Online Stability Improvement of Groebner Basis Solvers using Deep Learning

Over the past decade, the Gr\"obner basis theory and automatic solver generation have lead to a large number of solutions to geometric vision problems. In practically all cases, the derived solvers apply a fixed elimination template to calculate the Gr\"obner basis and thereby identify the zero-dimensional variety of the original polynomial constraints. However, it is clear that different variable or monomial orderings lead to different elimination templates, and we show that they may present a large variability in accuracy for a certain instance of a problem. The present paper has two contributions. We first show that for a common class of problems in geometric vision, variable reordering simply translates into a permutation of the columns of the initial coefficient matrix, and that -- as a result -- one and the same elimination template can be reused in different ways, each one leading to potentially different accuracy. We then prove that the original set of coefficients may contain sufficient information to train a classifier for online selection of a good solver, most notably at the cost of only a small computational overhead. We demonstrate wide applicability at the hand of generic dense polynomial problem solvers, as well as a concrete solver from geometric vision.

cs.CV

Event-Based Visual Odometry on Non-Holonomic Ground Vehicles

Despite the promise of superior performance under challenging conditions, event-based motion estimation remains a hard problem owing to the difficulty of extracting and tracking stable features from event streams. In order to robustify the estimation, it is generally believed that fusion with other sensors is a requirement. In this work, we demonstrate reliable, purely event-based visual odometry on planar ground vehicles by employing the constrained non-holonomic motion model of Ackermann steering platforms. We extend single feature n-linearities for regular frame-based cameras to the case of quasi time-continuous event-tracks, and achieve a polynomial form via variable degree Taylor expansions. Robust averaging over multiple event tracks is simply achieved via histogram voting. As demonstrated on both simulated and real data, our algorithm achieves accurate and robust estimates of the vehicle's instantaneous rotational velocity, and thus results that are comparable to the delta rotations obtained by frame-based sensors under normal conditions. We furthermore significantly outperform the more traditional alternatives in challenging illumination scenarios. The code is available at \url{https://github.com/gowanting/NHEVO}.

cs.CV

Cross-Modal Semi-Dense 6-DoF Tracking of an Event Camera in Challenging Conditions

Vision-based localization is a cost-effective and thus attractive solution for many intelligent mobile platforms. However, its accuracy and especially robustness still suffer from low illumination conditions, illumination changes, and aggressive motion. Event-based cameras are bio-inspired visual sensors that perform well in HDR conditions and have high temporal resolution, and thus provide an interesting alternative in such challenging scenarios. While purely event-based solutions currently do not yet produce satisfying mapping results, the present work demonstrates the feasibility of purely event-based tracking if an alternative sensor is permitted for mapping. The method relies on geometric 3D-2D registration of semi-dense maps and events, and achieves highly reliable and accurate cross-modal tracking results. Practically relevant scenarios are given by depth camera-supported tracking or map-based localization with a semi-dense map prior created by a regular image-based visual SLAM or structure-from-motion system. Conventional edge-based 3D-2D alignment is extended by a novel polarity-aware registration that makes use of signed time-surface maps (STSM) obtained from event streams. We furthermore introduce a novel culling strategy for occluded points. Both modifications increase the speed of the tracker and its robustness against occlusions or large view-point variations. The approach is validated on many real datasets covering the above-mentioned challenging conditions, and compared against similar solutions realised with regular cameras.

cs.RO

Self-passivated freestanding superconducting oxide film for flexible electronics

The integration of high-temperature superconducting YBa2Cu3O6+x (YBCO) into flexible electronic devices has the potential to revolutionize the technology industry. The effective preparation of high-quality flexible YBCO films therefore plays a key role in this development. We present a novel approach for transferring water-sensitive YBCO films onto flexible substrates without any buffer layer. Freestanding YBCO film on a polydimethylsiloxane substrate is extracted by etching the Sr3Al2O6 sacrificial layer from the LaAlO3 substrate. In addition to the obtained freestanding YBCO thin film having a Tc of 89.1 K, the freestanding YBCO thin films under inward and outward bending conditions have Tc of 89.6 K and 88.9 K, respectively. A comprehensive characterization involving multiple experimental techniques including high-resolution transmission electron microscopy, scanning electron microscopy, Raman and X-ray Absorption Spectroscopy is conducted to investigate the morphology, structural and electronic properties of the YBCO film before and after the extraction process where it shows the preservation of the structural and superconductive properties of the freestanding YBCO virtually in its pristine state. Further investigation reveals the formation of a YBCO passivated layer serves as a protective layer which effectively preserves the inner section of the freestanding YBCO during the etching process. This work plays a key role in actualizing the fabrication of flexible oxide thin films and opens up new possibilities for a diverse range of device applications involving thin-films and low-dimensional materials.

cond-mat.supr-con

Accelerating Globally Optimal Consensus Maximization in Geometric Vision

Branch-and-bound-based consensus maximization stands out due to its important ability of retrieving the globally optimal solution to outlier-affected geometric problems. However, while the discovery of such solutions caries high scientific value, its application in practical scenarios is often prohibited by its computational complexity growing exponentially as a function of the dimensionality of the problem at hand. In this work, we convey a novel, general technique that allows us to branch over an n-1 dimensional space for an n-dimensional problem. The remaining degree of freedom can be solved globally optimally within each bound calculation by applying the efficient interval stabbing technique. While each individual bound derivation is harder to compute owing to the additional need for solving a sorting problem, the reduced number of intervals and tighter bounds in practice lead to a significant reduction in the overall number of required iterations. Besides an abstract introduction of the approach, we present applications to four fundamental geometric computer vision problems: camera resectioning, relative camera pose estimation, point set registration, and rotation and focal length estimation. Through our exhaustive tests, we demonstrate significant speed-up factors at times exceeding two orders of magnitude, thereby increasing the viability of globally optimal consensus maximizers in online application scenarios.

cs.CV

Observing two-photon subwavelength interference of broadband chaotic light in polarization-selective Michelson interferometer

Differing from the traditional method of achieving subwavelength interference, we have demonstrated the two-photon subwavelength interference effect of broadband chaotic light in a polarization-selective Michelson interferometer with an ultrafast two-photon absorption detector the first time, which is achieved by manipulating two-photon probability amplitudes involved in the interference. In theory, the two-photon polarization coherence matrix and probability amplitudes matrix are combined to develop polarized two-photon interference terms, which explains the experimental results well. In order to make better use of this interferometer to produce the subwavelength effect, we also make a series of error analyses to find out the relationship between the visibility and the degree of polarization error. Our experimental and theoretical results are helpful to understand the two-photon subwavelength interference, which sheds light on the development of the two-photon interference theory of vector light field based on quantum mechanics. These experimental results may help to develop future optical interferometry, optical polarimetry, and subwavelength lithography.

quant-ph