SearcharxivSearch

arXiv subjects

Yandong Wang

Publications and source records attributed to Yandong Wang.

11 recordsLinked to original sources

Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems

Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end learning, leveraging large vision-language backbones with continuous action prediction. However, as bimanual tasks become tightly coupled and execution constraints become critical, implicit coordination alone is insufficient to ensure reliable, interpretable, and stable behavior. In this work, we propose Co-VLA, a coordination-aware bimanual manipulation framework introducing explicit structural priors into VLA models. We instantiate our method on a state-of-the-art vision-language backbone by replacing its monolithic action head with a Structured Action Expert (SAE) designed for bimanual coordination. Specifically, we introduce explicit structure at the action generation level with a modular coordination-aware loss that shapes shared and residual latents according to task-specific structures. The shared latent encodes task-level coordination intent, while residual latents capture execution adjustments for each arm. At deployment, a Latent-Aware Controller (LAC) interprets the learned representations to modulate synchronization strength, execution asymmetry, smoothness, and safety constraints in real time. LAC operates at the joint-command level and remains compatible with standard control pipelines without requiring force or impedance control. Experiments across simulation and real-world benchmarks show Co-VLA significantly outperforms monolithic baselines, achieving a 27% success rate gain in tight-coordination tasks, more than doubling performance in OOD real-world scenarios (from 13% to 27%), and reducing task completion time by up to 25%.

cs.RO

TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression

The rapid progress of large language models (LLMs) has led to remarkable performance gains across a wide range of tasks. However, when handling long documents that exceed the model's context window limit, the entire context cannot be processed in a single pass, making chunk-wise processing necessary. This requires multiple turns to read different chunks and update memory. However, supervision is typically provided only by the final outcome, which makes it difficult to evaluate the quality of memory updates at each turn in the multi-turn training setting. This introduces a temporal credit assignment challenge. Existing approaches, such as LLM-as-a-judge or process reward models, incur substantial computational overhead and suffer from estimation noise. To better address the credit assignment problem in multi-turn memory training, we propose Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning (TAMTRL). TAMTRL leverages relevant documents as teacher signals by aligning them with each turn of model input and assigns rewards through normalized probabilities in a self-supervised manner. This provides fine-grained learning signals for each memory update and improves long-context processing. Experiments with multiple models of varying scales across seven long-context benchmarks show that TAMTRL consistently outperforms strong baselines, demonstrating its effectiveness. Our code is available at https://anonymous.4open.science/r/TAMTRL-F1F8.

cs.CL

Dynamical Origin of (469219) Kamo`oalewa of Tianwen-2 Mission from the Main-Belt: $ν_6$ Secular Resonance, Flora Family or 3:1 Resonance with Jupiter

China's Tianwen-2 mission, launched on 29 May 2025, targets the near-Earth object (469219) Kamo`oalewa, an Earth quasi-satellite trapped in a 1:1 mean-motion resonance with our planet. Determining the origin of Kamo`oalewa is central to understanding the formation pathways and dynamical evolution of Earth's quasi-satellite population. Here we show a strong possibility of main-belt origin for Kamo`oalewa using long-term dynamical simulations. We examine three candidate source regions: the $ν_6$ secular resonance ($ν_6$), the 3:1 mean-motion resonance with Jupiter (3:1J MMR), and the Flora family. A total of 42,825 test particles were integrated over 100 Myr. We find that asteroids from all three regions can be transported onto Kamo`oalewa-like orbits, albeit with markedly different efficiencies. Particles originating near the $ν_6$ show the highest transfer probability (3.31%), followed by the Flora family (2.54%) and the 3:1J MMR (0.39%). We further identify representative dynamical pathways linking these source regions to Earth quasi-satellite orbits. The Tianwen-2 spacecraft is expected to rendezvous with Kamo`oalewa in 2026, performing close-proximity operations and returning samples. The mission will provide decisive observational constraints on the asteroid's composition and physical properties, offering a critical test of its proposed origin.

astro-ph.EP

Chern numbers of topological phonon band crossing determined with inelastic neutron scattering

Topological invariants in the band structure, such as Chern numbers, are crucial for the classification of topological matters and dictate the occurrence of exotic properties, yet their direct spectroscopic determination has been largely limited to electronic bands. Here, we use inelastic neutron scattering in conjunction with ab initio calculations to identify a variety of topological phonon band crossings in MnSi and CoSi single crystals. We find a distinct relation between the Chern numbers of a band-crossing node and the scattering intensity modulation in momentum space around the node. Given sufficiently high resolution, our method can be used to determine arbitrarily large Chern numbers of topological phonon band-crossing nodes.

cond-mat.mes-hall

Magnetic molecular orbitals in MnSi

A large body of knowledge about magnetism is attained from models of interacting spins, which usually reside on magnetic ions. Proposals beyond the ionic picture are uncommon and seldom verified by direct observations in conjunction with microscopic theory. Here, using inelastic neutron scattering to study the itinerant near-ferromagnet MnSi, we find that the system's fundamental magnetic units are interconnected, extended molecular orbitals consisting of three Mn atoms each, rather than individual Mn atoms. This result is further corroborated by magnetic Wannier orbitals obtained by ab initio calculations. It contrasts the ionic picture with a concrete example, and presents a novel regime of the spin waves where the wavelength is comparable to the spatial extent of the molecular orbitals. Our discovery brings important insights into not only the magnetism of MnSi, but also a broad range of magnetic quantum materials where structural symmetry, electron itinerancy and correlations act in concert.

cond-mat.str-el

Identifying Cause-and-Effect Relationships of Manufacturing Errors using Sequence-to-Sequence Learning

In car-body production the pre-formed sheet metal parts of the body are assembled on fully-automated production lines. The body passes through multiple stations in succession, and is processed according to the order requirements. The timely completion of orders depends on the individual station-based operations concluding within their scheduled cycle times. If an error occurs in one station, it can have a knock-on effect, resulting in delays on the downstream stations. To the best of our knowledge, there exist no methods for automatically distinguishing between source and knock-on errors in this setting, as well as establishing a causal relation between them. Utilizing real-time information about conditions collected by a production data acquisition system, we propose a novel vehicle manufacturing analysis system, which uses deep learning to establish a link between source and knock-on errors. We benchmark three sequence-to-sequence models, and introduce a novel composite time-weighted action metric for evaluating models in this context. We evaluate our framework on a real-world car production dataset recorded by Volkswagen Commercial Vehicles. Surprisingly we find that 71.68% of sequences contain either a source or knock-on error. With respect to seq2seq model training, we find that the Transformer demonstrates a better performance compared to LSTM and GRU in this domain, in particular when the prediction range with respect to the durations of future actions is increased.

cs.LG

In-situ synchrotron based high energy X-ray diffraction study of the deformation mechanism of δ-hydride in a commercially pure titanium

We used by in-situ high energy X-ray diffraction to inestigate the deformation behavior of Grade 2 commercially pure titanium that was hydrogen charged to form hydrides. The results showed that the peak broadening in the diffraction patterns are due to the high internal and interphase stresses generated within and around hydrides due to the volume expansion induced by the phase transformation. The hydrides exhibit typical high strength but brittle secondary phase behavior, which undertakes more elastic strain than matrix and is the location where cracks are first generated. Interestingly, the δ-hydrides sustain larger strains than the matrix, especially after the matrix yields. This study on the deformation mechanism of hydrides in pure titanium provides insight into the hydride deformation behavior and hydrogen embrittlement in both titanium and zirconium.

cond-mat.mtrl-sci

Record high $T_{\rm c}$ and robust superconductivity in transition metal $δ$-Ti phase at megabar pressure

We report a record high superconducting transition temperature ($T_{\rm c}$) up to 23.6 K under high pressure in the elemental metal Ti, one of the top ten most abundant elements in Earth's crust. The $T_{\rm c}$ increases monotonically from 2.3 K at 40.3 GPa to 23.6 K at 144.9 GPa, which surpasses all known records from elemental metals reported so far. With further compression, a robust $T_{\rm c}$ of ~23 K is observed between 144.9 and 183 GPa in the $δ$-Ti phase. The pressure-dependent $T_{\rm c}$ can be well described by the conventional electron-phonon coupling (EPC) mechanism. Density Functional Theory calculations show the Fermi nesting and the phonon softening of optical branches at the $γ$-Ti to $δ$-Ti phase transition pressure enhance EPC, which results in the record high $T_{\rm c}$. We attribute the robust superconductivity in $δ$-Ti to the apparent robustness of its strong EPC against lattice compression. These results provide new insight into exploring new high-$T_{\rm c}$ elemental metals and Ti-based superconducting alloys.

cond-mat.supr-con

Distributed Localization without Direct Communication Inspired by Statistical Mechanics

Distributed localization is essential in many robotic collective tasks such as shape formation and self-assembly.Inspired by the statistical mechanics of energy transition, this paper presents a fully distributed localization algorithm named as virtual particle exchange (VPE) localization algorithm, where each robot repetitively exchanges virtual particles (VPs) with neighbors and eventually obtains its relative position from the virtual particle (VP) amount it owns. Using custom-designed hardware and protocol, VPE localization algorithm allows robots to achieve localization using sensor readings only, avoiding direct communication with neighbors and keeping anonymity. Moreover, VPE localization algorithm determines the swarm center automatically, thereby eliminating the requirement of fixed beacons to embody the origin of coordinates. Theoretical analysis proves that the VPE localization algorithm can always converge to the same result regardless of initial state and has low asymptotic time and memory complexity. Extensive localization simulations with up to 10000 robots and experiments with 52 lowcost robots are carried out, which verify that VPE localization algorithm is scalable, accurate and robust to sensor noises. Based on the VPE localization algorithm, shape formations are further achieved in both simulations and experiments with 52 robots, illustrating that the algorithm can be directly applied to support swarm collaborative tasks.

cs.RO

GaDei: On Scale-up Training As A Service For Deep Learning

Deep learning (DL) training-as-a-service (TaaS) is an important emerging industrial workload. The unique challenge of TaaS is that it must satisfy a wide range of customers who have no experience and resources to tune DL hyper-parameters, and meticulous tuning for each user's dataset is prohibitively expensive. Therefore, TaaS hyper-parameters must be fixed with values that are applicable to all users. IBM Watson Natural Language Classifier (NLC) service, the most popular IBM cognitive service used by thousands of enterprise-level clients around the globe, is a typical TaaS service. By evaluating the NLC workloads, we show that only the conservative hyper-parameter setup (e.g., small mini-batch size and small learning rate) can guarantee acceptable model accuracy for a wide range of customers. We further justify theoretically why such a setup guarantees better model convergence in general. Unfortunately, the small mini-batch size causes a high volume of communication traffic in a parameter-server based system. We characterize the high communication bandwidth requirement of TaaS using representative industrial deep learning workloads and demonstrate that none of the state-of-the-art scale-up or scale-out solutions can satisfy such a requirement. We then present GaDei, an optimized shared-memory based scale-up parameter server design. We prove that the designed protocol is deadlock-free and it processes each gradient exactly once. Our implementation is evaluated on both commercial benchmarks and public benchmarks to demonstrate that it significantly outperforms the state-of-the-art parameter-server based implementation while maintaining the required accuracy and our implementation reaches near the best possible runtime performance, constrained only by the hardware limitation. Furthermore, to the best of our knowledge, GaDei is the only scale-up DL system that provides fault-tolerance.

stat.ML

IBM Deep Learning Service

Deep learning driven by large neural network models is overtaking traditional machine learning methods for understanding unstructured and perceptual data domains such as speech, text, and vision. At the same time, the "as-a-Service"-based business model on the cloud is fundamentally transforming the information technology industry. These two trends: deep learning, and "as-a-service" are colliding to give rise to a new business model for cognitive application delivery: deep learning as a service in the cloud. In this paper, we will discuss the details of the software architecture behind IBM's deep learning as a service (DLaaS). DLaaS provides developers the flexibility to use popular deep learning libraries such as Caffe, Torch and TensorFlow, in the cloud in a scalable and resilient manner with minimal effort. The platform uses a distribution and orchestration layer that facilitates learning from a large amount of data in a reasonable amount of time across compute nodes. A resource provisioning layer enables flexible job management on heterogeneous resources, such as graphics processing units (GPUs) and central processing units (CPUs), in an infrastructure as a service (IaaS) cloud.

cs.DC