SearcharxivSearch

arXiv subjects

Menglin Zou

Publications and source records attributed to Menglin Zou.

7 recordsLinked to original sources

RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robot control. However, their performance remains fundamentally constrained by the availability of high-quality robot trajectory data. In current robot learning practice, such data are primarily collected through human teleoperation, which is labor-intensive, costly, and difficult to scale. In this paper, we propose RDGen, a sim-to-real reinforcement learning framework for generating high-quality robot demonstrations. Rather than employing reinforcement learning solely as the final control policy, RDGen leverages trained RL policies as a structured trajectory generator. The system consists of a VLM-based task parser that identifies task-relevant objects, a Grounding DINO-based object localizer, and an RL policy transferred from simulation to the real robot. Successful rollouts are then harvested as clean, high-quality demonstrations for downstream VLA training, while the simulation stage further provides a scalable source of additional trajectories at little marginal cost. Experiments on a pick-and-place task demonstrate that the transferred RL policy achieves a high task success rate. Compared with human teleoperation, RDGen produces significantly smoother trajectories and yields superior downstream VLA performance. These results indicate that RL-generated demonstrations can serve as more reliable and consistent supervisory signals for robot policy learning.

cs.RO

ROI-Driven Foveated Attention for Unified Egocentric Representations in Vision-Language-Action Systems

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data collection cost, limited cross-embodiment alignment, and poor transfer from internet-scale visual data to robot control. We propose a region-of-interest (ROI) driven engineering workflow that introduces an egocentric, geometry-grounded data representation. By projecting end-effector poses via forward kinematics (FK) into a single external camera, we derive movement-aligned hand-centric ROIs without requiring wrist-mounted cameras or multi-view systems. Unlike directly downsampling the full frame, ROI is cropped from the original image before resizing, preserving high local information density for contact-critical regions while retaining global context. We present a reproducible pipeline covering calibration, synchronization, ROI generation, deterministic boundary handling, and metadata governance. The resulting representation is embodiment-aligned and viewpoint-normalized, enabling data reuse across heterogeneous robots. We argue that egocentric ROI serves as a practical data abstraction for scalable collection and cross-embodiment learning, bridging internet-scale perception and robot-specific control.

cs.RO

Beyond-Ten-Hour Coherence in a Decoherence-Free Trapped-Ion Clock Qubit

Quantum systems promise to revolutionize information processing science and technology [1-3]. The preservation of quantum coherence, the defining property of qubits, fundamentally constrains the performance of quantum information processing with quantum memories [4]. While trapped atomic ions theoretically support million-year coherence based on spontaneous emission [5-7], experimental demonstrations have reached far less, only about an hour [8-13]. Here we combine clock-state qubits with decoherence-free subspace (DFS) encoding to achieve coherence exceeding ten hours. Using correlation-based phase tracking in 171Yb+ ion pairs sympathetically cooled by 138Ba+ ion, we demonstrate this without magnetic shielding or enhanced microwave phase stabilization that previously limited coherence times. DFS encoding references the qubit phase to the inter-ion energy difference to reject microwave phase noise and common-mode magnetic fluctuations, while clock states provide environmental insensitivity. Throughout measurements extended to 1600 seconds, we observe minimal coherence decay, with exponential fits yielding a coherence time of (3.77 +/- 1.09) x 10^4 seconds. Our results establish DFS encoding as a form of passive error correction that eliminates technical noise constraints, unlocking the million-year coherence potential of atomic ions for scalable quantum information processing.

quant-ph

SaiVLA-0: Cerebrum--Pons--Cerebellum Tripartite Architecture for Compute-Aware Vision-Language-Action

We revisit Vision-Language-Action through a neuroscience-inspired triad. Biologically, the Cerebrum provides stable high-level multimodal priors and remains frozen; the Pons Adapter integrates these cortical features with real-time proprioceptive inputs and compiles intent into execution-ready tokens; and the Cerebellum (ParaCAT) performs fast, parallel categorical decoding for online control, with hysteresis/EMA/temperature/entropy for stability. A fixed-ratio schedule and two-stage feature caching make the system compute-aware and reproducible. Inspired by active, foveated vision, our wrist ROIs are geometrically tied to the end-effector via calibrated projection, providing a movement-stabilized, high-resolution view that is sensitive to fine-grained pose changes and complements the global context of the main view. The design is modular: upgrading the Cerebrum only retrains the Pons; changing robots only trains the Cerebellum; cerebellum-only RL can further refine control without touching high-level semantics. As a concept-and-protocol paper with preliminary evidence, we outline a timing protocol under matched conditions (GPU, resolution, batch) to verify anticipated efficiency gains. We also report preliminary LIBERO evidence showing that split feature caching reduces training time (7.5h to 4.5h) and improves average success (86.5% to 92.5%) under official N1.5 head-only training, and that SaiVLA0 reaches 99.0% mean success.

cs.RO

Realization of Trapped Ion Dynamics in the Strong-Field Regime and Non-Markovianity

We experimentally investigate trapped ion dynamics in the strong-driving regime, where the Rabi frequency (Omega) is comparable to the vibrational mode frequency (nu). In the conventional weak-driving regime (Omega << nu), the dynamics is well described by effective Hamiltonians for the carrier and motional sidebands, associated with detunings (delta = n nu (n = 0, +/- 1, ...)). In the strong-driving regime (Omega ~ nu), these interactions can no longer be treated independently. We characterize this physics through the reduced dynamics of the qubit, where non-Markovian behavior emerges as an operational probe of the spin-motion coupling. We observe a structured non-Markovian response, with well-defined maxima that follow the generalized resonance condition (delta^2 + Omega^2 = nu^2), reflecting the strong underlying spin-motion hybridization characteristic of the strong-driving regime.

quant-ph

Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm

Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to unstable training. Recent non-adversarial IRL approaches improve stability by jointly learning reward and policy via energy-based formulations but lack formal guarantees. This work bridges this gap. We first present a unified view showing canonical non-adversarial methods explicitly or implicitly maximize the likelihood of expert behavior, which is equivalent to minimizing the expected return gap. This insight leads to our main contribution: Trust Region Reward Optimization (TRRO), a framework that guarantees monotonic improvement in this likelihood via a Minorization-Maximization process. We instantiate TRRO into Proximal Inverse Reward Optimization (PIRO), a practical and stable IRL algorithm. Theoretically, TRRO provides the IRL counterpart to the stability guarantees of Trust Region Policy Optimization (TRPO) in forward RL. Empirically, PIRO matches or surpasses state-of-the-art baselines in reward recovery, policy imitation with high sample efficiency on MuJoCo and Gym-Robotics benchmarks and a real-world animal behavior modeling task.

cs.LG

Precision Polarization Tuning for Light Shift Mitigation in Trapped-Ion Qubits

Trapped-ion qubits are among the most promising candidates for quantum computing, quantum information processing, and quantum simulation. In general, trapped ions are considered to have sufficiently long coherence times, which are mainly characterized under laser-free conditions. However, in reality, essential laser fields for quantum manipulation introduce residual light shift, which seriously degrades the coherence due to power fluctuations. Here, we present a comprehensive study of AC Stark shifts in the hyperfine energy levels of the $^{171}\mathrm{Yb}^+$ ion, revealing an asymmetric light shift between two circular polarizations in the clock qubit and pronounced vector light shifts in the Zeeman qubits. By precisely tuning these polarizations, a remarkable enhancement in coherence time is observed, reaching over a hundredfold for the clock qubit and more than tenfold for the Zeeman qubits, when comparing conditions of maximum and minimum shifts. These findings advance the practical realization of scalable trapped-ion quantum processors, enabling deep quantum circuit execution and long duration adiabatic operations.

quant-ph