SearcharxivSearch

arXiv subjects

Xuan Du

Publications and source records attributed to Xuan Du.

6 recordsLinked to original sources

CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning

Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it remains unclear how to extend this paradigm to continual learning, where a single policy must acquire new skills without losing previously learned behaviors. Existing real-world continual learning methods do not explicitly constrain prior behaviors, leading to severe catastrophic forgetting. We introduce Continual Interactive Distillation for Embodied Reinforcement Learning (CIDER), a continual reinforcement learning framework that freezes the accumulated historical policy as a teacher before learning each new task and interleaves task learning with distillation-based retention. We further introduce gradient routing to separate the gradients used for acquiring new tasks from those used for preserving prior behaviors. We evaluate our method with a single shared actor on six real-world household and industrial manipulation tasks. Interactive Distillation maintains high measured success on previously learned tasks across our six-task real-robot sequence while acquiring each new task in 10 to 20 minutes, whereas every baseline forgets at least one previous task. Additional ablations reveal the key design choices that govern the tradeoff between stability and plasticity in real-world continual reinforcement learning.

cs.RO

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, tackles this by inserting reasoning steps only during natural speech generation. This requires high-quality data where reasoning and speech are precisely aligned, and the length ratio are under controlled. We introduce a novel pipeline to generate such seamlessly interleaved audio data. To train our model, we combine interleaved SFT with refined data and reinforcement learning with two new rewards: a TA-Balance Reward to manage timing and thinking-answer ratio, and a Linguistic Quality Reward to refine expression. Experiments show our approach achieves 13% better performance on mathmatical and logic benchmarks while generating instant response like a spoken-language instruct model which outputs fast CoT response. Furthermore, our method generates more natural and fluent answers than prior methods.

cs.CL

ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training

We study how to improve large foundation vision-language-action (VLA) systems through human-in-the-loop reinforcement learning (RL) in real-world environments. A key challenge is learning reliable value functions from heterogeneous real-world experience, as value estimation provides the primary learning signal for VLA training. In practice, replay buffers contain trajectories collected from historical policies, online rollouts, demonstrations, and intermittent human interventions. Because replay buffers mix trajectories generated by different behaviors, the observed returns can be mismatched with the quality of the current policy. Prior VLA post-training methods often rely on progress-style value signals, which reflect the average quality of historical behaviors, leading to mismatched learning signals for the current policy. In this paper, we propose ALOE, an off-policy evaluation framework whose value function directly evaluates current-policy behavior for each iteration. Specifically, ALOE combines chunked temporal-difference bootstrapping and conservative value aggregation to perform stable current-policy evaluation, then uses these estimates for advantage-weighted policy improvement. This design improves credit assignment to critical action chunks under sparse rewards and supports stable policy improvement. We evaluate ALOE on four real-world manipulation tasks encompassing long-horizon and high-precision scenarios: smartphone packing, laundry folding, multi-object sorting, and phone assembly. Across all tasks, ALOE outperforms other VLA post-training methods, highlighting the benefit of off-policy value estimates for real-world VLA post-training. Videos are available at our project website https://rooshy-yang.github.io/aloe.

cs.RO

Beam Measurements of Full Stokes Parameters for the FAST L-band 19-beam Receiver

The Five-hundred-meter Aperture Spherical radio Telescope (FAST) has been fully operational since 11 January 2020. We present a comprehensive analysis of the beam structure for each of the 19 feed horns on FAST's L-band receiver across the Stokes I, Q, U, and V parameters. Using an on-the-fly mapping pattern, we conducted simultaneous sky mapping using all 19 beams directed towards polarization calibrators J1407+2827 and J0854+2006 from 2020 to 2022. Electromagnetic simulations were also performed to model the telescope's beam patterns in all Stokes parameters. Our findings reveal a symmetrical Gaussian pattern in the Stokes I parameter of the central beam without strong sidelobes, while the off-center beams exhibit significant asymmetrical shapes that can be fitted using a combination of log-normal and Gaussian distributions. The inner beams have higher relative beam efficiencies and smaller beam sizes compared to those of the outer beams. The sidelobes of the inner beams contribute approximately 2% of the total flux in the main lobe, increasing to 5% for outer beams, with a peak at 6.8%. In Stokes U, a distinct four-lobed cloverleaf beam squash structure is observed, with similar intensity levels in both inner and outer beams. In Stokes V, a two-lobed beam squint structure is observed in the central beam, along with a secondary eight-lobed structure. The highest squint peak in Stokes V is about 0.3% of the Stokes I in the outer beams. These results align closely with the simulations, providing valuable insights for the design of radio multi-beam observations.

astro-ph.IM

Gain and Polarization Properties of a Large Radio Telescope from Calculation and Measurement: The John A. Galt Telescope

Measurement of the brightness temperature of extended radio emission demands knowledge of the gain (or aperture efficiency) of the telescope and measurement of the polarized component of the emission requires correction for the conversion of unpolarized emission from sky and ground to apparently polarized signal. Radiation properties of the John A. Galt Telescope at the Dominion Radio Astrophysical Observatory were studied through analysis and measurement in order to provide absolute calibration of a survey of polarized emission from the entire northern sky from 1280 to 1750 MHz, and to understand the polarization performance of the telescope. Electromagnetic simulators CST and GRASP-10 were used to compute radiation patterns of the telescope in all Stokes parameters, and aperture efficiency. Aperture efficiency was also evaluated using geometrical optics and was measured using Cyg A. Measured aperture efficiency varied smoothly with frequency between values of 0.49 and 0.54; GRASP-10 yielded values 6.5% higher but with closely similar variation with frequency. Overall error across the frequency band is 3%, but values at any two frequencies are relatively correct to ~1%. Dominant influences on aperture efficiency are illumination taper of the feed radiation pattern and shadowing by the feed-support struts. A model of ground emission was developed based on measurements and on empirical data from remote sensing of the Earth from satellite-borne telescopes. This model was convolved with the computed antenna response to estimate conversion of ground emission into spurious polarized signal. The computed spurious signal is comparable to measured values, but is not accurate enough to be used to correct observations. A simpler model, in which the ground is considered as an unpolarized emitter with a brightness temperature of ~240 K, is shown to have useful accuracy when compared to measurements.

astro-ph.IM

Generalized Full-Vector Multi-Mode Matching Analysis of Whispering-Gallery Microcavities

We outline a full-vectorial three-dimensional multi-mode matching technique in a cylindrical coordinate system that addresses the mutual coupling among multiple modes copropagating in a perturbed whispering-gallery-mode microcavity. In addition to its superior accuracy in respect to our previously implemented single-mode matching technique, this current technique is suitable for modelling waveguide-to-cavity coupling where the influence of multi-mode coupling is non-negligible. Using this methodology, a robust scheme for hybrid integration of a microcavity onto a silicon-on-insulator platform is proposed.

physics.optics