SearcharxivSearch

arXiv subjects

Yingkai Cai

Publications and source records attributed to Yingkai Cai.

3 recordsLinked to original sources

Cross-View Action Consistency for Camera-Robust Vision-Language-Action Policies

Vision-language-action (VLA) policies fine-tuned from a fixed scene camera can fail when the camera is moved, even when the task, objects, language, and robot state are unchanged. We study scene-camera viewpoint robustness using only a scene RGB image, language, and proprioception, without camera labels, extrinsics, depth, or point-cloud inputs. The wrist stream is masked throughout to prevent an unperturbed visual shortcut from confounding attribution to scene-camera variation. For flow-based VLAs, we propose to regularize the action-flow velocity field, the quantity directly integrated to generate continuous action chunks. We construct action-equivalent view pairs by resetting original LIBERO demonstrations to the same MuJoCo state and rendering nominal and perturbed scene-camera views. Both views are supervised by flow matching, while a cross-view loss encourages their predicted action-flow velocities to agree at the same sampled flow coordinates. On the LIBERO-Plus camera-perturbation track, our method reaches 87.2$\pm$0.4% (4,797 rollouts per seed across 3 training seeds), +7.4pp over flow-matching-only training on the same paired data (79.8$\pm$0.8%, also 3 seeds) and +12.5pp over naive mixed-camera SFT, while maintaining nominal-camera ID performance (95.0$\pm$0.8%; same-data FM-only: 95.0$\pm$4.3%). A shuffled-pair control collapses to 25.8%, showing that the gain depends on action-equivalent pairing. On a real robot, we evaluate three tabletop tasks with 10 rollouts per task and camera placement; held-out-camera success improves from 53.3% to 74.4% under the same single-scene-RGB inference interface.

cs.RO

Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed? The WAM denoising target mixes view-covariant coordinates, namely the predicted future scene, with view-invariant coordinates, namely the action chunk, future proprioception, and value. We show that consistency applied to the covariant block is provably harmful, shrinking legitimate view-specific content to a fraction $1/(1+4\lambda)$ of its true value, and we verify this shrinkage law in controlled experiments. Selective cross-view consistency (SCVC) therefore constrains only the invariant block, requires no camera labels, extrinsics, depth, or view synthesis at training or test time, and leaves the deployment interface unchanged. We introduce a carve-and-hold-out evaluation protocol on the LIBERO-Plus camera track that separates a distribution-matched ceiling from genuine interpolation and extrapolation to held-out viewpoints, with a matched pair-trained control isolating the effect of the consistency term from pair exposure. On held-out orbital viewpoints beyond the training envelope, SCVC improves closed-loop success over the matched control by 12.2 points (95% CI [7.4, 17.0]; +15.5, CI [11.7, 19.4], under an independent second seed) -- an effect two further camera axes replicate -- while interpolation within the envelope shows no gain in either seed (-1.2 and -4.3 points) and in-distribution competence is preserved (-0.6, -0.2). We also report a cross-backbone audit showing that published camera-robustness numbers are confounded by wrist-camera pose stability.

cs.RO

Atmospheric Density Model Optimization and Spacecraft Orbit Prediction Improvements Based on Q-Sat Orbit Data

Atmospheric drag calculation error greatly reduce the low-earth orbit spacecraft trajectory prediction fidelity. To solve the issue, the "correction - prediction" strategy is usually employed. In the method, one parameter is fixed and other parameters are revised by inverting spacecraft orbit data. However, based on a single spacecraft data, the strategy usually performs poorly as parameters in drag force calculation are coupled with each other, which result in convoluted errors. A gravity field recovery and atmospheric density detection satellite, Q-Sat, developed by xxxxx Lab at xxx University, is launched on August 6th, 2020. The satellite is designed to be spherical for a constant drag coefficient regardless of its attitude. An orbit prediction method for low-earth orbit spacecraft with employment of Q-Sat data is proposed in present paper for decoupling atmospheric density and drag coefficient identification. For the first step, by using a dynamic approach-based inversion, several empirical atmospheric density models are revised based on Q-Sat orbit data. Depending on the performs, one of the revised atmospheric density model would be selected for the next step in which the same inversion is employed for drag coefficient identification for a low-earth orbit operating spacecraft whose orbit needs to be predicted. Finally, orbit prediction is conducted by extrapolation with the dynamic parameters in the previous steps. Tests are carried out with the proposed method by using a GOCE satellite 15-day continuous orbit data. Compared with legacy "correction - prediction" method in which only GOCE data is employed, the accuracy of the 24-hour orbit prediction is improved by about 171m the highest for the proposed method. 14-day averaged 24-hour prediction precision is elevated by approximately 70m.

physics.geo-ph