SearcharxivSearch

arXiv subjects

Yong Pan

Publications and source records attributed to Yong Pan.

8 recordsLinked to original sources

One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.

cs.RO

Rethink Before You Execute: Adaptive Execution for World Action Models

World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed prefix before replanning. We argue that this fixed execution horizon is poorly matched to execution dynamics: the chunk reliability varies across task stages, so when to replan depends on the result of accumulated execution, not on the step counts. We propose TempoWAM (Timing Execution by Monitoring Progress Online), a lightweight plug-and-play execution scheme for WAMs. A Recurrent Progress Monitor first estimates task progress from the current observation, task instruction, remaining actions, and execution history; and an Adaptive Execution Protocol then evaluates whether the chunk is advancing the task to decide if replanning is needed. To bridge the training-deployment gap, the protocol is calibrated by a task-dependent calibration factor with online adaptation. Experiments on LIBERO, RoboTwin, and real-world tasks show that TempoWAM consistently improves the efficiency-success trade-off of WAM execution. On real robots, it reduces WAM inferences by 26.9% on easy tasks while maintaining success, and improves success by 13.3 points on difficult tasks.

cs.RO

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently limit representational capacity, our approach predicts multi-frame future occupancy in an end-to-end manner directly from raw image features. Inspired by the success of attention-based transformer architectures in foundational vision and language models such as GPT and VGGT, we employ a sparse occupancy representation that bypasses the intermediate bird's eye view (BEV) projection and its explicit geometric priors. This design allows the transformer to capture spatiotemporal dependencies more effectively. By avoiding both the finite-capacity constraint of discrete tokenization and the structural limitations of BEV representations, our method achieves state-of-the-art performance on the nuScenes benchmark for 1-3 second occupancy forecasting, outperforming existing approaches by a significant margin. Furthermore, it demonstrates robust scene dynamics understanding, consistently delivering high accuracy under arbitrary future trajectory conditioning.

cs.CV

QuadricFormer: Scene as Superquadrics for 3D Semantic Occupancy Prediction

3D occupancy prediction is crucial for robust autonomous driving systems as it enables comprehensive perception of environmental structures and semantics. Most existing methods employ dense voxel-based scene representations, ignoring the sparsity of driving scenes and resulting in inefficiency. Recent works explore object-centric representations based on sparse Gaussians, but their ellipsoidal shape prior limits the modeling of diverse structures. In real-world driving scenes, objects exhibit rich geometries (e.g., cuboids, cylinders, and irregular shapes), necessitating excessive ellipsoidal Gaussians densely packed for accurate modeling, which leads to inefficient representations. To address this, we propose to use geometrically expressive superquadrics as scene primitives, enabling efficient representation of complex structures with fewer primitives through their inherent shape diversity. We develop a probabilistic superquadric mixture model, which interprets each superquadric as an occupancy probability distribution with a corresponding geometry prior, and calculates semantics through probabilistic mixture. Building on this, we present QuadricFormer, a superquadric-based model for efficient 3D occupancy prediction, and introduce a pruning-and-splitting module to further enhance modeling efficiency by concentrating superquadrics in occupied regions. Extensive experiments on the nuScenes dataset demonstrate that QuadricFormer achieves state-of-the-art performance while maintaining superior efficiency.

cs.CV

GaussianAD: Gaussian-Centric End-to-End Autonomous Driving

Vision-based autonomous driving shows great potential due to its satisfactory performance and low costs. Most existing methods adopt dense representations (e.g., bird's eye view) or sparse representations (e.g., instance boxes) for decision-making, which suffer from the trade-off between comprehensiveness and efficiency. This paper explores a Gaussian-centric end-to-end autonomous driving (GaussianAD) framework and exploits 3D semantic Gaussians to extensively yet sparsely describe the scene. We initialize the scene with uniform 3D Gaussians and use surrounding-view images to progressively refine them to obtain the 3D Gaussian scene representation. We then use sparse convolutions to efficiently perform 3D perception (e.g., 3D detection, semantic map construction). We predict 3D flows for the Gaussians with dynamic semantics and plan the ego trajectory accordingly with an objective of future scene forecasting. Our GaussianAD can be trained in an end-to-end manner with optional perception labels when available. Extensive experiments on the widely used nuScenes dataset verify the effectiveness of our end-to-end GaussianAD on various tasks including motion planning, 3D occupancy prediction, and 4D occupancy forecasting. Code: https://github.com/wzzheng/GaussianAD.

cs.CV

Film Thickness Gauge Based on Interferometric Principle of Y-shaped Optical Fiber

In this paper, a thin film thickness gauge based on the interferometric principle of Y-shaped optical fiber is proposed to achieve accurate measurement of film thickness. In this paper, the optical fiber, the interferometric principle and the film thickness calculation principle are introduced, and the interferometric thickness measurement system based on Y-shaped optical fiber is constructed. The system uses the special structure of Y-shaped optical fiber to transmit the optical signal generated by the light source to the surface of the thin film, and obtains coherent optical signals of different wavelengths through reflection and interference. The spectrometer is used to receive and interpret these interference signals, and the thickness of the film is calculated according to the wavelength difference of the peak positions of the adjacent stages, combined with the refractive index of the film. In the specific design, the paper elaborates on the design of each part of the instrument, including the selection and parameter setting of the light source, Y-fiber and spectrometer. Among them, the Y-shaped optical fiber, as the core component of the instrument, has the function of transmitting optical signals and detecting optical signals on the surface of thin films. At the same time, the paper also introduces the housing packaging and internal assembly process of the instrument to ensure the portability and stability of the instrument. The results show that the thickness gauge has high measurement accuracy and stability, which can meet the needs of practical applications.

physics.optics

Multifunctional Portable Optical Measuring Instrument Based on Y-Fiber Optics

Based on grating diffraction principle, optical fiber transmission principle and optical interference principle, a multi-functional portable optical measuring instrument is constructed in this paper. The optical measurement visualization spectrometer based on CCD photoelectric image sensor is designed and assembled. The "Y" optical signal transmission fiber optical path suitable for multi-function measurement is improved and designed. The multi-function optical measurement system is built by combining with remote controlled multi-color LED lights. The spectral analysis, solution concentration monitoring and film thickness measurement are realized. The experimental results show that the observable wavelength range of the spectrometer is about 340-1050nm and the resolution is 1nm. The solution concentration can be obtained by measuring absorbance with optical fiber spectrometer. The film thickness measuring instrument can accurately measure the thickness of the micron film, and the measurement accuracy can reach 1.25 {\mu}m. It is proved that the instrument integrates multiple functions, has high measurement accuracy and wide range, and realizes non-contact measurement.

physics.optics

New measurement method for weak magnetic fields using magnetically induced deformation of chemical bonds in Co-CsPbBr3 quantum dots

The research on weak magnetic field detection is of great significance in advancing the development of bioscience, aerospace, chip manufacturing and other fields. However, the weak magnetic detecting still face some problems, including the large size of the detectors and the limited detection scale. To contribute to the detection of weak magnetic fields, the Co-CsPbBr3 colloidal quantum dots (QDs) composite magnetic material was synthesised on the basis of the theory of room temperature ferromagnetism, molecular polarisation and vibration level of chemical bond. The synthesis involved the mixing of Co2+ into CsPbBr3, an all-inorganic perovskite with activated ions. Subsequently, a weak magnetic field measurement system was devised, comprising working medium samples and a vibration level detection optical path. Following the acquisition, comparison, processing and analysis of multiple data sets, a Stokes displacement function model was established under different magnetic field sizes and the weak magnetic field intensity range of Pitsla (pT) was measured. The Pitsla weak magnetic field measurement system proposed in this paper provides a reference for the development of non-contact weak magnetic measurement methods and for the advancement of intelligent and low-dimensional weak signal measurement applications.

physics.optics