SearcharxivSearch

arXiv subjects

Guanglei Ding

Publications and source records attributed to Guanglei Ding.

2 recordsLinked to original sources

RhinoVLA Technical Report

Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify VLM visual and context tokens as a major source of deployment latency: for GEMM-dominated projection operators, computation grows linearly with the number of input tokens when model dimensions are fixed. Motivated by this observation, we propose RhinoVLA, a deployment-oriented VLA model co-designed with the Huixi R1 edge SoC. RhinoVLA adopts a token-efficient Qwen3-VL backbone and a continuous Action Expert, reducing the VLM-side token and computation burden while preserving pretrained multimodal capability. To support cross-robot learning, RhinoVLA further introduces a unified interface that combines View Registry, 72D physical state-action slot space, and robotinstance LoRA, allowing heterogeneous robot observations and action schemas to be aligned under a shared policy. On the deployment side, RhinoVLA is optimized through hardware-aware compilation, mixed-precision execution, and parallel visual encoding. Experiments show that RhinoVLA achieves downstream performance comparable to {\pi}0.5 at a similar parameter scale, while reaching 11.69 Hz end-to-end inference on Huixi R1, meeting the 10 Hz real-time closedloop control target. The project will be open-sourced at https://github.com/HuixiAI/RhinoVLA.

cs.RO

0.71-{\AA} resolution electron tomography enabled by deep learning aided information recovery

Electron tomography, as an important 3D imaging method, offers a powerful method to probe the 3D structure of materials from the nano- to the atomic-scale. However, as a grant challenge, radiation intolerance of the nanoscale samples and the missing-wedge-induced information loss and artifacts greatly hindered us from obtaining 3D atomic structures with high fidelity. Here, for the first time, by combining generative adversarial models with state-of-the-art network architectures, we demonstrate the resolution of electron tomography can be improved to 0.71 angstrom which is the highest three-dimensional imaging resolution that has been reported thus far. We also show it is possible to recover the lost information and remove artifacts in the reconstructed tomograms by only acquiring data from -50 to +50 degrees (44% reduction of dosage compared to -90 to +90 degrees full tilt series). In contrast to conventional methods, the deep learning model shows outstanding performance for both macroscopic objects and atomic features solving the long-standing dosage and missing-wedge problems in electron tomography. Our work provides important guidance for the application of machine learning methods to tomographic imaging and sheds light on its applications in other 3D imaging techniques.

cond-mat.mtrl-sci