SearcharxivSearch

arXiv subjects

Leiyu Wang

Publications and source records attributed to Leiyu Wang.

4 recordsLinked to original sources

Scaling by Diversified Experience for Vision-Language-Action Models

Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propose an Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts and a similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift. Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. Codes and Datasets is released on \href{https://sy-vla.github.io/}{project page}.

cs.CV

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform

Embodied AI in the real world requires both accurate hardware and robust vision-language-action (VLA) policies. We present OpenEAI-Platform, a fully open-source platform that integrates a low-cost 6+1 degree-of-freedom (dof) robotic arm (OpenEAI-Arm) and a reproducible VLA model (OpenEAI-VLA). OpenEAI-Arm provides open-source mechanical designs for low manufacturing cost and compliant control methods for higher accuracy. OpenEAI-VLA builds on Qwen3-VL-4B and uses a Diffusion Transformer action head, and is trained in two stages with only open-source robot and multimodal datasets. Across four real-world manipulation tasks, OpenEAI-Arm outperforms two commercial 6+1-dof arms under the same policy, and OpenEAI-VLA achieves success rates comparable to the large-scale pretrained pi0 baseline with only limited pretraining data. We will release the full hardware designs, drivers, models, and training/data pipelines to support reproducible research and scalable data collection. Our codes, layouts, and models will be released after the paper is accepted.

cs.RO

MO R-CNN: Multispectral Oriented R-CNN for Object Detection in Remote Sensing Image

Oriented object detection for multi-spectral imagery faces significant challenges due to differences both within and between modalities. Although existing methods have improved detection accuracy through complex network architectures, their high computational complexity and memory consumption severely restrict their performance. Motivated by the success of large kernel convolutions in remote sensing, we propose MO R-CNN, a lightweight framework for multi-spectral oriented detection featuring heterogeneous feature extraction network (HFEN), single modality supervision (SMS), and condition-based multimodal label fusion (CMLF). HFEN leverages inter-modal differences to adaptively align, merge, and enhance multi-modal features. SMS constrains multi-scale features and enables the model to learn from multiple modalities. CMLF fuses multimodal labels based on specific rules, providing the model with a more robust and consistent supervisory signal. Experiments on the DroneVehicle, VEDAI and OGSOD datasets prove the superiority of our method. The source code is available at:https://github.com/Iwill-github/MORCNN.

cs.CV

Generalized Beamspace Modulation Using Multiplexing: A Breakthrough in mmWave MIMO

Spatial multiplexing (SMX) multiple-input multiple-output (MIMO) over the best beamspace was considered as the best solution for millimeter wave (mmWave) communications regarding spectral efficiency (SE), referred as the best beamspace selection (BBS) solution. The equivalent MIMO water-filling (WF-MIMO) channel capacity was treated as an unsurpassed SE upper bound. Recently, researchers have proposed various schemes trying to approach the benchmark and the performance bound. But, are they the real limit of mmWave MIMO systems with reduced radio-frequency (RF) chains? In this paper, we challenge the benchmark and the corresponding bound by proposing a better transmission scheme that achieves higher SE, namely the Generalized Beamspace Modulation using Multiplexing (GBMM). Inspired by the concept of spatial modulation, besides the selected beamspace, the selection operation is used to carry information. We prove that GBMM is superior to BBS in terms of SE and can break through the well known `upper bound'. That is, GBMM renews the upper bound of the SE. We investigate SE-oriented precoder activation probability optimization, fully-digital precoder design, optimal power allocation and hybrid precoder design for GBMM. A gradient ascent algorithm is developed to find the optimal solution, which is applicable in all signal-to-noise-ratio (SNR) regimes. The best solution is derived in the high SNR regime. Additionally, we investigate the hybrid receiver design and deduce the minimum number of receive RF chains configured to gain from GBMM in achievable SE. We propose a coding approach to realize the optimized precoder activation. An extension to mmWave broadband communications is also discussed. Comparisons with the benchmark (i.e., WF-MIMO channel capacity) are made under different system configurations to show the superiority of GBMM.

eess.SP