arXiv · 2610.02898
MixVLA: Adaptive Mixing of Non-Invariant Information for Generalizable Vision-Language-Action Models
Abstract
Vision-Language-Action (VLA) models have achieved remarkable advances in robotic manipulation, yet their zero-shot generalization under out-of-distribution (OOD) conditions remains limited. These models often entangle task-relevant invariant structure with environment-specific non-invariant factors, causing policies to rely on spurious appearance cues during action prediction. In this work, we propose \textbf{MixVLA}, a model-agnostic training framework that improves the generalization of VLA models without requiring additional OOD data or architectural modifications. The key component of MixVLA is \textbf{Adaptive Mixing of Non-Invariant Information (AMI)}. AMI stochastically mixes non-invariant representations to regularize distribution-specific variability while preserving complementary predictive cues. The mixed non-invariant features are then fused with invariant representations for final action prediction, resulting in improved robustness without sacrificing policy expressiveness. Extensive experiments across challenging manipulation settings, including LIBERO, LIBERO-Plus, the RoboTwin perturbation suite, and real-world tasks, demonstrate that MixVLA improves overall zero-shot robustness while retaining strong in-domain performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pingrui Zhang, Yu Zhang, Pengyuan Wu, Bin Wang, Haoming Song, Xianqiang Gao, ZhaxiZhuoma, Zhigang Wang, Dong Wang, Bin Zhao, Xuelong Li. 2026-10-02. MixVLA: Adaptive Mixing of Non-Invariant Information for Generalizable Vision-Language-Action Models. https://arxiv.org/abs/2610.02898
Cite the original work for its findings. Save a collection to share your selection of sources.