arXiv · 2609.24337
LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba
Abstract
While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that synergizes the strengths of CNNs, Mamba, and Transformers. The core of our design is the Adaptive Bidirectional Mamba (ABM), which eliminates the directional bias of unidirectional SSMs through bidirectional selective scanning with learnable alpha blending, enabling content-adaptive directional fusion without the overhead of exhaustive multi-path routing. To further accelerate inference, we propose a deployment-friendly Reparameterized ConvSE (RepConvSE) module that leverages structural reparameterization to reduce latency and memory access overhead. Extensive experiments demonstrate that LiAuto-MindViT achieves state-of-the-art performance on image classification, object detection, and semantic segmentation while enabling efficient inference through reparameterization.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lifu Mu, Shuai Chen, Wen Zheng, Haoyi Sun, Xueyang Fu, Pengfei Yu, Ning Mao, Tao Wei, Zhou Pan. 2026-09-21. LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba. https://arxiv.org/abs/2609.24337
Cite the original work for its findings. Save a collection to share your selection of sources.