arXiv · 2609.22684
StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies
Abstract
Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with external memory banks but require explicit storage and retrieval. To address these limitations, we propose StateMem, a single-state residual memory framework for VLA policies that uses prediction error to update a persistent memory token through low-rank residuals and to adaptively route cached prefixes. A training-free controller adjusts the routing threshold online, while fast correction compensates for stale prefix features during cache reuse. We evaluate StateMem on LIBERO, RoboMemArena, and real-world manipulation tasks. On LIBERO, StateMem achieves an average success rate of 97.6% and reduces the average VLM prefix refresh rate by 20.25% relative to full refresh. In the Occlusion category of RoboMemArena, StateMem achieves the best performance among single-VLA methods, reaching 21.8% Task Success Rate (TSR) and 44.3% Cumulative Success Rate (CSR). Across six real-world manipulation tasks, it achieves +21% in average success rate.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenzhuo Li, Qiongfeng Shi, Yi Zhou. 2026-09-19. StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies. https://arxiv.org/abs/2609.22684
Cite the original work for its findings. Save a collection to share your selection of sources.