arXiv · 2609.29204
AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution
Abstract
Vision-language-action (VLA) models offer strong local control and instruction following but often struggle with long-horizon tasks requiring persistent memory and planning. Task harnesses provide persistent context for agent reasoning by retaining task history and tracking progress across execution stages. To bring these complementary capabilities together, we introduce AdaHVLA, an adaptive harness that refines code-based coordination policies through robot experience to better align agent reasoning and memory with VLA execution. Its decoupled multiagent adaptation process separates evidence analysis, harness revision, and behavioral assessment into distinct working contexts, using testable coordination hypotheses to guide revisions and subsequent rollouts to assess their predicted effects. A stateful revision graph links execution evidence, hypotheses, revisions, and observed effects, preserving alternative harnesses and adaptation memory to guide refinement across repeated attempts and continued adaptation across tasks and environments. In simulation, AdaHVLA raises mean test success on NaVILA-LH from 22.5\% to as high as 57.5\% and improves manipulation test success across three VLA backbones by up to 30.8 percentage points over the initial harness. Real-world deployment further illustrates how the adapted policies support stable execution across task stages.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junyi Tang, Jie Peng, Zezhen Ding, Yuan Shen, Tianlong Chen. 2026-09-24. AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution. https://arxiv.org/abs/2609.29204
Cite the original work for its findings. Save a collection to share your selection of sources.