arXiv · 2609.33177
DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model
Abstract
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the changes between them are what an action policy needs to anticipate. We introduce DeltaWAM, a change-centric WAM that makes a compact delta token the unit of future prediction. Each token is a single vector encoding changes between consecutive dense DINO feature maps. DeltaWAM builds on DeltaWorld, a latent world model pretrained on large-scale videos, to autoregressively predict one delta token per future frame. A flow-matching action expert then conditions on the predicted transitions and current DINO features, which serve as spatial anchors, to generate action chunks. Trained for 256 GPU hours on two H100 GPUs, DeltaWAM has 0.725B parameters and achieves 92.8% average success on LIBERO. It also shows robust generalization under procedural perturbations on LIBERO-Pro. Inference takes 142.1 ms per action chunk with 3.86 GB peak memory.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tianyun Jiang, Wenrui Bao, Bingxin Xu, Yu Tian, Yuzhang Shang. 2026-09-27. DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model. https://arxiv.org/abs/2609.33177
Cite the original work for its findings. Save a collection to share your selection of sources.