arXiv · 2609.19796
LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation
Abstract
During manipulation, robot and scene motion can move previously observed regions outside the camera's field of view. Geometry-aware RGB features encode visible structure, while control under partial observability requires scene memory that integrates observation history and grounds inferred content in current evidence. We introduce \lifd{} (Look, Imagine, Focus, and Do), a framework for persistent, 3D-aware scene memory. LIFD learns scene tokens through multi-view agreement, then completes them from a single RGB view and recurrent memory using rectified flow. Anchor-Guided Cross-Attention anchors generation to current geometry-aware features, and compact slot features condition a visuomotor policy. Multi-view and geometric supervision are used during representation learning; deployment requires one RGB camera, proprioception, and a task instruction. LIFD (Staged) reaches 91.6\% average success on LIBERO and 79.8\% on MetaWorld, improving LIBERO average success by 11.1 percentage points over Joint training. After policy-head adaptation with ten demonstrations per family, LIFD achieves 56.0\% mean success across four UR5e task families, compared with 40.5\% for OpenVLA-7B.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wenbo Li, Yiteng Chen, Wenhao Li, Qingyao Wu. 2026-09-18. LIFD: Anchored Diffusion for 3D-Aware Scene Memory in Robotic Manipulation. https://arxiv.org/abs/2609.19796
Cite the original work for its findings. Save a collection to share your selection of sources.