arXiv · 2602.14021
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
Abstract
Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motion: static multi-view reconstruction systems assume a rigid world, whereas dynamic tracking frameworks rely on explicit ego-motion estimation or separate object motion models. In this work, we propose Flow4R, a unified framework that treats relative scene flow as the central representation linking 3D structure, camera ego-motion, and dynamic object motion. Given a two-view input, Flow4R employs a shared Vision Transformer to predict a compact, pixel-aligned property set comprising 3D point positions, scene flow, pose weights, and confidence maps. This flow-centric formulation allows local geometry and bidirectional motion to be jointly inferred in a single feedforward pass, eliminating the need for explicit pose regression heads or complex bundle adjustment. By training jointly on static and dynamic datasets, Flow4R achieves state-of-the-art performance on 4D reconstruction and tracking benchmarks, demonstrating the power of the flow-centric formulation for spatiotemporal scene understanding.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shenhan Qian, Ganlin Zhang, Shangzhe Wu, Daniel Cremers. 2026-02-15. Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow. https://arxiv.org/abs/2602.14021
Cite the original work for its findings. Save a collection to share your selection of sources.