arXiv · 2604.14556
Controllable Video Object Insertion via Multi-View Priors
Abstract
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text or a single reference image. Consequently, object appearance is underconstrained under viewpoint changes, often leading to identity drift, incorrect foreground-background layering, boundary artifacts, and temporal flickering. In this paper, we propose a video object insertion framework that incorporates multi-view object priors to address these limitations. The framework lifts a 2D reference image into a multi-view representation and uses view-consistent conditioning to provide stable identity guidance and view-adaptive appearance cues. A quality-aware weighting mechanism reduces the influence of noisy or imperfect reconstructed views. We further introduce an Integration-Aware Consistency Module that promotes plausible occlusion, clean boundaries, and temporal continuity. Experiments demonstrate that the proposed framework improves visual quality, controllability, identity consistency, and foreground-background integration for video object insertion compared to the baseline methods. Project page: https://polarisxq.github.io/MOVI/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Qi Xia, Peishan Cong, Yichen Yao, Ziyi Wang, Yaoqin Ye, Yuexin Ma. 2026-04-16. Controllable Video Object Insertion via Multi-View Priors. https://arxiv.org/abs/2604.14556
Cite the original work for its findings. Save a collection to share your selection of sources.