arXiv · 2606.22042
IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance
Abstract
Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in multi-object scenarios, remains challenging due to attention leakage, identity drift, and unstable temporal dynamics. In this work, we propose IDAGEdit, a training-free framework for fine-grained multi-object video editing with strong temporal consistency. The framework adopts Layout-guided Attention Modulation to facilitate coherent multi-object editing, while Instance-level Masks are introduced to preserve individual object identity and enforce localized attention within each object region, thereby enabling fine-grained, object-level editing. Extensive qualitative and quantitative evaluations demonstrate that our method improves temporal stability and multi-object controllability over state-of-the-art video editing approaches.
Explore related subjects
Keep this discovery
Yuan-Zhih Lin, Huu-Thang Nguyen, Huu-Phu Do, Hong-Han Shuai, Ching-Chun Huang. 2026-06-20. IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance. https://arxiv.org/abs/2606.22042
Cite the original work for its findings. Save a collection to share your selection of sources.