arXiv · 2506.12520
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
Abstract
We propose VINO, the first zero-shot, training-free video editing method conditioned on both image and text. Our approach introduces $\rho$-start sampling and dilated dual masking to construct structured noise maps that enable coherent and accurate edits. To further enhance visual fidelity, we present zero image guidance, a controllable negative prompt strategy. Extensive experiments demonstrate that VINO faithfully incorporates the reference image into video edits, achieving strong performance compared to state-of-the-art baselines, all without any test-time or instance-specific training.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Saemee Choi, Sohyun Jeong, Hyojin Jang, Jaegul Choo, Jinhee Kim. 2025-06-14. Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts. https://arxiv.org/abs/2506.12520
Cite the original work for its findings. Save a collection to share your selection of sources.