arXiv · 2609.28342
Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement
Abstract
Removing an object from a real image requires more than synthesizing plausible content within a mask: the method must suppress residual object features, preserve the unedited scene, and generate replacement content that is consistent with the surrounding background. This paper approaches object removal from a stage-based perspective and proposes a zero-shot framework for constrained latent inpainting with a frozen pretrained Stable Diffusion model, requiring no task-specific training or model fine-tuning. The method integrates SAM-based mask construction, BLIP image-caption conditioning, DDIM inversion, background-weighted masked null-text optimization, decoder self-attention masking, hard outside-mask latent anchoring, and localized renoise--denoise refinement into a unified pipeline. The method is evaluated through qualitative examples, quantitative local-consistency metrics, and ablation studies. The results demonstrate effective object removal and context-consistent replacement content. The ablations indicate that background-weighted masked NTI is particularly beneficial for structurally complex backgrounds, whereas the no-NTI variant is sufficient in other evaluated examples. Repeated refinement further reduces object remnants and boundary artifacts remaining after the primary editing pass.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arman Taghizadeh, Ulf Krumnack, Kai-Uwe Kühnberger. 2026-09-23. Zero-Shot Object Removal via Attention Masking, Latent Anchoring, and Refinement. https://arxiv.org/abs/2609.28342
Cite the original work for its findings. Save a collection to share your selection of sources.