arXiv · 2603.23924
DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis
Abstract
Text-to-image models often struggle to synthesize correct occlusion relationships among multiple objects, especially in densely overlapping regions. Many training-free layout-guided methods enforce 2D spatial constraints but do not explicitly resolve depth-dependent attention competition, which can cause concept mixing and implausible occlusion. To address this problem, we propose DepthArb, a training-free framework that formulates occlusion generation as attention arbitration within a unified denoising trajectory. DepthArb employs two core occlusion-control mechanisms: Attention Arbitration Modulation suppresses background-object attention within foreground support according to relative depth, while Spatial Compactness Control limits attention dispersion to preserve object coherence. Because interference varies during generation, Occlusion Conflict Estimation constructs a shared spatial conflict field to adaptively weight both objectives. Through a unified spatial-text attention interface, DepthArb operates on U-Net cross-attention and the image-to-text component of MMDiT joint attention without model retraining. We further introduce OcclBench, a benchmark with continuous relative-depth specifications and occlusion-specific evaluation metrics. Experiments on OcclBench and public benchmarks show that DepthArb improves several layout and occlusion metrics over the evaluated baselines while maintaining competitive text-image alignment.
Explore related subjects
Keep this discovery
Hongjin Niu, Jiahao Wang, Xirui Hu, Weizhan Zhang, Lan Ma, Yuan Gao, Feng Lei. 2026-03-25. DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis. https://arxiv.org/abs/2603.23924
Cite the original work for its findings. Save a collection to share your selection of sources.