arXiv · 2604.12281
MAST: Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer
Abstract
Style transfer applies the appearance of a reference image to a content image while preserving its spatial structure. Recent diffusion-based methods achieve strong stylization but typically assume a single global style. We instead consider regional-multi style transfer, which assigns multiple references to user-specified regions of a content image. Extending them to this setting reveals two coupled shared-attention issues: ambiguous mass allocation among content and style partitions and degraded selectivity as more styles are jointly normalized, while style aggregation at the attention output further suppresses fine details. We propose MAST (Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer), a unified attention-control framework for frozen diffusion models. Logit-level Attention Mass Allocation enforces mask-derived partition masses, Sharpness-aware Temperature Scaling adaptively restores selectivity, and Discrepancy-aware Detail Injection recovers high-frequency content. MAST jointly processes all style--mask pairs in a single denoising pass without training, optimization, or post-hoc composition. Across two to five styles, MAST achieves the best average ArtFID, FID, and R-FID among all baselines, demonstrating regional style fidelity, content preservation, and scalability.
Explore related subjects
Keep this discovery
Dongkyung Kang, Jaeyeon Hwang, Junseo Park, Minji Kang, Yeryeong Lee, Beomseok Ko, Hanyoung Roh, Jeongmin Shin, Hyeryung Jang. 2026-04-14. MAST: Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer. https://arxiv.org/abs/2604.12281
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.