arXiv · 2602.00627
FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization
Abstract
Benefiting from the significant advancements in text-to-image diffusion models, research in personalized image generation, particularly customized portrait generation, has also made great strides recently. However, existing methods either require time-consuming fine-tuning and lack generalizability or fail to achieve high fidelity in facial details. To address these issues, we propose FaceSnap, a novel method based on Stable Diffusion (SD) that requires only a single reference image and produces extremely consistent results in a single inference stage. This method is plug-and-play and can be easily extended to different SD models. Specifically, we design a new Facial Attribute Mixer that can extract comprehensive fused information from both low-level specific features and high-level abstract features, providing better guidance for image generation. We also introduce a Landmark Predictor that maintains reference identity across landmarks with different poses, providing diverse yet detailed spatial control conditions for image generation. Then we use an ID-preserving module to inject these into the UNet. Experimental results demonstrate that our approach performs remarkably in personalized and customized portrait generation, surpassing other state-of-the-art methods in this domain.
Explore related subjects
Keep this discovery
Benxiang Zhai, Yifang Xu, Guofeng Zhang, Yang Li, Sidan Du. 2026-01-31. FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization. https://arxiv.org/abs/2602.00627
Cite the original work for its findings. Save a collection to share your selection of sources.