arXiv · 2605.20237
AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation
Abstract
We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editing conditions. Instead of relying on large-scale vision-language models or per-subject fine-tuning, our method injects fine-grained visual features from a single reference image into the diffusion process. Based on CLIP emergent local spatialization, we develop semantic-selective local attention. To further disentangle character appearance from spatial layout, we incorporate pose-aware conditioning during adapter training. The resulting pretrained adapter remains compact, modular, and fully compatible with Stable Diffusion community workflows, while requiring no additional fine-tuning at deployment time. Furthermore, we present a high-quality anime character dataset based on curated and restructured Danbooru prompts, and evaluate our method across several practical character editing scenarios. Our code, model weights, and dataset will be publicly released upon acceptance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yixuan Han. 2026-05-17. AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation. https://arxiv.org/abs/2605.20237
Cite the original work for its findings. Save a collection to share your selection of sources.