arXiv · 2509.14685
DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images
Abstract
Automatic colorization of line drawings has been widely studied to reduce the labor cost of hand-drawn anime production. Deep learning approaches, including image/video generation and feature-based correspondence, have improved accuracy but struggle with occlusions, pose variations, and viewpoint changes. To address these challenges, we propose DACoN, a framework that leverages foundation models to capture part-level semantics, even in line drawings. Our method fuses low-resolution semantic features from foundation models with high-resolution spatial features from CNNs for fine-grained yet robust feature extraction. In contrast to previous methods that rely on the Multiplex Transformer and support only one or two reference images, DACoN removes this constraint, allowing any number of references. Quantitative and qualitative evaluations demonstrate the benefits of using multiple reference images, achieving superior colorization performance. Our code and model are available at https://github.com/kzmngt/DACoN.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kazuma Nagata, Naoshi Kaneko. 2025-09-18. DACoN: DINO for Anime Paint Bucket Colorization with Any Number of Reference Images. https://arxiv.org/abs/2509.14685
Cite the original work for its findings. Save a collection to share your selection of sources.