arXiv · 2609.32389
RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion
Abstract
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales memory linearly with reference count $N$ and generated content often departs from the given references rather than reproducing them. We propose \textbf{RefCompose}, a pixel space compositional conditioning framework that decouples \emph{where} things go from \emph{what} they look like, via a single fixed resolution reference canvas that keeps conditioning size constant regardless of reference count. Spatial layout is induced at inference time from a frozen diffusion transformer and extracted via Grounding DINO, requiring no LLM or dedicated layout model, while dual stream LoRA adapters inject a layout derived depth map and the encoded canvas through separate low rank streams, disentangling geometric scaffolding from localized appearance. On the Dense Layout protocol, RefCompose consistently outperforms layout based and state of the art multi reference baselines on color, texture, shape, spatial accuracy, and identity/content preservation at higher reference counts, all with constant inference memory, making it a practical building block for multi subject cinematic composition at production scale.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sai Sri Teja Kuppa, Parth Shinde, Priyadharsan Balaji S, Jinka Harshavardhan, Sriprabha Ramanarayanan. 2026-09-26. RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion. https://arxiv.org/abs/2609.32389
Cite the original work for its findings. Save a collection to share your selection of sources.