arXiv · 2604.06757
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
Abstract
Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions, spatial layouts, and editing instructions, can be unified into a single visual representation. We present FlowInOne, a framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean image-in, image-out pipeline governed by a single flow matching model. This vision-centric formulation naturally eliminates cross-modal alignment bottlenecks, noise scheduling, and task-specific architectural branches, unifying text-to-image generation, layout-guided editing, and visual instruction following under one coherent paradigm. To support this, we introduce VisPrompt-5M, a large-scale dataset of 5 million visual prompt pairs spanning diverse tasks including physics-aware force dynamics and trajectory prediction, alongside VP-Bench, a rigorously curated benchmark assessing instruction faithfulness, spatial precision, visual realism, and content consistency. Extensive experiments demonstrate that FlowInOne achieves state-of-the-art performance among open-source models across all unified generation tasks while remaining competitive with leading commercial systems, thereby establishing a new foundation for fully vision-centric generative modeling, in which perception and creation coexist within a unified continuous visual space. Our code and models are released on https://csu-jpg.github.io/FlowInOne.github.io/
Explore related subjects
Keep this discovery
Junchao Yi, Rui Zhao, Jiahao Tang, Weixian Lei, Linjie Li, Qisheng Su, Zhengyuan Yang, Lijuan Wang, Xiaofeng Zhu, Alex Jinpeng Wang. 2026-04-08. FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching. https://arxiv.org/abs/2604.06757
Cite the original work for its findings. Save a collection to share your selection of sources.