arXiv · 2512.13677
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
Abstract
In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on fragmented, task-specific architectures or complex fusion mechanisms, JoVA employs native joint representation learning for direct video, audio, and text interaction in a dual-branch architecture. This design eliminates redundant alignment modules and effectively unifies diverse multimodal tasks within a single model. Furthermore, we utilize channel-wise conditioning for flexible image and video reference to avoid massive token expansion, alongside a mouth-area loss to enhance lip alignment. To fully empower and systematically evaluate this framework, we construct a comprehensive training corpus encompassing video-audio generation and editing datasets, and introduce unified benchmarks tailored for these multimodal tasks. Extensive experiments demonstrate that JoVA achieves state-of-the-art performance across benchmarks, establishing it as an extensible framework for versatile content creation. Project page: https://visual-ai.github.io/jova
Explore related subjects
Keep this discovery
Xiaohu Huang, Haoyang He, Hao Zhou, Qiangpeng Yang, Shilei Wen, Kai Han. 2025-12-15. JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing. https://arxiv.org/abs/2512.13677
Cite the original work for its findings. Save a collection to share your selection of sources.