arXiv · 2412.03758
ARCON: Advancing Auto-Regressive Continuation for Driving Videos
Abstract
Recent advancements in auto-regressive large language models (LLMs) have led to their application in video generation. This paper explores the use of Large Vision Models (LVMs) for video continuation, a task essential for building world models and predicting future frames. We introduce ARCON, a scheme that alternates between generating semantic and RGB tokens, allowing the LVM to explicitly learn high-level structural video information. We find high consistency in the RGB images and semantic maps generated without special design. Moreover, we employ an optical flow-based texture stitching method to enhance visual quality. Experiments in autonomous driving scenarios show that our model can consistently generate long videos.
Explore related subjects
Keep this discovery
Ruibo Ming, Jingwei Wu, Zhewei Huang, Zhuoxuan Ju, Jianming HU, Lihui Peng, Shuchang Zhou. 2024-12-04. ARCON: Advancing Auto-Regressive Continuation for Driving Videos. https://arxiv.org/abs/2412.03758
Cite the original work for its findings. Save a collection to share your selection of sources.