arXiv · 2406.04814
Lifelong Learning of Video Diffusion Models From a Single Video Stream
Abstract
This work demonstrates that training autoregressive video diffusion models from a single video stream$\unicode{x2013}$resembling the experience of embodied agents$\unicode{x2013}$is not only possible, but can also be as effective as standard offline training given the same number of gradient steps. Our work further reveals that this main result can be achieved using experience replay methods that only retain a subset of the preceding video stream. To support training and evaluation in this setting, we introduce four new datasets for streaming lifelong generative video modeling: Lifelong Bouncing Balls, Lifelong 3D Maze, Lifelong Drive, and Lifelong PLAICraft, each consisting of one million consecutive frames from environments of increasing complexity.
Explore related subjects
Keep this discovery
Jason Yoo, Yingchen He, Saeid Naderiparizi, Dylan Green, Gido M. van de Ven, Geoff Pleiss, Frank Wood. 2024-06-07. Lifelong Learning of Video Diffusion Models From a Single Video Stream. https://arxiv.org/abs/2406.04814
Cite the original work for its findings. Save a collection to share your selection of sources.