arXiv · 2610.03473
UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
Abstract
We propose UniDynamics, a diffusion-based framework for future 4D dynamic scenes (RGB, depth, and optical flow) generation from a single event-RGB pair, without requiring long histories or control priors as in existing methods, while explicitly modeling future motion fields. The core idea is to leverage event streams to offer an alternative motion prior for single-RGB extrapolation, and to enforce geometric and motion constraints throughout generation via multimodal modeling. Specifically, we design an Event Latent Enhancement (ELE) module to align and enhance event latents into diffusion-injectable conditioning features, providing robust initial motion priors and reliable texture/structure cues. We further introduce a Perceptual Dynamics Space (PDS) embedded in the multi-scale U-Net, which decouples and adaptively interacts depth and flow while continuously feeding back constraints to appearance features, improving geometric-motion consistency for physically plausible and spatiotemporally coherent prediction. Experiments on VKitti2 and DSEC demonstrate state-of-the-art performance, producing high-quality, temporally coherent, and 4D-consistent future predictions, especially under challenging high-speed motion blur.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daikun Liu, Xin Zhan, Teng Wang, Xiaoping Wang, Changyin Sun. 2026-10-02. UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation. https://arxiv.org/abs/2610.03473
Cite the original work for its findings. Save a collection to share your selection of sources.