arXiv · 2410.10719
4-LEGS: 4D Language Embedded Gaussian Splatting
Abstract
The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views. Recently, several techniques have been proposed for connecting these low-level representations with the high-level semantics understanding embodied within the scene. These methods elevate the rich semantic understanding from 2D imagery to 3D representations, distilling high-dimensional spatial features onto 3D space. In our work, we are interested in connecting language with a dynamic modeling of the world. We show how to lift spatio-temporal features to a 4D representation based on 3D Gaussian Splatting. This enables an interactive interface where the user can spatiotemporally localize events in the video from text prompts. We demonstrate our system on public 3D video datasets of people and animals performing various actions.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gal Fiebelman, Tamir Cohen, Ayellet Morgenstern, Peter Hedman, Hadar Averbuch-Elor. 2024-10-14. 4-LEGS: 4D Language Embedded Gaussian Splatting. https://arxiv.org/abs/2410.10719
Cite the original work for its findings. Save a collection to share your selection of sources.