arXiv · 2607.22717
TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians
Abstract
While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Marek Lisowski, Łukasz Smoliński, Kornel Howil, Piotr Biliński, Marcin Mazur, Przemysław Spurek. 2026-07-21. TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians. https://arxiv.org/abs/2607.22717
Cite the original work for its findings. Save a collection to share your selection of sources.