arXiv · 2602.13185
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
Abstract
Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more robust and scalable pathway. We propose FlexAM, a unified framework built upon a novel 3D control signal. This signal represents video dynamics as a point cloud, introducing three key enhancements: multi-frequency positional encoding to distinguish fine-grained motion, depth-aware encoding, and a flexible control signal for balancing precision and generalization. This representation allows FlexAM to effectively disentangle appearance and motion, enabling a wide range of tasks including I2V/V2V editing, camera control, and spatial object editing. Extensive experiments demonstrate that FlexAM achieves superior performance across all evaluated tasks.
Explore related subjects
Keep this discovery
Mingzhi Sheng, Zekai Gu, Peng Li, Cheng Lin, Hao-Xiang Guo, Ying-Cong Chen, Yuan Liu. 2026-02-13. FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control. https://arxiv.org/abs/2602.13185
Cite the original work for its findings. Save a collection to share your selection of sources.