SearcharxivSearch

arXiv subjects

Huaize Liu

Publications and source records attributed to Huaize Liu.

3 recordsLinked to original sources

MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding

First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the appearance redundancy, the remaining motion is typically compressed with a homogeneous latent geometry. Such representations use the same temporal support for broad scene evolution and fine-grained, frame-specific details. We introduce MotionStrata, which organizes a fixed motion budget into Global Motion and Detailed Motion. Temporally compressed Global queries summarize broad evolution, whereas frame-aligned Detailed queries preserve fine-grained structures whose configuration varies across frames. Frequency-guided routing and coarse-to-fine training establish this hierarchy without increasing motion dimensionality. Experiments show that MotionStrata maintains high reconstruction quality under aggressive compression and outperforms uniform and alternative grouped representations. Additional experiments evaluate hierarchical representation, downstream generation, and decoding cost. These results support hierarchical motion organization as a useful design principle for compact video autoencoding.

cs.CV

A Self-supervised Motion Representation for Portrait Video Generation

Recent advancements in portrait video generation have been noteworthy. However, existing methods rely heavily on human priors and pre-trained generative models, Motion representations based on human priors may introduce unrealistic motion, while methods relying on pre-trained generative models often suffer from inefficient inference. To address these challenges, we propose Semantic Latent Motion (SeMo), a compact and expressive motion representation. Leveraging this representation, our approach achieve both high-quality visual results and efficient inference. SeMo follows an effective three-step framework: Abstraction, Reasoning, and Generation. First, in the Abstraction step, we use a carefully designed Masked Motion Encoder, which leverages a self-supervised learning paradigm to compress the subject's motion state into a compact and abstract latent motion (1D token). Second, in the Reasoning step, we efficiently generate motion sequences based on the driving audio signal. Finally, in the Generation step, the motion dynamics serve as conditional information to guide the motion decoder in synthesizing realistic transitions from reference frame to target video. Thanks to the compact and expressive nature of Semantic Latent Motion, our method achieves efficient motion representation and high-quality video generation. User studies demonstrate that our approach surpasses state-of-the-art models with an 81% win rate in realism. Extensive experiments further highlight its strong compression capability, reconstruction quality, and generative potential.

cs.CV

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals-such as audio, text, and labels-to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.

cs.CV