Searcharxiv⌕ Search

arXiv · 2610.09846

DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion

Abstract

Holistic co-speech animation is prone to averaging in both motion representation and speech conditioning. In coordinate-space diffusion, slow body posture, mid-frequency gesture strokes, and fast hand or facial details are entangled in one prediction target, often producing low-variance, over-smoothed motion. Meanwhile, dense rhythmic and acoustic cues can dominate sparse content-specific information under fixed multimodal fusion. We present DynaConTalk, a wavelet-constrained diffusion framework for long-form and controllable holistic co-speech motion generation. Diffusion operates in stationary wavelet transform (SWT) coefficient space, whose temporally aligned bands separate coarse posture evolution, gesture strokes, and fine expressive details. Our dynamic gating network preserves HuBERT and speaker identity as a base and selectively adds rhythm, mel, and transcript features through motion-state- and noise-aware residual gates. Attention pooling and learned depth routing deliver complementary conditions to each denoising stage, while a frame-resolution rhythm path preserves precise timing. A signed proposal-consensus update then reconciles these conditions with the evolving motion state. Matched-noise constraint injection uses the same sampling interface for history continuation and localized keypose repair, and extends to reference-guided control. Separate body-hand and facial denoisers, followed by inverse SWT and a pose-driven root regressor, produce holistic motion. Experiments evaluate generation quality, facial accuracy, temporal continuity, and controllable editing. Code, models, and the interactive editing interface are available at https://github.com/zhuyifeiabcd1/DynaConTalk.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yifei Zhu, Yangyang Cai, Mingyi Shi, Miao Cheng, Lin Gu, Taku Komura, Yoshifumi Kitamura. 2026-10-07. DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion. https://arxiv.org/abs/2610.09846

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HistCAD: Constraint-Aware Parametric CAD Histories for Evaluating Editability

Sketch constraints specify geometric conditions for constructing and modifying parametric CAD models. We study whether predicted constraints allow a given history to reproduce the required initial model and support prescribed dimensional edits. We introduce HistCAD, an executable representation and dataset whose Academic and Industrial collections contain 180,495 parametric construction histories with entity-referenced sketch constraints and retained feature operations. Predictors receive these histories with the geometry and feature definitions retained and explicit sketch constraints removed. They generate constraints for every sketch without seeing the edit request. The benchmark compares models built with alternative constraint sets for the same history under the same dimensional edit. An edit succeeds when the model reproduces the required initial geometry, reaches the target value, preserves specified relations and unedited dimensions in the target sketch, and rebuilds through the complete history. A predictor trained on both collections and supplied with descriptions of the input histories achieves overall edit success of 52.4% on Academic and 29.0% on Industrial. Models retaining only endpoint-connectivity constraints in the target sketch and the history's constraints elsewhere can reach the target and rebuild while failing preservation. For all-sketch predictions, we retain the target-sketch prediction and restore the history's constraints in other sketches. More models then reproduce the required initial geometry, and some of these newly matched models complete the edit. HistCAD connects constraint learning to the construction and revision of parametric CAD models.

cs.GR↗

Event-T2M: Event-level Conditioning for Complex Text-to-Motion Synthesis

Text-to-motion generation has advanced with diffusion models, yet existing systems often collapse complex multi-action prompts into a single embedding, leading to omissions, reordering, or unnatural transitions. In this work, we shift perspective by introducing a principled definition of an event as the smallest semantically self-contained action or state change in a text prompt that can be temporally aligned with a motion segment. Building on this definition, we propose Event-T2M, a diffusion-based framework that decomposes prompts into events, encodes each with a motion-aware retrieval model, and integrates them through event-based cross-attention in Conformer blocks. Existing benchmarks mix simple and multi-event prompts, making it unclear whether models that succeed on single actions generalize to multi-action cases. To address this, we construct HumanML3D-E, the first benchmark stratified by event count. Experiments on HumanML3D, KIT-ML, and HumanML3D-E show that Event-T2M matches state-of-the-art baselines on standard tests while outperforming them as event complexity increases. Human studies validate the plausibility of our event definition, the reliability of HumanML3D-E, and the superiority of Event-T2M in generating multi-event motions that preserve order and naturalness close to ground-truth. These results establish event-level conditioning as a generalizable principle for advancing text-to-motion generation beyond single-action prompts.

cs.GR↗

SECA: Strict-Elastic Compact Activation for Differentiable Elastoplastic Deformation

Return mapping keeps an exact elastic domain but switches its response tangent at yield, while smoothing laws with a sub-yield tail admit plastic flow below the threshold. We present SECA, a Strict-Elastic Compact Activation method for differentiable elastoplastic simulation. SECA prescribes a bounded $C^2$ activation over a finite positive-overstress interval and integrates it into a Perzyna-type flow law, so the implicit update returns exactly zero plastic increment for sub-yield trials while smoothing the onset. For scalar and proportional isotropic $J_2$ updates we derive an exact decomposition of the plastic-increment discrepancy into finite-width and viscosity contributions, and an explicit algebraic budget for the transition width. The update is embedded in finite-deformation loading histories with contact and complete unloading, and differentiated to released equilibria. Controlled indentation results stay closely matched. In the tested coarse cube-indentation run, where the transition is crossed without intermediate increments, SECA reduces whole-run Newton iterations by 47% and backtracking halvings by 93% against return mapping. A four-method comparison evaluates released-shape error and computational cost across loading partitions, and on a contact scene the path derivative drives a bounded control task to its target in three iterations. Further examples show elastic recovery and residual deformation under compression, tension and bending.

cs.GR↗