SearcharxivSearch

arXiv subjects

Changhai Chen

Publications and source records attributed to Changhai Chen.

4 recordsLinked to original sources

HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation

We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion Transformer (DiT)-based flow matching models to the billion-parameter scale within the motion generation domain, delivering instruction-following capabilities that significantly outperform current open-source benchmarks. Uniquely, we introduce a comprehensive, full-stage training paradigm -- including large-scale pretraining on over 3,000 hours of motion data, high-quality fine-tuning on 400 hours of curated data, and reinforcement learning from both human feedback and reward models -- to ensure precise alignment with the text instruction and high motion quality. This framework is supported by our meticulous data processing pipeline, which performs rigorous motion cleaning and captioning. Consequently, our model achieves the most extensive coverage, spanning over 200 motion categories across 6 major classes. We release HY-Motion 1.0 to the open-source community to foster future research and accelerate the transition of 3D human motion generation models towards commercial maturity.

cs.CV

Learning Audio-Driven Viseme Dynamics for 3D Face Animation

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme curves, and supports multilingual speech inputs. The core of our approach is a novel parametric viseme fitting algorithm that utilizes phoneme priors to extract viseme parameters from speech videos. With the guidance of phonemes, the extracted viseme curves can better correlate with phonemes, thus more controllable and friendly to animators. To support multilingual speech inputs and generalizability to unseen voices, we take advantage of deep audio feature models pretrained on multiple languages to learn the mapping from audio to viseme curves. Our audio-to-curves mapping achieves state-of-the-art performance even when the input audio suffers from distortions of volume, pitch, speed, or noise. Lastly, a viseme scanning approach for acquiring high-fidelity viseme assets is presented for efficient speech animation production. We show that the predicted viseme curves can be applied to different viseme-rigged characters to yield various personalized animations with realistic and natural facial motions. Our approach is artist-friendly and can be easily integrated into typical animation production workflows including blendshape or bone based animation.

cs.GR

A Selection Region Based Routing Protocol for Random Mobile ad hoc Networks with Directional Antennas

In this paper, we propose a selection region based multihop routing protocol with directional antennas for wireless mobile ad hoc networks, where the selection region is defined by two parameters: a reference distance and the beamwidth of the directional antenna. At each hop, we choose the nearest node to the transmitter within the selection region as the next hop relay. By maximizing the expected density of progress, we present an upper bound for the optimum reference distance and derive the relationship between the optimum reference distance and the optimum transmission probability. Compared with the results with routing strategy using omnidirectional antennas in \cite{Di:Relay-Region}, we find interestingly that the optimum transmission probability is a constant independent of the beamwidth, the expected density of progress with the new routing strategy is increased significantly, and the computational complexity involved in the relay selection is also greatly reduced.

cs.IT

A Selection Region Based Routing Protocol for Random Mobile ad hoc Networks

We propose a selection region based multi-hop routing protocol for random mobile ad hoc networks, where the selection region is defined by two parameters: a reference distance and a selection angle. At each hop, a relay is chosen as the nearest node to the transmitter that is located within the selection region. By assuming that the relay nodes are randomly placed, we derive an upper bound for the optimum reference distance to maximize the expected density of progress and investigate the relationship between the optimum selection angle and the optimum reference distance. We also note that the optimized expected density of progress scales as $Θ(\sqrtλ)$, which matches the prior results in the literature. Compared with the spatial-reuse multi-hop protocol in \cite{Baccelli:Aloha} recently proposed by Baccelli \emph{et al.}, in our new protocol the amount of nodes involved and the calculation complexity for each relay selection are reduced significantly, which is attractive for energy-limited wireless ad hoc networks (e.g., wireless sensor networks).

cs.IT