arXiv · 2609.35728
FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning
Abstract
We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini reference-to-video backbone to accept rolling action prompts, streaming audio, and dynamically updated image, audio, and video references. Video-driven rotary positional embeddings align reference chunks with the generation timeline, while reference-plus-image conditioning and partially noised historical motion frames preserve appearance and avoid accumulated drift. Second, a Proactive Interaction Agent separates pre-online planning from online scheduling and response: it prepares a persona, a long-horizon agenda, and reusable multimodal skills in advance, then autonomously schedules behaviors, responds to audience input, and handles interruptions during a live session. FlowAct-R2 supports real-time 720p generation and hour-scale streaming across entertainment streaming, live shopping, video chatting, and live vlogging.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang, Wentao Hu, Yuxuan Luo, Yuan Zhang, Mingyuan Gao. 2026-09-28. FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning. https://arxiv.org/abs/2609.35728
Cite the original work for its findings. Save a collection to share your selection of sources.