SearcharxivSearch

arXiv subjects

Xingbei Chen

Publications and source records attributed to Xingbei Chen.

2 recordsLinked to original sources

EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes

Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typically rely on coarse labels or prompt-only control. Although modern diffusion transformers (DiTs) such as FLUX improve visual fidelity, they are not designed to incorporate structured affective representations. We present EmoSpace, an immersive affective content generation framework guided by fine-grained emotion prototypes, transforming free-form text and emotion descriptions into fine-grained affective imagery through three coordinated components. First, to represent sub-emotion variation beyond conventional categorical models, EmoSpace learns a hierarchical bank of 256 prototypes with input-conditioned adaptation through vision-language alignment. Second, Prototype-Conditioned Steering converts these prototypes into DiT-compatible generation signals through multi-pathway injection and temporal blending, while Iterative Prompt Refinement enriches prompts with prototype-aligned sub-emotion descriptors. Third, Affect-Grounded Modulation coordinates emotion conditioning with controllable LoRAs for panoramic, stylized, and multi-conditional generation. Through quantitative and qualitative evaluations, EmoSpace improves fine-grained emotional alignment while maintaining high aesthetic quality. Our user study shows that EmoSpace outputs are perceived as more emotionally aligned than baseline results and more suitable for immersive scene design. Additionally, we find that immersive presentation alters emotional perception and increases emotional engagement. Together, these findings inform the design of emotion-aware generative systems for immersive media, with potential applications including education, immersive storytelling, and artistic creation. We will release our code and models to facilitate future research along this line.

cs.CV

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistically styled videos, but also provides practical methods for enhancing emotional expression in video generation.

cs.CV