SearcharxivSearch

arXiv subjects

Huilai Li

Publications and source records attributed to Huilai Li.

6 recordsLinked to original sources

Identity-Aware Human-Object Interaction Motion Captioning

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subject identity. To address this limitation, we introduce Identity-Aware Human-Object Interaction Motion Captioning task. This task requires each generated caption to specify both the subject identity and the corresponding HOI motion. For example, the model generates "Sub_ID lifts the chair" rather than "A person lifts the chair". For this task, we design identity-aware HOI motion captions based on the BEHAVE and InterCap datasets. We further propose ID-HOINet, which learns from multi-view videos while supporting single-view identity-aware HOI motion caption generation. ID-HOINet contains two core components: Multi-View Identity-Motion Learning Module (MVIML) and Two-Stage Caption Rewriting Strategy (TSCR). MVIML learns from multi-view videos by modeling dependencies across temporal stages and camera viewpoints, capturing identity and interaction motion features. At inference, the TSCR first retrieves the subject identity and generates identity-agnostic HOI motion captions. TSCR then rewrites these captions with the predicted identity to produce the final identity-aware HOI motion captions. Experiments demonstrate that ID-HOINet achieves state-of-the-art performance. Code will be released upon acceptance.

cs.CV

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, existing research advances along two main paths: pre-training pseudo-label generators for fine-grained cross-modal semantic guidance, or refining AVVP model architectures to enhance audio-visual fusion. However, since audio and visual signals are typically unaligned, achieving accurate video parsing fundamentally relies on precise perception of uni-modal events. Yet these multi-modal focused strategies excessively emphasize multi-modal fusion while inadequately guiding and preserving uni-modal semantics, resulting in noisy pseudo-labels and sub-optimal video parsing performance. This paper proposes a novel framework that enhances uni-modal representations for both the pseudo-label generator and the AVVP model. Specifically, we introduce a similarity-based label migration approach to annotate pre-training data, thereby enabling the pseudo-label generator to better understand uni-modal events. We also employ a soft-constrained manner to refine modeling of uni-modal features in parallel with multi-modal fusion. These designs enable coordinated attention to both uni-modal and cross-modal representations, thus boosting the localization performance for events. Extensive experiments show that our method outperforms state-of-the-art methods in both pseudo-label and AVVP performance.

cs.CV

DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition

Micro expression recognition (MER) is crucial for inferring genuine emotion. Applying a multimodal large language model (MLLM) to this task enables spatio-temporal analysis of facial motion and provides interpretable descriptions. However, there are still two core challenges: (1) The entanglement of static appearance and dynamic motion cues prevents the model from focusing on subtle motion; (2) Textual labels in existing MER datasets do not fully correspond to underlying facial muscle movements, creating a semantic gap between text supervision and physical motion. To address these issues, we propose DEFT-LLM, which achieves motion semantic alignment by multi-expert disentanglement. We first introduce Uni-MER, a motion-driven instruction dataset designed to align text with local facial motion. Its construction leverages dual constraints from optical flow and Action Unit (AU) labels to ensure spatio-temporal consistency and reasonable correspondence to the movements. We then design an architecture with three experts to decouple facial dynamics into independent and interpretable representations (structure, dynamic textures, and motion-semantics). By integrating the instruction-aligned knowledge from Uni-MER into DEFT-LLM, our method injects effective physical priors for micro expressions while also leveraging the cross modal reasoning ability of large language models, thus enabling precise capture of subtle emotional cues. Experiments on multiple challenging MER benchmarks demonstrate state-of-the-art performance, as well as a particular advantage in interpretable modeling of local facial motion.

cs.CV

ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lacking cross-modal semantic bridging in intermediate layers. This causes modality semantic gap for further fusion, making it difficult to distinguish between event-related content and irrelevant background content. Moreover, they rarely consider the correlations between events, which limits the model to infer concurrent events among complex scenarios. In this paper, we incorporate multi-stage semantic guidance and multi-event relationship modeling, which respectively enable hierarchical semantic understanding of audio-visual events and adaptive extraction of event dependencies, thereby better focusing on event-related information. Specifically, our eventaware semantic guided network (ESG-Net) includes a early semantics interaction (ESI) module and a mixture of dependency experts (MoDE) module. ESI applys multi-stage semantic guidance to explicitly constrain the model in learning semantic information through multi-modal early fusion and several classification loss functions, ensuring hierarchical understanding of event-related content. MoDE promotes the extraction of multi-event dependencies through multiple serial mixture of experts with adaptive weight allocation. Extensive experiments demonstrate that our method significantly surpasses the state-of-the-art methods, while greatly reducing parameters and computational load. Our code will be released on https://github.com/uchiha99999/ESG-Net.

cs.MM

Multicycle dynamics and high-codimension bifurcations in SIRS epidemic models with cubic psychological saturated incidence

This study investigates bifurcation dynamics in an SIRS epidemic model with cubic saturated incidence, extending the quadratic saturation framework established by Lu, Huang, Ruan, and Yu (Journal of Differential Equations, 267, 2019). We rigorously prove the existence of codimension-three Bogdanov-Takens bifurcations and degenerate Hopf bifurcations, demonstrating the coexistence of three limit cycles within a single epidemiological model, a phenomenon that is rarely documented and exhibits significant dynamical complexity. Our analysis reveals that both the infection rate $κ$ (through specific inequality conditions) and psychological effect thresholds critically govern disease dynamics: from complete eradication to various persistence patterns, including multiple periodic oscillations and coexistent steady states. By innovatively applying singularity theory, we characterize the topology of the bifurcation set through the local unfolding of singularities and the identification of nondegenerate singularities for fronts. Numerical simulations verify the emergence of three limit cycles in monotonic parameter regimes and two limit cycles in nonmonotonic regimes. This work advances existing bifurcation research by incorporating higher-order interactions and comprehensive singularity analysis, thereby providing a mathematical foundation for decoding complex transmission mechanisms critical to the design of public health strategies.

math.DS

Optimal resource control in reaction diffusion advection population model

This paper is to investigate the control problem of maximizing the net benefit of a single species while the cost of the resource allocation is minimized in a population model which can be described by a reaction diffusion advection equation of logistic type with spatial-temporal resource control coefficient. The existence of an optimal control is established and the uniqueness and characterization of the optimal control are investigated. Numerical simulation illustrate several cases with Dirichlet and Neumann boundary conditions.

math.OC