SearcharxivSearch

arXiv subjects

Jiameng Zhang

Publications and source records attributed to Jiameng Zhang.

4 recordsLinked to original sources

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-temporal video grounding. Our results suggest that visual feature extraction is the dominant bottleneck under sparse-frame inputs. Adapting only the final three ViT layers, 4% of total parameters, achieves 68.8% temporal mIoU and surpasses a zero-shot 8B model using dense inputs by 12.8 points. Language model fine-tuning, by contrast, offers negligible or negative returns. A boundary-aware sampling strategy, Hybrid16, further improves temporal mIoU by 26 points over uniform sampling when temporal boundaries are available. We conclude that for sparse-frame video grounding, training strategy dominates model scale: a fine-tuned 2B model consistently outperforms a zero-shot 8B model, with or without dense frame access.

cs.CV

Time-independent theoretical framework for stroboscopic nonlinear dynamics based on time-nonlocal response

Recent experiments have demonstrated the ability to manipulate nonlinear interactions via time modulation, giving rise to the so-called stroboscopic nonlinearity. To date, however, this phenomenon has not been subjected to a rigorous theoretical analysis. In this work, we clarify the physical mechanism underlying stroboscopic nonlinear dynamics based on time-nonlocal response and establish an effective time-independent model under suitable modulation conditions. The proposed model almost exactly reproduces the full time-dependent dynamics in the quasi-steady state and significantly outperforms empirical descriptions used previously. Our results provide a clear physical picture of stroboscopic nonlinear dynamics, and can be extended to other systems with time-nonlocal response, establishing a general framework for engineering nonlinear interactions through temporal modulation.

physics.optics

Towards Enhanced Immersion and Agency for LLM-based Interactive Drama

LLM-based Interactive Drama is a novel AI-based dialogue scenario, where the user (i.e. the player) plays the role of a character in the story, has conversations with characters played by LLM agents, and experiences an unfolding story. This paper begins with understanding interactive drama from two aspects: Immersion, the player's feeling of being present in the story, and Agency, the player's ability to influence the story world. Both are crucial to creating an enjoyable interactive experience, while they have been underexplored in previous work. To enhance these two aspects, we first propose Playwriting-guided Generation, a novel method that helps LLMs craft dramatic stories with substantially improved structures and narrative quality. Additionally, we introduce Plot-based Reflection for LLM agents to refine their reactions to align with the player's intentions. Our evaluation relies on human judgment to assess the gains of our methods in terms of immersion and agency.

cs.CL

Novel transition and Bellerophon state in coupled Stuart-Landau oscillators

We study synchronization in a system of Stuart-Landau oscillators with frequency-weighted coupling. For three typical unimodal frequency distributions, namely, the Lorentzian, the triangle, and the uniform, we found that the first-order transition occurs when the frequency distribution is relatively compact, while the synchronization transition is continuous when it is relatively wide. In both cases, there is a regime of Bellerophon state between the incoherent state and the synchronized state. Remarkably, we revealed novel transition behavior for such coupled oscillators with amplitudes, i.e., the regime of Bellerophon state actually contains two stages. In the first stage, the oscillators achieve chaotic phase synchronization; while in the second stage, oscillators form periodical phase synchronization. Our results suggest that Bellerophon state also exists in coupled oscillators with amplitude dynamics.

nlin.AO