SearcharxivSearch

arXiv subjects

Chunjiang He

Publications and source records attributed to Chunjiang He.

3 recordsLinked to original sources

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios

Spoken Language Understanding (SLU) is moving from task-specific pipelines toward large audio language models (LALMs) that generate natural-language responses. However, existing speech benchmarks mainly focus on single-speaker settings or isolated subtasks, leaving speaker-centric understanding in realistic multi-speaker conversations insufficiently evaluated. We introduce MSU-Bench, a diagnostic benchmark for multi-speaker conversational understanding, covering 16 speaker-centric tasks and 2,300 QA instances in a two-tier framework from speaker grounding to dialogue reasoning. We build a Gemini-assisted annotation and QA generation pipeline with human-in-the-loop verification, achieving high QA validity and strong agreement between human answers and verified labels. We further analyze speaker-referencing schemes and diagnostic error types to reveal bottlenecks in speaker grounding and reasoning. Experiments reveal clear gaps across model families, with closed-source systems leading overall but all models still facing challenges in complex speaker grounding and multi-speaker reasoning. The benchmark annotations, metadata, and evaluation scripts will be available at the GitHub repository: https://github.com/ASLP-lab/MSU-Bench.

eess.AS

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

Recent advances in AudioLLMs have enabled spoken dialogue systems to move beyond turn-based interaction toward real-time full-duplex communication, where the agent must decide when to speak, yield, or interrupt while the user is still talking. Existing full-duplex approaches either rely on voice activity cues, which lack semantic understanding, or on ASR-based modules, which introduce latency and degrade under overlapping speech and noise. Moreover, available datasets rarely capture realistic interaction dynamics, limiting evaluation and deployment. To mitigate the problem, we propose \textbf{FastTurn}, a unified framework for low-latency and robust turn detection. To advance latency while maintaining performance, FastTurn combines streaming CTC decoding with acoustic features, enabling early decisions from partial observations while preserving semantic cues. We also release a test set based on real human dialogue, capturing authentic turn transitions, overlapping speech, backchannels, pauses, pitch variation, and environmental noise. Experiments show FastTurn achieves higher decision accuracy with lower interruption latency than representative baselines and remains robust under challenging acoustic conditions, demonstrating its effectiveness for practical full-duplex dialogue systems.

cs.SD

Optical microcavity characterization via resonance spectra and modes

This paper describes how resonance spectra and mode profiles can be used to characterize and quantify the mode-shaping effects in open-access plano-concave optical microcavities. The presented semi-analytic theory is based on the application of perturbation theory to the roundtrip evolution of the optical field. It includes various mirror-shape and nonparaxial effects and extends the nonparaxial theory presented by van Exter et al. (2022, Phys. Rev. A 106, 013501) and verified by Koks et al. (2022, Phys. Rev. A 105, 063502) to the common case of an anisotropic Gaussian mirror. The presented measurements and analyses of resonance spectra and mode profiles demonstrate how the different mode-shaping effects can be individually distinguished and quantified. Spin-orbit coupling, which is one of the nonparaxial effects, is prominently visible in the intriguing polarization patterns of the resonant modes, while polarization tomography yields the shape-induced birefringence and associated polarization splitting of the fundamental modes.

physics.optics