SearcharxivSearch

arXiv subjects

Yujie Liao

Publications and source records attributed to Yujie Liao.

7 recordsLinked to original sources

Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods

This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.

eess.AS

On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin

Recent large-scale ASR models already achieve strong Mandarin recognition accuracy and have some ability to recognize Chinese dialects. However, their dialect recognition accuracy is still limited in real-world speech. Direct dialect adaptation can lower dialect CER, but it may also raise Mandarin CER. We therefore study how to adapt a capable ASR model to improve multi-dialect recognition without degrading Mandarin recognition. We adopt an adaptation pipeline where continual pre-training (CPT) and dialect supervised fine-tuning (SFT) provide a strong foundation, and On-Policy Self-Distillation (OPSD) serves as the final refinement. OPSD addresses the train--test mismatch in autoregressive ASR by training the student model on its own decoded prefixes while a frozen teacher, conditioned on the reference transcript as privileged context, provides soft token-level targets. This replaces hard cross-entropy updates on dialect data with distillation, preserving Mandarin ability while refining dialect recognition. We instantiate the framework with Qwen3-ASR-1.7B and evaluate it on public and internal Mandarin and dialect test sets. Under matched refinement data and schedule, OPSD improves dialect recognition without raising Mandarin CER, whereas continued teacher-forced fine-tuning increases Mandarin CER. We will release the model weights and evaluation scripts.

eess.AS

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

Recent advances in large language models (LLMs) and multimodal LLMs (MLLMs) have created new opportunities for wearable speech interfaces, with smart glasses providing an egocentric platform for continuous audio sensing and assistance. However, speech recognition and understanding in this setting remain challenging because of dynamic acoustic conditions, speaker overlap, and the spatial ambiguity introduced by wearer-centered recording geometry. To support systematic evaluation in this setting, we introduce the IEEE SLT 2026 SmartGlasses Challenge for egocentric multi-speaker speech processing. The challenge consists of two tracks, Dyadic Dialogue Understanding and Multi-party Meeting Understanding, and jointly evaluates Time-Stamped Speaker-Attributed Automatic Speech Recognition (TSA-ASR) and Spoken Language Understanding (SLU). It is built on a 106-hour four-channel egocentric speech dataset containing 714 sessions collected in real-world scenarios. This paper describes challenge tasks, dataset construction, submissions, and summarizes the main findings from the shared evaluation. The results show that heavy speaker overlap remains a major factor affecting TSA-ASR performance, while paralinguistic acoustic understanding continues to be difficult for current audio-language models in complex SLU settings. Further details can be found on the official challenge website.

eess.AS

OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs

Recent advancements in Speech Large Language Models have significantly enhanced multi-dimensional speech understanding. However, the majority of high-performance frameworks are predominantly optimized for GPU centric ecosystems and proprietary backbones, creating a significant gap for deployment on non-CUDA computing infrastructures. In this paper, we present OSUM-Pangu, a fully open-source speech understanding foundation model developed on a completely non-CUDA software and hardware stack. By integrating an audio encoder with the openPangu-7B LLM backbone, we successfully implement the entire training and inference pipeline on the Ascend NPU platform. To facilitate efficient task alignment under non-CUDA resource constraints, we adopt a practical training process that sequentially bridges speech perception and user intent recognition. Experimental results demonstrate that OSUM-Pangu achieves task accuracy comparable to mainstream GPU-based models while maintaining robust natural language interaction capabilities. Our work provides a reproducible, non-CUDA baseline for the open-source speech community, promoting the independent evolution of multimodal intelligence.

cs.SD

Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning

Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and on-screen text. To tackle such CAVSR (AVSR including rich visual Context), we propose VASR designed to "see" and reason the visual context to improve speech recognition. Specifically, we construct an Audio-Visual Chain-of-Thought (AV-CoT) that explicitly enforces intermediate cross-modal grounding between acoustic signals and visual evidence. This evidence-driven reasoning mitigates the "single-modality dominance" problem, where models either over-rely on visual context or fail to utilize it. Besides, to address the data scarcity, we construct and release a corresponding data pipeline and test set. Experiments show that AV-CoT effectively mitigates the single-modality dominance, achieving state-of-the-art performance in CAVSR. The project is open-sourced.

cs.SD

Generalized Varying Coefficient Mediation Models

Motivated by an analysis of causal mechanism from economic stress to entrepreneurial withdrawals through depressed affect, we develop a two-layer generalized varying coefficient mediation model. This model captures the bridging effects of mediators that may vary with another variable, by treating them as smooth functions of this variable. It also allows various response types by introducing the generalized varying coefficient model in the first layer. The varying direct and indirect effects are estimated through spline expansion. The theoretical properties of the estimated direct and indirect coefficient functions, including estimation biases, asymptotic distributions, and so forth, are explored. Simulation studies validate the finite-sample performance of the proposed estimation method. A real data analysis based on the proposed model discovers some interesting behavioral economic phenomenon, that self-efficacy influences the deleterious impact of economic stress, both directly and indirectly through depressed affect, on business owners' withdrawl intentions.

math.ST

Flat-band based ferromagnetic semiconducting state in the graphitic C$_4$N$_3$ monolayer

A new set of lattice-models based on the hexagonal $\sqrt{N}\times\sqrt{N}$ super-cells of the well-known honeycomb lattice with single-hole defect (HL-D-1/2N) are proposed to realize the nontrivial isolated flat-bands. Through performing both tight-binding and density functional theory calculations, we demonstrate that the experimentally realized graphitic carbon nitride (Adv. Mater., 22, 1004, 2010; Nat. Commun., 9, 3366, 2018), the HL-D-1/8 based C$_4$N$_3$, is a perfect system to host such flat bands. For the flat high-energy P-6m2 C$_4$N$_3$ structure, it displays the ferromagnetic half-metallicity which is not related to the isolated flat bands. However, the P-6m2 C$_4$N$_3$ structure is dynamically unstable. Using a structure searching method based on group and graph theory, we find that a new corrugated Pca21 C4N3 structure has the lowest energy among all known C$_4$N$_3$ structures. This Pca21 C$_4$N$_3$ structure is an intrinsic ferromagnetic half-semiconductor (Tc$\approx$241 K) with one semiconducting spin-channel (1.75 eV) and one insulating spin-channel (3.64 eV), which is quite rare in the two-dimensional (2D) systems. Its ferromagnetic semiconducting property originates from the isolated p$_z$-state flat-band as the corrugation shift the flat band upward to the Fermi level. Interestingly, this Pca21 C$_4$N$_3$ structure is found to be piezoelectric and ferroelectric, which makes C$_4$N$_3$ an unusual transition-metal-free 2D multiferroic.

cond-mat.mtrl-sci