Searcharxiv⌕ Search

arXiv · 2610.11150

SmoothConv and DuplexConv: Complementary Mandarin Multi-Party Conversational Speech Corpora for Speech Interaction

Abstract

Recent advances in large audio language models (LALMs) have driven the development of natural and intelligent speech interaction systems. Such systems need to model complex conversational behaviors, including turn-taking, overlapping speech, and speaker coordination. Multi-party conversations provide a realistic setting for studying these behaviors, yet existing Mandarin conversational corpora often lack synchronized participant-level speech tracks and comprehensive annotations. In this work, we introduce SmoothConv and DuplexConv, two complementary Mandarin multi-party conversational speech corpora totaling 2,100 hours. SmoothConv provides human-verified conversations for reliable analysis and evaluation, while DuplexConv offers large-scale automatically annotated conversations through a scalable pipeline for model training. Both corpora provide synchronized participant-level speech tracks and multi-dimensional fine-grained annotations. We further release the SmoothConv Benchmark and evaluate these resources on speech separation, multi-speaker automatic speech recognition (MSASR), and turn detection tasks. Experimental results demonstrate the utility of the proposed resources for multi-party speech interaction modeling. The datasets, benchmark, and related resources are publicly available.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Chengyou Wang, Mingchen Shao, Chunjiang He, Zeyu Zhu, Jierui Guo, Bingshen Mu, Zikai Liu, Hanke Xie, Yuhang Dai, Zhou Zhu, Lei Xie. 2026-10-08. SmoothConv and DuplexConv: Complementary Mandarin Multi-Party Conversational Speech Corpora for Speech Interaction. https://arxiv.org/abs/2610.11150

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Sound Field Interpolation Using Physics-Informed Extreme Learning Machine with Pre-Training

Numerous machine learning-based sound field interpolation methods have been proposed. In particular, physics-informed neural networks (PINNs) can accurately interpolate sound fields from a small number of microphones. However, their high computational cost and long training time pose practical challenges for applications requiring real-time processing or online learning. To address this, we propose a hybrid framework that combines PINN-based pre-training with a physics-informed extreme learning machine (PIELM) tailored for acoustic fields. By replacing iterative PINN fine-tuning for each target sound field with closed-form output-layer adaptation using hidden-layer weights pre-trained by PINN, the proposed method efficiently interpolates unknown sound fields from limited observations. Simulation results under simplified one-dimensional free-field conditions demonstrate that, given a pre-trained model, the proposed method achieves interpolation accuracy comparable to that of PINN-based fine-tuning while reducing the adaptation time by more than three orders of magnitude.

eess.AS↗

CTC-Seeded Token Edit Refinement for Non-Autoregressive Speech Recognition

Non-autoregressive automatic speech recognition (ASR) enables parallel decoding, but many refinement-based methods begin from random, fully masked, or fixed-length token sequences, requiring multiple iterations to reconstruct the complete transcript. We instead formulate ASR decoding as a variable-length edit refinement of a greedy connectionist temporal classification (CTC) hypothesis. An acoustic-conditioned Edit Flow decoder operates directly on the collapsed CTC hypothesis, predicting insertion, deletion, and substitution operations in parallel. The Edit Flow decoder is jointly trained with a CTC model using a continuous-time discrete diffusion loss. During inference, we find that just two edit steps yield substantial Word Error Rate (WER) reductions, and classifier-free guidance (CFG) further enhances recognition quality by focusing the model on audio features. We also constrain edit proposals using CTC confidence to improve accuracy. Finally, ablation studies validate our design choices, while decoder pretraining and pretrained encoder integration yield significant additional performance gains.

eess.AS↗

NormToken: Speaker- and Duration-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement

Semantic speech tokens should preserve linguistic content while suppressing utterance-specific acoustic and duration variation. However, existing speech-to-unit (S2U) tokenizers often retain speaker-related acoustic characteristics and duration information. To address this issue, we propose NormToken, an iterative semantic token purification framework that alternates S2U and text-to-unit (T2U) training. In each iteration, the T2U model produces text-derived tokens to supervise a newly initialized S2U tokenizer. The resulting S2U tokens are then used as targets for the next T2U iteration. This cycle drives the two models toward a shared, text-predictable token space. Experiments on Mandarin and English demonstrate improved S2U--T2U agreement and parallel-utterance token consistency. De-tokenizers trained on the initial and refined tokens further show that refined tokens maintain comparable WER and CER while improving speaker similarity in both voice cloning and text-to-speech synthesis. In voice cloning, refined tokens produce speaking rates closer to the acoustic reference, suggesting reduced dependence on source duration information. Audio samples are available at https://hanlin1004.github.io/normtoken_demopage.

eess.AS↗