Searcharxiv⌕ Search

arXiv · 2610.02582

Acoustic gap placement in second-language read speech production

Abstract

Speakers can differ not only in how much silence they produce but also in where interruptions fall, information that global pause counts obscure. We analyzed 115 publicly available Speech Accent Archive recordings of the same read passage: 57 speakers whose archive metadata listed Mandarin as their first language and a mainland-China birthplace, and 58 American English first-language speakers born in the U.S. Midwest. Word-level forced alignment was used to extract interword acoustic gaps and classify each position as punctuation-marked or unpunctuated. At a 250 ms threshold, punctuation-marked gaps were similar for the Mandarin and English groups (means 5.56 and 5.47), whereas unpunctuated-position gaps were more frequent in the Mandarin group (means 3.67 and 0.43; adjusted rate ratio 11.98, 95% confidence interval 6.53-22.00). A crossed-effects logistic model with random intercepts for speakers and passage positions confirmed that the contrast was disproportionately concentrated at unpunctuated positions (interaction odds ratio 9.44, 95% confidence interval 5.23-17.02). The pattern persisted at a 500 ms threshold, after removing positions adjacent to punctuation, after excluding severe transcript-deviation cases, and when analysis was restricted to acoustically confirmed silence. Nearly half of the Mandarin group's unpunctuated gaps occurred at positions classified as within-phrase in an exploratory single-coder annotation. In this fixed-passage corpus, acoustic gap placement relative to textual and syntactic structure distinguished the groups more clearly than punctuation-marked pausing. The results support placement-sensitive measurement as a reproducible dimension of breakdown fluency while not establishing a causal mechanism, proficiency difference, or perceptual consequence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Peyman Jahanbin. 2026-10-01. Acoustic gap placement in second-language read speech production. https://arxiv.org/abs/2610.02582

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Controllable Accent Normalization via Discrete Diffusion

Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control. The implementation is available at https://github.com/P1ping/DLM-AN

eess.AS↗

FASTDIAR: Frame-level speaker encoder for Streaming Diarization

Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.

eess.AS↗

Unsupervised Instantaneous Phase and Frequency Tracking by Inverse Voice Synthesis

Knowledge-driven neural vocoders struggle to learn reliable fundamental frequency end-to-end, because spectral objectives provide weak supervision of periodic structure and lack phase information. We address this with a source-filter model whose alias-free additive source makes the instantaneous phase of the glottal cycle explicit; differentiating it yields the instantaneous frequency, and thus $F_0$, without an external tracker. Waveform error supervises only the deterministic harmonic path, while a spectral loss covers the full signal. On M4Singer and LM-SSD, the reconstruction is phase-aligned, reaching a signal-to-reconstruction-error ratio of 8.1 dB, while neural baselines remain negative. However, GOLF, given an external $F_0$, still reaches lower spectral distortion. On LM-SSD, the recovered $F_0$ attains the highest overall accuracy of any method tested, including supervised neural pitch trackers applied off the shelf, and the glottal closure instants come within 0.53 points of REAPER's identification rate, without any $F_0$ label.

eess.AS↗