Searcharxiv⌕ Search

arXiv · 2609.36737

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

Abstract

The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann. 2026-09-29. Reconstructing the Vocal Tract with Differentiable Acoustic Simulation. https://arxiv.org/abs/2609.36737

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come with several constraints such as increased sensory requirements, computational cost, and modality synchronization, to mention a few. These challenges constrain the direct uses of these multimodal solutions in real-world applications. In this work, we develop approaches where the learning happens with all available modalities but the deployment or inference is done with just one or reduced modalities. To do so, we propose a Multimodal Training and Unimodal Deployment (MUTUD) framework which includes a Temporally Aligned Modality feature Estimation (TAME) module that can estimate information from missing modality using modalities present during inference. This innovative approach facilitates the integration of information across different modalities, enhancing the overall inference process by leveraging the strengths of each modality to compensate for the absence of certain modalities during inference. We apply MUTUD to various audiovisual speech tasks and show that it can reduce the performance gap between the multimodal and corresponding unimodal models to a considerable extent. MUTUD can achieve this while reducing the model size and compute compared to multimodal models, in some cases by almost 80%.

cs.SD↗

Beyond Acoustic Prefixes: Persistent Access to Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

Large language models (LLMs) provide strong linguistic priors for serialized output training (SOT), yet LLM-based multi-talker ASR degrades substantially as the number of overlapping talkers increases. Conventional systems expose acoustic evidence primarily through an initial projected mixture prefix, requiring the decoder to preserve and recover talker-relevant information indirectly throughout autoregressive generation. We first examine whether this limitation can be resolved by enriching the static prefix using discrete connectionist temporal classification (CTC) tokens, hybrid token--acoustic prompts, and continuous talker-specific representations. The results suggest that acoustic content alone does not fully address the conditioning bottleneck. We therefore extend onset-based serialization from the output target to the acoustic-conditioning pathway and introduce persistent decoder-side access to SOT-aligned serialized acoustic memory. Talker-specific representations are organized in utterance-onset order and retained as external acoustic memory, while the conventional mixture prefix provides a complementary global-conditioning path. LLM layers query this memory throughout generation through gated residual cross-attention. We further introduce a second adaptation stage that jointly applies low-rank updates to the acoustic-retrieval pathway and selected LLM self-attention projections. Experiments on LibriMix show consistent improvements over static-prefix prompting. These results indicate that effective LLM-based multi-talker ASR depends not only on providing richer acoustic representations, but also on maintaining persistent access to acoustic evidence structured according to the serialized output.

cs.SD↗

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which extends Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a preliminary increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni on 18 of 21 comparable turn-taking, overlap-behavior, and timing metrics and MiniCPM-o 4.5 on 19 of 22, with gains in both interaction decisions and response timing. On human-recorded HumDial-FDBench, it attains the top Final score (69.6) of the compared duplex models.

cs.SD↗