Searcharxiv⌕ Search

arXiv · 2609.30975

A Comprehensive Study of Content Representations for Speech Synthesis

Abstract

Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Diego Torres, Axel Roebel, Nicolas Obin. 2026-09-25. A Comprehensive Study of Content Representations for Speech Synthesis. https://arxiv.org/abs/2609.30975

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

HuPER: A Human-Inspired Framework for Phonetic Perception

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic-phonetics evidence and linguistic knowledge. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER is also the first framework to enable adaptive, multi-path phonetic perception under diverse acoustic conditions. All training data, models, and code are open-sourced. Code and demo avaliable at https://github.com/Berkeley-Speech-Group/HuPER.

eess.AS↗

SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering

Spatial audio research for virtual and augmented reality, teleconferencing and hearing devices often represents sound fields in the Spherical Harmonics (SH) domain, known as Ambisonics. A typical study simulates a room, renders what a listener or a microphone array would capture in it, and processes those signals in the SH domain. We present SHroom (Spherical Harmonics ROOM), an open-source Python library that performs this whole workflow in one package, from room simulation to binaural rendering, head rotation, microphone-array simulation and Ambisonics encoding. Existing tools cover either the room simulation or the downstream processing, so researchers bridge them with ad-hoc code that is hard to reproduce and compare. In SHroom, every step operates on one shared signal type through one processing interface, so a simulated room flows through the complete chain without format conversions. Built on the image-source engine of `pyroomacoustics`, SHroom reproduces the Ambisonic Room Impulse Response (ARIR) of its spherical-harmonic receivers while computing it about 3x faster for SH orders 4 to 12. SHroom is available at 'https://github.com/Yhonatangayer/shroom' and installable via `pip install pyshroom`.

eess.AS↗

Exploring Second-Order Pattern Recognition in Speaker Recognition

In traditional pattern recognition tasks, neural networks are trained to recognise human-defined patterns (e.g. audio categories) in model inputs (e.g. audio). Meanwhile, some Explainable AI (XAI) methods explain latent patterns characterising the network's recognition of inputs as human-defined patterns; this work calls these latent patterns second-order patterns and proposes to discover them. Accordingly, we apply a hierarchical clustering algorithm to analyse whether our speaker recognition network's representations learned from known utterances naturally form hierarchical clusters. Each resulting cluster is a second-order pattern that characterises a context in our network's recognition of the known utterances as speaker identities. All discovered second-order patterns are then interpreted using the Hierarchical Cluster-Class Matching (HCCM) method. Moreover, we propose a new task, second-order pattern recognition, to identify which of the discovered second-order patterns characterising the recognition of known utterances also apply to unseen utterances, thereby characterising the recognition of unseen utterances. Accordingly, we design the Hierarchical Cluster Navigation and Assignment (HCNA) method. HCNA recognises a second-order pattern as applying to an unseen utterance when the utterance's network representation lies within the extrapolation space of the cluster regarded as that second-order pattern. Experimental results demonstrate that the extrapolation space introduced in HCNA substantially improves task performance.

eess.AS↗