Searcharxiv⌕ Search

arXiv · 2609.36324

Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

Abstract

Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their generative trajectories. Recent few-step flow-map distillation approaches for TTS construct targets from numerically integrated teacher trajectories, creating a trade-off between target accuracy and training cost. We propose Local Flow-Map Distillation (LFMD), which adapts Eulerian Map Distillation to conditional TTS and avoids teacher trajectory integration during target construction. For inference, we derive a sampling schedule (TD-DP) from teacher dynamics and consistency of the learned maps, with a single cost graph supporting multiple NFE budgets without external audio-metric evaluation. Because scheduling offers no flexibility at one NFE, we refine this regime with alignment-aware temporal self-distillation using soft-DTW. Across Seed-TTS and LibriSpeech-PC, LFMD improves low-NFE synthesis over a matched integral-distillation baseline. On Seed-TTS, the refined student reaches 1.80% WER with 1-NFE, compared with 1.76% for its 32-NFE teacher.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang. 2026-09-28. Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech. https://arxiv.org/abs/2609.36324

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension

Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.

cs.SD↗

InterBias-SV: Compound Conditions in Speaker Verification

Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trained encoders, their mean contrasts are +0.0026, +0.0088, and +0.0024 in equal error rate (EER), with larger variation across settings. These descriptive averages do not establish equivalence to additivity: trial matching, checkpoint identity, and parts of the condition metadata remain unverified. We also examine two interpretation problems. Near-chance EER can make additive predictions difficult to interpret, but chance performance is not a hard EER ceiling, and correlation with the prediction does not identify a saturation mechanism. Ratios of demographic gaps are unstable when their clean reference is near zero; absolute gaps provide a more direct summary. The benchmark provides condition definitions, analysis scripts, and explicit requirements for interpretable compound-condition comparisons, while separating recomputable summaries from claims that require further experimental validation.

cs.SD↗

Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models

Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code

cs.SD↗