Searcharxiv⌕ Search

arXiv · 2610.01695

Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs

Abstract

Encoder-based speech large language models (Speech-LLMs) commonly employ pretrained speech encoders that prioritize linguistic content but may discard fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding. Encoder-free Speech-LLMs instead map Mel-spectrogram features directly into the LLM input space through lightweight embedding layers, enabling the LLM to learn from low-level acoustic features. However, systematic pretraining strategies for encoder-free Speech-LLMs remain underexplored, limiting their ability to compensate for the absence of large-scale pretrained speech encoders. We propose metadata-supervised pretraining (MSP), which leverages speech attributes such as speaker identity and emotion to develop speaker-discriminative and paralinguistic capabilities. We further introduce speaker-aware utterance composition (SAUC) to strengthen speaker discrimination and apply random span masking to regularize pretraining. We primarily evaluate our approach on joint ASR and speaker diarization in multi-speaker conversations, complemented by experiments on paralinguistic speech-understanding tasks. Under matched training-data conditions, our encoder-free model outperforms its randomly initialized encoder-based counterpart. With limited metadata-annotated data, it is competitive with models using speech encoders pretrained on substantially larger corpora, outperforming them in several settings. These results demonstrate the potential of encoder-free architectures for building native multimodal LLMs that acquire diverse speech capabilities.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu Li. 2026-10-01. Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs. https://arxiv.org/abs/2610.01695

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

ParaCalib: Semantically Calibrated Paralinguistic Modeling for Depression Detection

Vocal behavior provides important signals for speech-based depression detection, but its interpretation often depends on what is being said and how it functions in context. However, existing methods typically treat acoustic cues as context-independent markers, making it difficult to distinguish vocal form from its context-dependent communicative function. We propose ParaCalib, a framework that semantically calibrates paralinguistic behavior by interpreting vocal patterns relative to utterance-level semantic context and inferred communicative function and representing them in a comparable state space. Concretely, ParaCalib uses an Audio-Language Model (ALM) to generate contextualized vocal descriptions and an LLM-based Paralinguistic State Extractor (PSE) to map these descriptions into a structured Semantically Calibrated Paralinguistic (SC-Para) representation. ParaCalib achieves the highest mean Macro-F1 among the evaluated methods, reaching 71.9\% on DAIC-WOZ and 90.5\% on MODMA. Our controlled analysis provides direct evidence of semantic calibration at the PSE stage: under a fixed caption-level acoustic description, varying the accompanying semantic context changes the inferred depression-related paralinguistic evidence. Our exploratory analysis further identifies recurring configurations of paralinguistic states associated with depression labels rather than a single uniformly dominant state.

eess.AS↗

FedCFM: Federated Continual Domain Generalization for Fake Speech Detection via Conditional Flow Matching

The generalization ability of Fake Speech Detection (FSD) models is crucial for real-world deployment. Existing multi-dataset co-training methods rely on fixed training sets and cannot adapt to emerging spoofing types. Although con-tinual learning has been explored, many approaches overlook limited data storage at individual devices, thereby restricting practical applicability. To address this, we propose FedCFM, a Federated continual domain generalization framework via Conditional Flow Matching (CFM) for collaboration without sharing raw speech data across distributed clients facing diverse and evolving spoofing attacks. Each client trains a CFM-based generator to model spoof-type-specific embedding distributions, and cross-client generator exchange enables synthesis of unseen spoof-type embeddings for continual classifier updating through generative replay and knowledge distillation. With the same training datasets, FedCFM achieves lower EER than the eval-uated centralized and federated domain generalization baselines, demonstrating strong cross-domain generalization. Code will be released on https://github.com/jspycpp/FedCFM.

eess.AS↗

A Federated Deepfake Speech Detection Method Based on Layer-Wise Center-Guided Weighting Aggregation

The advancement of deep learning-based speech synthesis has significantly increased the diversity of deepfake speech, posing threats to voice authentication. While centralized training is effective for deepfake speech detection (DSD), it requires considerable computational resources and raises privacy concerns. To address these issues, we propose a Federated DSD (FedDSD) method that enables collaborative model training across decentralized speech datasets without sharing raw audio. Specifically, each client trains a local model using the FedProx algorithm to mitigate the effects of data heterogeneity and uploads model parameters to a central server. To improve global model aggregation, we further propose a layer-wise center-guided weighting aggregation (L-CGWA) strategy that adjusts each client's contribution per layer based on its distance to a reference center, capturing inter-client and inter-layer discrepancies and enhancing the robustness of model aggregation. Experimental results demonstrate that models trained under the proposed FedDSD method achieve equal error rates (EERs) comparable to those obtained via centralized co-training, while significantly out-performing models trained on individual corpora. Furthermore, the proposed FedDSD method demonstrates robust generalization capabilities across diverse cross-domain datasets.

eess.AS↗