arXiv · 2610.04887
TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models
Abstract
Audio large language models (ALLMs) can understand speech content, yet their ability to use speaker identity for verification remains limited. We propose TS-SP (Two-Stage Speaker Preservation), a parameter-efficient framework for learning speaker-preserving representations and making them accessible to an ALLM's language-model component. We instantiate and evaluate TS-SP on Qwen2.5-Omni-7B. First, we adapt the audio encoder with speaker identity supervision. We then freeze the adapted encoder and train the language model to compare speakers. Both stages use low-rank adaptation (LoRA), keeping the pretrained base weights fixed. On Vox1-O, TS-SP reduces the equal error rate (EER) from 7.01\% for the Paired Loss Adaptation Baseline to 4.37\%. EER remains within 4.31--4.79\% under unseen prompts. Cross-domain evaluation on CN-Celeb yields a similar EER to the baseline, but lower accuracy at the native decision threshold. These findings support two-stage adaptation for improving speaker verification on the evaluated backbone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junjie Li, Zheng Liang, Zhe Li, Tianchi Liu, Kong Aik Lee. 2026-10-04. TS-SP: Learning Speaker-Preserving Representations in Audio Large Language Models. https://arxiv.org/abs/2610.04887
Cite the original work for its findings. Save a collection to share your selection of sources.