arXiv · 2606.21979
Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings
Abstract
Understanding speaker attributes is crucial for voice-related applications, yet conventional approaches rely on fixed categorical labels, lacking semantic richness and zero-shot generalizability. We propose a novel framework for open-set speaker attribute prediction leveraging Large Language Model (LLM) embeddings to represent attributes in a continuous semantic space. To bridge the cross-modal gap, we introduce a keyword-appending strategy that structures broad semantic representations into a compact, discriminative manifold. Furthermore, we employ a top-k negative loss to establish robust decision boundaries in crowded semantic regions. Experimental results on LibriTTS-P demonstrate that our method outperforms closed-set benchmarks and generalizes effectively to unseen synonyms. Geometric analysis suggests that our strategies regularize the embedding manifold, balancing semantic cohesion with predictive clarity.
Explore related subjects
Keep this discovery
Byoungjun So, Jaejun Lee, Kyogu Lee. 2026-06-20. Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings. https://arxiv.org/abs/2606.21979
Cite the original work for its findings. Save a collection to share your selection of sources.