arXiv · 2509.15082
From Who Said What to Who They Are: Modular Training-free Identity-Aware LLM Refinement of Speaker Diarization
Abstract
Speaker diarization (SD) remains challenging in real-world scenarios due to dynamic environments and unknown speaker numbers. SD is rarely used alone and is typically paired with automatic speech recognition (ASR). However, existing non-modular SD+ASR frameworks lack flexibility and do not provide true speaker identities. We propose a training-free modular pipeline combining off-the-shelf SD, ASR, and a large language model (LLM) to determine who spoke, what was said, and who they are. Using structured LLM prompting on reconciled SD and ASR outputs, our method leverages semantic continuity in conversational context to refine low-confidence speaker labels and assigns role identities while correcting split speakers. On a real-world patient-clinician dataset, our approach achieves a 29.7% relative error reduction over baseline reconciled SD and ASR. It enhances diarization performance without additional training and delivers a complete pipeline for SD, ASR, and speaker identity detection in practical applications.
Explore related subjects
Keep this discovery
Yu-Wen Chen, William Ho, Maxim Topaz, Julia Hirschberg, Zoran Kostic. 2025-09-18. From Who Said What to Who They Are: Modular Training-free Identity-Aware LLM Refinement of Speaker Diarization. https://arxiv.org/abs/2509.15082
Cite the original work for its findings. Save a collection to share your selection of sources.