arXiv · 2608.17585
The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
Abstract
Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Anton Firc, Kamil Malinka, Vojtěch Staněk, Miroslav Hlaváček, Marek Bartoň. 2026-08-18. The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report. https://arxiv.org/abs/2608.17585
Cite the original work for its findings. Save a collection to share your selection of sources.