arXiv · 2607.03418
DETECT-3B-Omni is Agnostic of Content and Demographics
Abstract
A trustworthy and GDPR-compliant deepfake audio detector must base its decisions on acoustic artifacts, not on what is being said or who is speaking. We present a large-scale study of semantic independence for Resemble AI's detector, DETECT-3B-Omni. Using 10,240 audio samples from diverse US English speakers across 30 states, generated through 8 different AI voice-cloning systems, we test whether detection accuracy depends on spoken content (benign versus malicious), speaker gender, speaker age, or speaker region. Using equivalence testing, our results show that the accuracy difference between any two of these groups is at most 2 percentage points, at 99% confidence. The detector therefore identifies AI-generated audio with equivalent accuracy regardless of what the audio says or who the speaker is.
Explore related subjects
Keep this discovery
Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Dominik Schnieders, Zohaib Ahmed. 2026-07-03. DETECT-3B-Omni is Agnostic of Content and Demographics. https://arxiv.org/abs/2607.03418
Cite the original work for its findings. Save a collection to share your selection of sources.