arXiv · 2609.19067
What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting
Abstract
The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers' training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers' training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yixuan Xiao, Ngoc Thang Vu. 2026-07-22. What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting. https://doi.org/10.1109/icassp49660.2025.10888114
Cite the original work for its findings. Save a collection to share your selection of sources.