arXiv · 2609.23830
DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline
Abstract
Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski. 2026-09-20. DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline. https://arxiv.org/abs/2609.23830
Cite the original work for its findings. Save a collection to share your selection of sources.