Searcharxiv⌕ Search

arXiv · 2609.39651

Neural Audio Codec for Robust Audio Deepfake Detection

Abstract

Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8~pp compared with the original DAC at 0.5~kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jungwoo Kim, Joonyong Park, Junyoung Koh, Jong-Seok Lee. 2026-09-30. Neural Audio Codec for Robust Audio Deepfake Detection. https://arxiv.org/abs/2609.39651

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Dislikes may also arise from ambiguous, out-of-domain, or non-request turns, making it hard to isolate ASR effects. We release WASIL (it denotes connection or linking in Arabic): in-the-wild Arabic spoken interaction prompts with audio, ASR hypotheses, assistant responses, and explicit like/dislike feedback (8,529 turns; 14.2% dislikes), plus a 2,000-turn test set covering Modern Standard Arabic (MSA) and four major dialects with their labels. We provide low-cost gold transcripts via multi-ASR agreement-guided post-editing and annotate answerability (answerable, ambiguous/needs-clarification, unsupported, not-a-request/noise) to separate intrinsic unanswerability from ASR-induced degradation. Finally, we describe scalable reference-free evaluation of responses from ASR vs. gold transcripts using multi-judge LLM scoring.

cs.SD↗

Trigger Sound Suppression for Misophonia

Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.

cs.SD↗

Bad: Taming the Bioacoustic Data Deluge with a Bat Activity Detector

Passive Acoustic Monitoring of bats generates massive ultrasonic datasets (>27 GB/night per node), straining edge storage and battery life. Legacy triggers fail against acoustic confusers, while deep models exceed microcontroller limits. We present a hardware-aware Bat Activity Detector (BAD) specifically designed to discriminate bat calls from hard biological and environmental confusers across variable sampling rates (192-384 kHz). Tailored for the Silicon Labs EFM32PG26 (MVP) in 8-bit integer precision, our model achieves 100 percent hardware offload across all 14 layers (17.2 KB Flash, 73.1 KB RAM). End-to-end preprocessing (74.00 ms for 76 frames) and inference (30.00 ms) of 100 ms clips at 192 kHz require 104.00 ms per clip. On spatially out-of-domain recordings under a realistic low-prevalence regime (r_pos = 0.05), BAD achieves an AUC-ROC of 0.9748 and suppresses 99.4% of non-target noise frames while retaining 65.3% of bat calls - delivering a >33x precision gain over classical Goertzel baselines.

cs.SD↗