Searcharxiv⌕ Search

arXiv · 2610.07216

Exposing and Mitigating Neural Codec Vulnerabilities in Audio Deepfake Detection

Abstract

Existing audio deepfake detection (ADD) datasets and detectors are primarily built for vocoder-based synthesis, evaluated against traditional post-hoc perturbations such as MP3/AAC compression or additive noise, applied independently of generation. However, recent speech synthesizers, particularly ALM-based systems, use neural audio codecs both for compression and as the resynthesis reconstructing waveforms from generated tokens, producing artifacts distinct from post-hoc compression. Neural codecs thus play a dual role: some are designed for pure compression under low-bandwidth communication, while others serve as resynthesis components. Despite this dual role, robustness to codec-based compression, unlike post-hoc compression, remains largely unexplored. We expose this gap, showing that state-of-the-art (SOTA) ADD models degrade drastically on codec-compressed speech; in particular, systems trained on Codec Resynthesized data as a proxy for codec-based generation prove most vulnerable, with legitimately compressed bona fide speech often misclassified as fake. To investigate this, we construct the Audio Neural Codec-Spoof dataset by applying seven neural codec algorithms to existing ADD benchmarks, isolating codec-induced resynthesis artifacts as a controlled proxy for codec-based generation. As baseline mitigation, we propose PCL-NET (Pairwise Consistency Learned Network), fine-tuning a pretrained XLS-R (300M) encoder with a pairwise consistency objective that minimizes the representation distance between an utterance's uncompressed and codec-compressed versions, disentangling codec artifacts from the real-versus-fake decision. As a result, PCL-NET reduces average EER under neural codec compression from 28.67% to 12.77%, while preserving competitive CoSG-based deepfake detection performance. We will also make the dataset publicly available on Hugging Face upon acceptance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Abdullah, Awais Khan, Khalid Mahmood Malik. 2026-10-05. Exposing and Mitigating Neural Codec Vulnerabilities in Audio Deepfake Detection. https://arxiv.org/abs/2610.07216

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Auditing generative audio calls for known-task audio-llm evaluation

Speech and audio LLMs are evaluated by comparing waveform predictions with predictions from an automatic speech recognition (ASR) transcript. For fixed closed-set tasks, this conflates acoustic evidence with the need to invoke a generative audio model. We estimate incremental call value with matched selectors sharing pre-call evidence. Each policy may retain the transcript label, use a local encoder, or invoke a generative model; matched control removes generative actions but preserves pre-call evidence and development selection. On VocalSound, transcript-only accuracy is 0.296, while supervised CLAP and WavLM controls reach 0.850 and 0.854 without calls. Full selector reaches 0.925 at 12.5% calls versus 0.921 for matched No-call selector (difference 0.004; 95% CI [-0.025, 0.033]). Thus, results do not show a call gain after transcript and encoder evidence are available. Relevant quantity is incremental accuracy from allowing calls, not the waveform-transcript gap.

cs.SD↗

Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

Audiobook narration, conversational agents, and audiovisual dubbing require speech that conveys changing emotions and adapts its pacing within a single utterance. But most existing TTS systems typically rely on utterance-level style conditioning, making such fine-grained control difficult to achieve. In light of this, and inspired by the success of post-training in large language models, we propose a unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration. Supervised fine-tuning establishes instruction-conditioned speech generation, while reinforcement learning with group relative policy optimization refines control accuracy using emotion and duration rewards alongside content and speaker preservation objectives. By reusing the pretrained architecture, our approach avoids additional inference-time control modules. Experiments demonstrate significantly improved fine-grained controllability while maintaining speech intelligibility and speaker identity, highlighting post-training as a practical approach to extending existing speech synthesis models.

cs.SD↗

Do Language Models Need Music Supervision? Verifiable Rewards for Multi-Constraint Symbolic Music Generation

Language models now generate symbolic music from text, and research has focused on musicality. However, many applications require a score that meets explicit constraints, which models struggle to satisfy jointly: on MusicConstraintBench, our benchmark of 2,180 items over eight families of programmatically verifiable constraints, Llama-3.1-70B satisfies 0.630 of single-constraint items but only 0.044 of four-constraint ones. As a remedy, we introduce MusicRLVR, which trains a language model with group relative policy optimisation (GRPO) on verifier rewards alone, needing no human annotation, reward model or music-domain supervised fine-tuning. MusicRLVR incorporates (1) a hard validation gate that rejects malformed scores, (2) graded per-family credit that, unlike a binary reward, separates partially correct outputs, and (3) an all-satisfied bonus for meeting every constraint at once. Extensive experiments show that, in under four hours of training, MusicRLVR raises Qwen3-4B-Instruct-2507 from 0.160 to 0.797 on mixed constraints, outperforming Llama-3.1-70B, and generalises to unseen property combinations, out-of-range parameters and more constraints than any training prompt. The recipe transfers to Qwen3-8B, and neither trained model loses significant accuracy on general benchmarks.

cs.SD↗