arXiv · 2610.00706
AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
Abstract
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan. 2026-09-30. AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models. https://arxiv.org/abs/2610.00706
Cite the original work for its findings. Save a collection to share your selection of sources.