arXiv · 2509.22363
Investigating Faithfulness in Large Audio Language Models
Abstract
Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks. While these models can generate Chain-of-Thought (CoT) explanations, the faithfulness of these reasoning chains remains unclear. In this work, we propose a systematic framework to evaluate CoT faithfulness in LALMs with respect to both the input audio and the final model prediction. We define three criteria for audio faithfulness: hallucination-free, holistic, and attentive listening. We also introduce a benchmark based on both audio and CoT interventions to assess faithfulness\footnote{The benchmarking interface and evaluation results are available at https://poonehmousavi.github.io/faithfulness/. Experiments on Audio Flamingo 3 and Qwen2.5-Omni suggest a potential multimodal disconnect: reasoning often aligns with the final prediction but is not always strongly grounded in the audio and can be vulnerable to hallucinations or adversarial perturbations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pooneh Mousavi, Lovenya Jain, Mirco Ravanelli, Cem Subakan. 2025-09-26. Investigating Faithfulness in Large Audio Language Models. https://arxiv.org/abs/2509.22363
Cite the original work for its findings. Save a collection to share your selection of sources.