SearcharxivSearch

arXiv subjects

Jeremy Stoddard

Publications and source records attributed to Jeremy Stoddard.

3 recordsLinked to original sources

AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently activate in response to several unrelated concepts. We introduce the first mechanistic interpretability framework for AudioLLMs, leveraging sparse autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features. Our pipeline identifies representative audio clips, assigns meaningful names via automated captioning, and validates concepts through human evaluation and steering. Experiments show that AudioLLMs encode structured and interpretable features, enhancing transparency and control. This work provides a foundation for trustworthy deployment in high-stakes domains and enables future extensions to larger models, multilingual audio, and more fine-grained paralinguistic features. Project URL: https://townim-faisal.github.io/AutoInterpret-AudioLLM/

cs.SD

Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features

Estimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequency bands using first-order Ambisonics (FOA) speech recordings as inputs. The proposed framework utilizes a novel feature named Spectro-Spatial Covariance Vector (SSCV), efficiently representing temporal, spectral as well as spatial information of the FOA signal. Our models significantly outperform existing single-channel methods with only spectral information, reducing estimation errors by more than half for all three acoustic parameters. Additionally, we introduce FOA-Conv3D, a novel back-end network for effectively utilising the SSCV feature with a 3D convolutional encoder. FOA-Conv3D outperforms the convolutional neural network (CNN) and recurrent convolutional neural network (CRNN) backends, achieving lower estimation errors and accounting for a higher proportion of variance (PoV) for all 3 acoustic parameters.

eess.AS

Gaussian Process Regression for Generalized Frequency Response Function Estimation

Kernel-based modeling of dynamic systems has garnered a significant amount of attention in the system identification literature since its introduction to the field. While the method was originally applied to linear impulse response estimation in the time domain, the concepts have since been extended to the frequency domain for estimation of frequency response functions (FRFs), as well as to the estimation of the Volterra series in time domain. In the latter case, smoothness and exponential decay was imposed along the hypersurfaces of the multidimensional impulse responses, allowing lower variance estimates than could be obtained in a simple least squares framework. The Volterra series can also be expressed in a frequency domain context, however there are several competing representations which all possess some unique advantages. Perhaps the most natural representation is the generalized frequency response function (GFRF), which is defined as the multidimensional Fourier transform of the corresponding Volterra kernel in the time-domain series. The representation leads to a series of frequency domain functions with increasing dimension.

eess.SY