arXiv · 2605.12225
On the Interpretability of Whisper Encodings Using Sparse Autoencoders
Abstract
While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understanding text-based transformer models, leaving ASR systems largely unexplored. In order to address this gap, we examine the internal representations of Whisper's encoder using a sparse autoencoder. We find diverse monosemantic features across linguistic and non-linguistic boundaries, spanning a hierarchy from phonetic to semantic representations, and conduct a causal feature-steering campaign across this hierarchy, including cross-lingual steering. We further find that steering is more reliable for higher-level features than lower-level ones, an asymmetry that may reflect redundant encoding of lower-level information. Altogether, this work demonstrates that Whisper's encoder represents a surprisingly rich hierarchy of linguistic information that extends well beyond what is strictly necessary for transcription.
Explore related subjects
Keep this discovery
Dan Pluth, Zachary Nicholas Houghton, Yu Zhou, Vijay K. Gurbani. 2026-05-12. On the Interpretability of Whisper Encodings Using Sparse Autoencoders. https://arxiv.org/abs/2605.12225
Cite the original work for its findings. Save a collection to share your selection of sources.