arXiv · 2506.23845
Position: Use Sparse Autoencoders to Discover Unknowns
Abstract
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. We argue that even if SAEs may be less effective for \textit{acting on known concepts}, SAEs are especially powerful tools for \textit{discovering unknown concepts}. This distinction separates existing negative results from positive results, and suggests several classes of SAE applications. Specifically, we outline use cases for SAEs in (i) ML interpretability, explainability, fairness, auditing, and safety, and (ii) social and health sciences.
Explore related subjects
Keep this discovery
Kenny Peng, Rajiv Movva, Jon Kleinberg, Emma Pierson, Nikhil Garg. 2025-06-30. Position: Use Sparse Autoencoders to Discover Unknowns. https://arxiv.org/abs/2506.23845
Cite the original work for its findings. Save a collection to share your selection of sources.