arXiv · 2608.24007
Revenge of Monosemanticity: Neuron Specialization as a New Form of Feature Learning in MLPs
Abstract
Understanding how neural networks learn and organize features is central to understanding their behavior. Much existing theory of feature learning has focused on the emergence of a global low-dimensional representation. We show that this picture is incomplete. In regression problems with clustered data, we demonstrate that multilayer perceptrons (MLPs) naturally develop monosemantic specialized neurons: individual neurons become strongly aligned with a specific predictive feature relevant to a particular region of the input space. Rather than learning a single global low-dimensional representation, MLPs learn a collection of local low-dimensional representations. We show that this ability to specialize gives MLPs a provable data-efficiency advantage over feature-learning methods based on a global low-dimensional representation.
Explore related subjects
Keep this discovery
Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis, Mikhail Belkin. 2026-08-25. Revenge of Monosemanticity: Neuron Specialization as a New Form of Feature Learning in MLPs. https://arxiv.org/abs/2608.24007
Cite the original work for its findings. Save a collection to share your selection of sources.