arXiv · 2609.13198
SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation
Abstract
One of the pivotal recent challenges in neural network interpretability is polysemanticity, where a single neuron is activated by multiple, often unrelated concepts, hindering clear functional understanding. Although prior work has explored this phenomenon, existing approaches remain architecture-specific and depend on manual heuristics such as a fixed number of concept clusters ($K$), limiting their generality and scalability--especially for modern Transformer-based models. To address these limitations, we introduce SPICE (\textbf{S}imple \textbf{P}olysemantic Feature \textbf{I}nterpretation via \textbf{C}lustering-based \textbf{E}xplanation), a generalizable framework for analyzing polysemanticity in deep vision architectures. SPICE avoids architecture-dependent propagation rules, enabling the first systematic comparison of polysemanticity across both CNNs and Transformers, and automatically determines the number of concept clusters per neuron, eliminating reliance on a preset $K$ and supporting scalable analysis for large models. Using SPICE, we conduct a comprehensive investigation into how polysemanticity emerges, varies across depth and architecture, and forms through distinct computational pathways.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sehyun Lee, Dahee Kwon, Damin Lee, Jaesik Choi. 2026-08-14. SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation. https://arxiv.org/abs/2609.13198
Cite the original work for its findings. Save a collection to share your selection of sources.