arXiv · 2509.17196
Evolution of Concepts in Language Model Pre-Training
Abstract
Language models obtain extensive capabilities through pre-training. However, the pre-training process remains a black box. In this work, we track linear interpretable feature evolution across pre-training snapshots using a sparse dictionary learning method called crosscoders. We find that most features begin to form around a specific point, while more complex patterns emerge in later training stages. Feature attribution analyses reveal causal connections between feature evolution and downstream performance. Our feature-level observations are highly consistent with previous findings on Transformer's two-stage learning process, which we term a statistical learning phase and a feature learning phase. Our work opens up the possibility to track fine-grained representation progress during language model learning dynamics.
Explore related subjects
Keep this discovery
Xuyang Ge, Wentao Shu, Jiaxing Wu, Yunhua Zhou, Zhengfu He, Xipeng Qiu. 2025-09-21. Evolution of Concepts in Language Model Pre-Training. https://arxiv.org/abs/2509.17196
Cite the original work for its findings. Save a collection to share your selection of sources.