arXiv · 2411.11465
Re-examining learning linear functions in context
Abstract
In-context learning (ICL) has emerged as a powerful paradigm for easily adapting Large Language Models (LLMs) to various tasks. However, our understanding of how ICL works remains limited. We explore a simple model of ICL in a controlled setup with synthetic training data to investigate ICL of univariate linear functions. We experiment with a range of GPT-2-like transformer models trained from scratch. Our findings challenge the prevailing narrative that transformers adopt algorithmic approaches like linear regression to learn a linear function in-context. These models fail to generalize beyond their training distribution, highlighting fundamental limitations in their capacity to infer abstract task structures. Our experiments lead us to propose a mathematically precise hypothesis of what the model might be learning.
Explore related subjects
Keep this discovery
Omar Naim, Guilhem Fouilhé, Nicholas Asher. 2024-11-18. Re-examining learning linear functions in context. https://doi.org/10.1007/978-3-032-02813-6_8
Cite the original work for its findings. Save a collection to share your selection of sources.