arXiv · 2410.08255
Investigating Representation Universality: Case Study on Genealogical Representations
Abstract
Motivated by interpretability and reliability, we investigate whether large language models (LLMs) deploy universal geometric structures to encode discrete, graph-structured knowledge. To this end, we present two complementary experimental evidence that might support universality of graph representations. First, on an in-context genealogy Q&A task, we train a cone probe to isolate a tree-like subspace in residual stream activations and use activation patching to verify its causal effect in answering related questions. We validate our findings across five different models. Second, we conduct model stitching experiments across models of diverse architectures and parameter counts (OPT, Pythia, Mistral, and LLaMA, 410 million to 8 billion parameters), quantifying representational alignment via relative degradation in the next-token prediction loss. Generally, we conclude that the lack of ground truth representations of graphs makes it challenging to study how LLMs represent them. Ultimately, improving our understanding of LLM representations could facilitate the development of more interpretable, robust, and controllable AI systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
David D. Baek, Yuxiao Li, Max Tegmark. 2024-10-10. Investigating Representation Universality: Case Study on Genealogical Representations. https://arxiv.org/abs/2410.08255
Cite the original work for its findings. Save a collection to share your selection of sources.