arXiv · 2503.15421
Probing the topology of the space of tokens with structured prompts
Abstract
This article presents a general and flexible method for prompting a large language model (LLM) to reveal its (hidden) token input embedding up to homeomorphism. Moreover, this article provides strong theoretical justification -- a mathematical proof for generic LLMs -- for why this method should be expected to work. With this method in hand, we demonstrate its effectiveness by recovering the token subspace of Llemma-7B. The results of this paper apply not only to LLMs but also to general nonlinear autoregressive processes.
Explore related subjects
Keep this discovery
Michael Robinson, Sourya Dey, Taisa Kushner. 2025-03-19. Probing the topology of the space of tokens with structured prompts. https://arxiv.org/abs/2503.15421
Cite the original work for its findings. Save a collection to share your selection of sources.