arXiv · 2601.22231
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
Abstract
This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectively function as geometric priors that shape the spatial structure of the representation. We introduce token-level diagnostics that measure how multi-view geometric consistency in ViT representation depends on consitent PEs. Through extensive experiments on 14 foundation ViT models, we reveal how PEs influence multi-view geometry and spatial reasoning. Our findings clarify the role of PEs as a causal mechanism that governs spatial structure in ViT representations. Our code is provided in https://github.com/shijianjian/vit-geometry-probes
Explore related subjects
Keep this discovery
Jian Shi, Michael Birsak, Wenqing Cui, Zhenyu Li, Peter Wonka. 2026-01-29. Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning. https://arxiv.org/abs/2601.22231
Cite the original work for its findings. Save a collection to share your selection of sources.