arXiv · 2601.04098
Layer-wise Positional Bias in Short-Context Language Modeling
Abstract
Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as positional bias. Prior work characterizes this bias in model behavior through performance drops in long-context tasks or in model architecture through attention-based analyses. However, it remains unmeasured how input positions actually drive predictions layer by layer. We introduce a layer conductance framework within a sliding-window design, applied to short-context next-word prediction to isolate model-internal behavior from task and context-window pressure. The resulting layer-wise positional importance profiles are stable across diverse texts and lexical scrambling, confirming they reflect model-internal structure. Characterizing how these profiles evolve across depth, we find recency bias increases monotonically while primacy bias is subtle and diminishes. We also find that this positional bias is not uniform across word types: function words exhibit higher recency bias while content words show higher primacy bias.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Maryam Rahimi, Mahdi Nouri, Yadollah Yaghoobzadeh. 2026-01-07. Layer-wise Positional Bias in Short-Context Language Modeling. https://arxiv.org/abs/2601.04098
Cite the original work for its findings. Save a collection to share your selection of sources.