arXiv · 2609.21320
Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction
Abstract
Modern text and image representations are often matrix-valued, with rows corresponding to tokens, patches, or other local feature vectors. Predictive information is often sparse but sample-specific, making classical sparse regression methods with a common support poorly suited to this heterogeneity. This paper formalizes an individualized sparse regression framework for matrix-valued covariates in which each observation has its own rows of interest, while the associated regression effects are shared across the population. To estimate this model, we introduce a diagonalized attention mechanism that uses query--key scores to localize sample-specific signal rows and a value matrix for downstream regression. The proposed method has a parameter dimension independent of sample size and can identify rows of interest for new observations without their responses. We establish existence theorems showing that, under suitable score-separation and concentration conditions, single-head and multi-head diagonalized attention models recover the latent rows with high probability, yielding prediction risk bounds. Our theory therefore provides a statistical explanation of how attention-based scoring localizes sample-specific signals in heterogeneous matrix-valued data. Simulations demonstrate strong prediction and localization in regression and misspecified classification across varying sample sizes, dimensions, and signal cardinalities. Real sentiment analyses show improved classification accuracy and interpretable token selection.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Borui Peng, Liwei Lin, Feifei Wang, Long Feng. 2026-09-18. Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction. https://arxiv.org/abs/2609.21320
Cite the original work for its findings. Save a collection to share your selection of sources.