arXiv · 2602.19481
A Selection Premium Decomposition for the Expected Maximum of Random Walks
Abstract
When $K$ models are evaluated on the same validation set of size $n$, the selected winner's apparent performance is biased upward. Suppose $K$ models are evaluated on a shared sequence of i.i.d. observations $X_1,\dots, X_n$, where model $k$ achieves response $f_k(X_i)$ with mean $\mu_k = \mathbb E[f_k(X)]$. Writing $Y_{i,k} = f_k(X_i)-\mu_k$ for the centered increment and $S_{n,k} = \sum_{i=1}^n Y_{i,k}$ for the centered cumulative score, the expected maximum satisfies $0\le\mathbb E\bigl[\max_k S_{n,k}\bigr] = \sum_{i=1}^n \mathbb E\bigl[\varphi_K(S_{i-1})\bigr]$ where $\varphi_K(u) = \mathbb{E}\bigl[\max_k(u_k + Y_k)\bigr] - \max_k u_k$, $u\in \mathbb R^K$, is the selection premium function. This formula corresponds to the null hypothesis case (all models are equal in the sense that they have the same mean), which clarifies that the bias arises from selection. While this decomposition follows from elementary conditioning and telescoping, we develop the analytical consequences in five directions. (i) structural properties of $\varphi_K$; (ii) extension to stopping times, recovering Wald's equation at $K=1$; (iii) a winner's curse decomposition for heterogeneous means; (iv) a universal bias concentration law showing that the first $\alpha$-fraction of observations generates a $\sqrt\alpha$-fraction of total bias.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Victor H. de la Pena, Fangyuan Lin, Victor K. de la Pena. 2026-02-23. A Selection Premium Decomposition for the Expected Maximum of Random Walks. https://arxiv.org/abs/2602.19481
Cite the original work for its findings. Save a collection to share your selection of sources.