arXiv · 2608.01383
An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
Abstract
Masked prediction learns by inferring missing variables from visible context. When does optimizing this conditional task recover the true joint data distribution? We study this question using an $\varepsilon$-identifiability modulus, which measures the worst-case joint-distribution error permitted by excess risk at most $\varepsilon$. For distributions with separated global modes, schedules retaining large visible contexts can permit substantial mode-weight errors at exponentially small excess risk. An exact information decomposition explains why: for a fixed mask, the loss penalizes only the mode-weight mismatch that remains unresolved by the visible context. For small mode-weight perturbations, the objective's sensitivity is proportional to residual mode uncertainty averaged over masks. Under joint masked-block log loss, low-visibility masks that retain mode uncertainty restore this sensitivity, while positive full-mask probability bounds joint-distribution error in terms of excess risk. We empirically validate these predictions through exact calculations and controlled stochastic optimization.
Explore related subjects
Keep this discovery
Yichao Cai, Javen Qinfeng Shi. 2026-08-02. An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules. https://arxiv.org/abs/2608.01383
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.