arXiv · 2609.37789
Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance
Abstract
Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fabian A. Mikulasch, Friedemann Zenke. 2026-09-29. Predictive Self-Supervised Learning Provably Identifies Stochastic Signals under Nuisance. https://arxiv.org/abs/2609.37789
Cite the original work for its findings. Save a collection to share your selection of sources.