arXiv · 2503.02110
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
Abstract
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, based on an arbitrary prior or description language. Grunwald and Langford [2004] previously established the lack of asymptotic consistency, from an agnostic PAC (frequentist worst case) perspective, of the MDL rule with a penalty parameter of $\lambda=1$, suggesting that it underegularizes. Driven by interest in understanding how benign or catastrophic under-regularization and overfitting might be, we obtain a precise quantitative description of the worst case limiting error as a function of the regularization parameter $\lambda$ and noise level (or approximation error), significantly tightening the analysis of Grunwald and Langford for $\lambda=1$ and extending it to all other choices of $\lambda$.
Explore related subjects
Keep this discovery
Xiaohan Zhu, Nathan Srebro. 2025-03-03. Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification. https://arxiv.org/abs/2503.02110
Cite the original work for its findings. Save a collection to share your selection of sources.