SearcharxivSearch

arXiv subjects

Mingzhi Song

Publications and source records attributed to Mingzhi Song.

5 recordsLinked to original sources

Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample

We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO fluctuation, and deletion-to-full risk transfer. On each fixed finite horizon, bounded centered training-loss gradients, a one-sided Hessian lower bound, locally Lipschitz Hessians, and a strict tube-closure condition yield an explicit $(n-1)^{-2}$ bound for the deletion-response error. Bounded evaluation-loss gradients transfer the deletion-response bound to the score without requiring the Hessian to be invertible. Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve. For bounded smooth two-layer mean-field networks training both layers, the score-error bound is uniform in width.

stat.ML

Convergence of difference inclusions: a diameter criterion and step-size conditions

We study bounded realizations of discrete difference inclusions with set-valued increments and additive errors. Our results have two parts. First, we give a general convergence criterion in which changes in a convergent scalar quantity control the diameters of local sequence segments near an accumulation point. We also develop a stratified descent framework for verifying this control. For convergent realizations, the framework yields a stationarity condition for the limit from outer limits of the scaled update map. When applied to first-order methods for minimizing locally Lipschitz objectives definable in polynomially bounded o-minimal structures, the framework yields convergence to critical points for bounded sequences generated by the inexact subgradient, momentum, and stochastic subgradient methods with step sizes of order $k^{-1}$, under the corresponding error and noise conditions. Second, we study first-order methods with polynomial step sizes of order $k^{-a}$. At the square-summability boundary \(a=1/2\), we construct a locally Lipschitz semialgebraic objective for which the exact subgradient method generates a bounded, nonconvergent sequence whose accumulation points satisfy our active-geometry assumptions. Under these assumptions, bounded sequences generated by the momentum method converge for \(1/2<a\leq1\), while bounded sequences generated by the stochastic subgradient method converge almost surely for \(2/3<a\leq1\). The momentum range is sharp for this method class, while the stochastic range matches recent full-sequence convergence results for smooth nonconvex objectives.

math.OC

Cross-Calibrated Confidence Fields for Local Risk Updates

How can training data be used to compare local updates to the current model, choose an update, and retain valid bounds for the selected update's population-risk change? We construct lower and upper confidence fields that jointly cover the population-risk change of every update in a possibly continuous local update space. The fields can therefore be used both to compare the updates and to select an update; a negative upper endpoint certifies improvement over the reference model. For linear risk changes in possibly infinite-dimensional feature spaces, cross-calibration uses the discrepancy between two balanced folds to calibrate the full-sample estimation error. Under covariance-aligned sub-Gaussian tails, covariance-estimation stability, and sufficient sample size, the cross-calibrated field has finite-sample simultaneous coverage. The confidence field's directional widths are governed by a population ridge effective dimension rather than the ambient feature dimension. For losses formed locally by continuous selection among finitely many smooth branches, separate uniform bounds for the linear Taylor field, Taylor remainder, and branch-interface discrepancy extend the field to every local update.

stat.ML

On the diameter of subgradient sequences in o-minimal structures

We study subgradient sequences of locally Lipschitz functions definable in a polynomially bounded o-minimal structure. We show that the diameter of any subgradient sequence is related to the variation in function values, with error terms dominated by a double summation of step sizes. Consequently, we prove that bounded subgradient sequences converge if the step sizes are of order $1/k$. The proof uses Lipschitz $L$-regular stratifications in o-minimal structures to analyze subgradient sequences via their projections onto different strata.

math.OC

Feature Selection in High-dimensional Spaces Using Graph-Based Methods

High-dimensional feature selection is a central problem in a variety of application domains such as machine learning, image analysis, and genomics. In this paper, we propose graph-based tests as a useful basis for feature selection. We describe an algorithm for selecting informative features in high-dimensional data, where each observation comes from one of $K$ different distributions. Our algorithm can be applied in a completely nonparametric setup without any distributional assumptions on the data, and it aims at outputting those features in the data, that contribute the most to the overall distributional variation. At the heart of our method is the recursive application of distribution-free graph-based tests on subsets of the feature set, located at different depths of a hierarchical clustering tree constructed from the data. Our algorithm recovers all truly contributing features with high probability, while ensuring optimal control on false-discovery. We show the superior performance of our method over other existing ones through synthetic data, and demonstrate the utility of this method on several real-life datasets from the domains of climate change and biology, wherein our algorithm is not only able to detect known features expected to be associated with the underlying process, but also discovers novel targets that can be subsequently studied.

stat.ME