SearcharxivSearch

arXiv subjects

Shuofeng Zhang

Publications and source records attributed to Shuofeng Zhang.

4 recordsLinked to original sources

Position: Many generalization measures for deep learning are fragile

In this position paper, we argue that many post-mortem generalization measures -- those computed on trained networks -- are \textbf{fragile}: small training modifications that barely affect the performance of the underlying deep neural network can substantially change a measure's value, trend, or scaling behavior. For example, minor hyperparameter changes, such as learning rate adjustments or switching between SGD variants, can reverse the slope of a learning curve in widely used generalization measures such as the path norm. We also identify subtler forms of fragility. For instance, the PAC-Bayes origin measure is regarded as one of the most reliable, and is indeed less sensitive to hyperparameter tweaks than many other measures. However, it completely fails to capture differences in data complexity across learning curves. This data fragility contrasts with the function-based marginal-likelihood PAC-Bayes bound, which does capture differences in data-complexity, including scaling behavior, in learning curves, but which is not a post-mortem measure. Beyond demonstrating that many post-mortem bounds are fragile, this position paper also argues that developers of new measures should explicitly audit them for fragility.

cs.LG

Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias

For overparameterized linear regression with isotropic Gaussian design and minimum-$\ell_p$ interpolator $p\in(1,2]$, we give a unified, high-probability characterization for the scaling of the family of parameter norms $ \\{ \lVert \widehat{w_p} \rVert_r \\}_{r \in [1,p]} $ with sample size. We solve this basic, but unresolved question through a simple dual-ray analysis, which reveals a competition between a signal *spike* and a *bulk* of null coordinates in $X^\top Y$, yielding closed-form predictions for (i) a data-dependent transition $n_\star$ (the "elbow"), and (ii) a universal threshold $r_\star=2(p-1)$ that separates $\lVert \widehat{w_p} \rVert_r$'s which plateau from those that continue to grow with an explicit exponent. This unified solution resolves the scaling of *all* $\ell_r$ norms within the family $r\in [1,p]$ under $\ell_p$-biased interpolation, and explains in one picture which norms saturate and which increase as $n$ grows. We then study diagonal linear networks (DLNs) trained by gradient descent. By calibrating the initialization scale $\alpha$ to an effective $p_{\mathrm{eff}}(\alpha)$ via the DLN separable potential, we show empirically that DLNs inherit the same elbow/threshold laws, providing a predictive bridge between explicit and implicit bias. Given that many generalization proxies depend on $\lVert \widehat {w_p} \rVert_r$, our results suggest that their predictive power will depend sensitively on which $l_r$ norm is used.

cs.LG

Why flatness does and does not correlate with generalization for deep neural networks

The intuition that local flatness of the loss landscape is correlated with better generalization for deep neural networks (DNNs) has been explored for decades, spawning many different flatness measures. Recently, this link with generalization has been called into question by a demonstration that many measures of flatness are vulnerable to parameter re-scaling which arbitrarily changes their value without changing neural network outputs. Here we show that, in addition, some popular variants of SGD such as Adam and Entropy-SGD, can also break the flatness-generalization correlation. As an alternative to flatness measures, we use a function based picture and propose using the log of Bayesian prior upon initialization, $\log P(f)$, as a predictor of the generalization when a DNN converges on function $f$ after training to zero error. The prior is directly proportional to the Bayesian posterior for functions that give zero error on a test set. For the case of image classification, we show that $\log P(f)$ is a significantly more robust predictor of generalization than flatness measures are. Whilst local flatness measures fail under parameter re-scaling, the prior/posterior, which is global quantity, remains invariant under re-scaling. Moreover, the correlation with generalization as a function of data complexity remains good for different variants of SGD.

cs.LG

First-principles study of the layered thermoelectric material TiNBr

Layer-structured materials are often considered to be good candidates for thermoelectric materials, because they tend to exhibit intrinsically low thermal conductivity as a result of atomic interlayer interactions. The electrical properties of layer-structured materials can be easily tuned using various methods, such as band modification and intercalation. We report TiNBr, as a member of the layer-structured metal nitride halide system MNX (M = Ti, Zr, Hf; X = Cl, Br, I), and it exhibits an ultrahigh Seebeck coefficient of 2215 $μV/K$ at 300K. The value of the dimensionless figure of merit, ZT, along A axis can be as high as 0.661 at 800K, corresponding to a lattice thermal conductivity as low as 1.34 W/(m K). The low ${κ_l}$ of TiNBr is associated with a collectively low phonon group velocity ($2.05\times 10^3 $ m/s on average) and large phonon anharmonicity that can be quantified using the Grüneisen parameter and three-phonon processes. Animation of the atomic motion in highly anharmonic modes mainly involves the motion of N atoms, and the charge density difference reveals that the N atoms become polarized with the merging of anharmonicity. Moreover, the fitting procedure of the energy-displacement curve verifies that in addition to the three-phonon processes, the fourth-order anharmonic effect is also important in the integral anharmonicity of TiNBr. Our work is the first study of the thermoelectric properties of TiNBr and may help establish a connection between the low lattice thermal conductivity and the behavior of phonon vibrational modes.

cond-mat.mtrl-sci