SearcharxivSearch

arXiv subjects

P. Baglioni

Publications and source records attributed to P. Baglioni.

3 recordsLinked to original sources

Large fluctuations in NSPT computations: a lesson from $O(N)$ non-linear sigma models

In the last three decades, Numerical Stochastic Perturbation Theory (NSPT) has proven to be an excellent tool for calculating perturbative expansions in theories such as Lattice QCD, for which standard, diagrammatic perturbation theory is known to be cumbersome. Despite the significant success of this stochastic method and the improvements made in recent years, NSPT apparently cannot be successfully implemented in low-dimensional models due to the emergence of huge statistical fluctuations: as the perturbative order gets higher, the signal to noise ratio is simply not good enough. This does not come as a surprise, but on very general grounds, one would expect that the larger the number of degrees of freedom, the less severe the fluctuations will be. By simulating $2D$ $O(N)$ non-linear sigma models for different values of $N$, we show that indeed the fluctuations are tamed in the large $N$ limit, meeting our expectations: for a large number of internal degrees of freedom (i.e. for large enough $N$), NSPT perturbative computation can be pushed to large perturbative orders $n$. By re-expressing our perturbative expansions as power series in the $gN$ ('t Hooft) coupling, we show some evidence that at any given order $n$ there is a tendency to gaussianity for the stochastic process distributions at large $N$. By summing our series, we can verify leading order results for the energy and its (field theoretic) variance in the large $N$ limit. We finally establish general relationships between the various perturbative orders in the expansion of the (field theoretic) variance of a given observable and combinations of variances and covariances of given orders NSPT stochastic processes. Having established all this, we conclude discussing interesting applications of NSPT computations in the context of theories similar to $O(N)$ (i.e. $CP(N-1)$ models).

hep-lat

Kernel shape renormalization explains output-output correlations in finite Bayesian one-hidden-layer networks

Finite-width one hidden layer networks with multiple neurons in the readout layer display non-trivial output-output correlations that vanish in the lazy-training infinite-width limit. In this manuscript we leverage recent progress in the proportional limit of Bayesian deep learning (that is the limit where the size of the training set $P$ and the width of the hidden layers $N$ are taken to infinity keeping their ratio $α= P/N$ finite) to rationalize this empirical evidence. In particular, we show that output-output correlations in finite fully-connected networks are taken into account by a kernel shape renormalization of the infinite-width NNGP kernel, which naturally arises in the proportional limit. We perform accurate numerical experiments both to assess the predictive power of the Bayesian framework in terms of generalization, and to quantify output-output correlations in finite-width networks. By quantitatively matching our predictions with the observed correlations, we provide additional evidence that kernel shape renormalization is instrumental to explain the phenomenology observed in finite Bayesian one hidden layer networks.

cond-mat.dis-nn

Predictive power of a Bayesian effective action for fully-connected one hidden layer neural networks in the proportional limit

We perform accurate numerical experiments with fully-connected (FC) one-hidden layer neural networks trained with a discretized Langevin dynamics on the MNIST and CIFAR10 datasets. Our goal is to empirically determine the regimes of validity of a recently-derived Bayesian effective action for shallow architectures in the proportional limit. We explore the predictive power of the theory as a function of the parameters (the temperature $T$, the magnitude of the Gaussian priors $λ_1$, $λ_0$, the size of the hidden layer $N_1$ and the size of the training set $P$) by comparing the experimental and predicted generalization error. The very good agreement between the effective theory and the experiments represents an indication that global rescaling of the infinite-width kernel is a main physical mechanism for kernel renormalization in FC Bayesian standard-scaled shallow networks.

cond-mat.dis-nn