SearcharxivSearch

arXiv subjects

Francesco Camilli

Publications and source records attributed to Francesco Camilli.

16 recordsLinked to original sources

Variational Bounds for Perceptron Learning from Structured Data

We introduce a variational approach to a finite-temperature continuous-spin perceptron trained on a Gaussian mixture. The model allows for a broad class of concave utilities and log-concave separable prior measures on the spins. By combining the interpolation method with log-concavity and concentration estimates, we derive lower and upper minimax variational bounds for the limiting quenched pressure. Remarkably, the two bounds differ only in the order of optimization of two variational parameters, while all remaining extrema are controlled by the concave--convex structure of the variational potential. Whenever the two optimizations commute, the two bounds match and identify the solution of the model. The same potential yields the fixed-point equations as stationarity conditions and provides a unified route to the computation of the ground-state energy, training loss, and generalization error.

cs.LG

From entropic constraints to reinforced processes: a probabilistic origin of multiscale measures

We investigate multiscale Gibbs measures from a variational and probabilistic viewpoint, focusing on the structural asymmetry among conditional entropies that characterizes their construction. We show how this asymmetry emerges both from variational principles with entropic constraints and from stochastic processes with reinforcement. We thus introduce the reinforced multinomial process and prove a large-deviation principle for its empirical histogram. The associated rate function reproduces precisely the entropy imbalance defining multiscale measures, thereby providing a genuine probabilistic mechanism for their emergence. The reinforced multinomial process thus offers a simple and rigorous stochastic foundation for multiscale Gibbs structures.

math-ph

Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime

We rigorously analyse fully-trained neural networks of arbitrary depth in the Bayesian optimal setting in the so-called proportional scaling regime where the number of training samples and width of the input and all inner layers diverge proportionally. We prove an information-theoretic equivalence between the Bayesian deep neural network model trained from data generated by a teacher with matching architecture, and a simpler model of optimal inference in a generalized linear model. This equivalence enables us to compute the optimal generalization error for deep neural networks in this regime. We thus prove the "deep Gaussian equivalence principle" conjectured in Cui et al. (2023) (arXiv:2302.00375). Our result highlights that in order to escape this "trivialisation" of deep neural networks (in the sense of reduction to a linear model) happening in the strongly overparametrized proportional regime, models trained from much more data have to be considered.

math.ST

Limit theorems for the non-convex multispecies Curie-Weiss model

We study the thermodynamic properties of the generalized non-convex multispecies Curie-Weiss model, where interactions among different types of particles (forming the species) are encoded in a generic matrix. For spins with a generic prior distribution, we compute the pressure in the thermodynamic limit using simple interpolation techniques. For Ising spins, we further analyze the fluctuations of the magnetization in the thermodynamic limit under the Boltzmann-Gibbs measure. It is shown that a central limit theorem holds for a rescaled and centered vector of species magnetizations, which converges to either a centered or non-centered multivariate normal distribution, depending on the rate of convergence of the relative sizes of the species.

math-ph

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size $n$ scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters $kd+k$ and of data $n$ are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate $n/d^2$, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.

stat.ML

On the phase diagram of the multiscale mean-field spin-glass

In this paper we study the phase diagram of a Sherrington-Kirkpatrick (SK) model where the couplings are forced to thermalize at different time scales. Besides being a challenging generalization of the SK model, such settings may arise naturally in physics whenever part of the many degrees of freedom of a system relaxes to equilibrium considerably faster than the others. For this model we compute the asymptotic value of the second moment of the overlap distribution. Furthermore, we provide a rigorous sufficient condition for an annealed solution to hold, identifying a high temperature, or weak coupling, region. In addition, we also prove that for sufficiently strong couplings the solution must present a number of replica symmetry breaking levels at least equal to the number of time scales already present in the multiscale model. Finally, we give a sufficient condition for the existence of gaps in the support of the functional order parameters.

math-ph

On the phase diagram of extensive-rank symmetric matrix denoising beyond rotational invariance

Matrix denoising is central to signal processing and machine learning. Its statistical analysis when the matrix to infer has a factorised structure with a rank growing proportionally to its dimension remains a challenge, except when it is rotationally invariant. In this case the information theoretic limits and an efficient Bayes-optimal denoising algorithm, called rotational invariant estimator [1,2], are known. Beyond this setting few results can be found. The reason is that the model is not a usual spin system because of the growing rank dimension, nor a matrix model (as appearing in high-energy physics) due to the lack of rotation symmetry, but rather a hybrid between the two. Here we make progress towards the understanding of Bayesian matrix denoising when the signal is a factored matrix $XX^\intercal$ that is not rotationally invariant. Monte Carlo simulations suggest the existence of a \emph{denoising-factorisation transition} separating a phase where denoising using the rotational invariant estimator remains Bayes-optimal due to universality properties of the same nature as in random matrix theory, from one where universality breaks down and better denoising is possible, though algorithmically hard. We argue that it is only beyond the transition that factorisation, i.e., estimating $X$ itself, becomes possible up to irresolvable ambiguities. On the theory side, we combine mean-field techniques in an interpretable multiscale fashion in order to access the minimum mean-square error and mutual information. Interestingly, our alternative method yields equations reproducible by the replica approach of [3]. Using numerical insights, we delimit the portion of phase diagram where we conjecture the mean-field theory to be exact, and correct it using universality when it is not. Our complete ansatz matches well the numerics in the whole phase diagram when considering finite size effects.

cond-mat.dis-nn

Information limits and Thouless-Anderson-Palmer equations for spiked matrix models with structured noise

We consider a prototypical problem of Bayesian inference for a structured spiked model: a low-rank signal is corrupted by additive noise. While both information-theoretic and algorithmic limits are well understood when the noise is a Gaussian Wigner matrix, the more realistic case of structured noise still proves to be challenging. To capture the structure while maintaining mathematical tractability, a line of work has focused on rotationally invariant noise. However, existing studies either provide sub-optimal algorithms or are limited to special cases of noise ensembles. In this paper, using tools from statistical physics (replica method) and random matrix theory (generalized spherical integrals) we establish the first characterization of the information-theoretic limits for a noise matrix drawn from a general trace ensemble. Remarkably, our analysis unveils the asymptotic equivalence between the rotationally invariant model and a surrogate Gaussian one. Finally, we show how to saturate the predicted statistical limits using an efficient algorithm inspired by the theory of adaptive Thouless-Anderson-Palmer (TAP) equations.

cs.IT

The Decimation Scheme for Symmetric Matrix Factorization

Matrix factorization is an inference problem that has acquired importance due to its vast range of applications that go from dictionary learning to recommendation systems and machine learning with deep networks. The study of its fundamental statistical limits represents a true challenge, and despite a decade-long history of efforts in the community, there is still no closed formula able to describe its optimal performances in the case where the rank of the matrix scales linearly with its size. In the present paper, we study this extensive rank problem, extending the alternative 'decimation' procedure that we recently introduced, and carry out a thorough study of its performance. Decimation aims at recovering one column/line of the factors at a time, by mapping the problem into a sequence of neural network models of associative memory at a tunable temperature. Though being sub-optimal, decimation has the advantage of being theoretically analyzable. We extend its scope and analysis to two families of matrices. For a large class of compactly supported priors, we show that the replica symmetric free entropy of the neural network models takes a universal form in the low temperature limit. For sparse Ising prior, we show that the storage capacity of the neural network models diverges as sparsity in the patterns increases, and we introduce a simple algorithm based on a ground state search that implements decimation and performs matrix factorization, with no need of an informative initialization.

cond-mat.dis-nn

Fundamental limits of overparametrized shallow neural networks for supervised learning

We carry out an information-theoretical analysis of a two-layer neural network trained from input-output pairs generated by a teacher network with matching architecture, in overparametrized regimes. Our results come in the form of bounds relating i) the mutual information between training data and network weights, or ii) the Bayes-optimal generalization error, to the same quantities but for a simpler (generalized) linear model for which explicit expressions are rigorously known. Our bounds, which are expressed in terms of the number of training samples, input dimension and number of hidden units, thus yield fundamental performance limits for any neural network (and actually any learning procedure) trained from limited data generated according to our two-layer teacher neural network model. The proof relies on rigorous tools from spin glasses and is guided by ``Gaussian equivalence principles'' lying at the core of numerous recent analyses of neural networks. With respect to the existing literature, which is either non-rigorous or restricted to the case of the learning of the readout weights only, our results are information-theoretic (i.e. are not specific to any learning algorithm) and, importantly, cover a setting where all the network parameters are trained.

cs.LG

Central limit theorem for the overlaps on the Nishimori line

The overlap distribution of the Sherrington-Kirkpatrick model on the Nishimori line has been proved to be self averaging for large volumes. Here we study the joint distribution of the rescaled overlaps around their common mean and prove that it converges to a Gaussian vector.

math-ph

Matrix factorization with neural networks

Matrix factorization is an important mathematical problem encountered in the context of dictionary learning, recommendation systems and machine learning. We introduce a new `decimation' scheme that maps it to neural network models of associative memory and provide a detailed theoretical analysis of its performance, showing that decimation is able to factorize extensive-rank matrices and to denoise them efficiently. We introduce a decimation algorithm based on ground-state search of the neural network, which shows performances that match the theoretical prediction.

cond-mat.dis-nn

Bayes-optimal limits in structured PCA, and how to reach them

How do statistical dependencies in measurement noise influence high-dimensional inference? To answer this, we study the paradigmatic spiked matrix model of principal components analysis (PCA), where a rank-one matrix is corrupted by additive noise. We go beyond the usual independence assumption on the noise entries, by drawing the noise from a low-order polynomial orthogonal matrix ensemble. The resulting noise correlations make the setting relevant for applications but analytically challenging. We provide the first characterization of the Bayes-optimal limits of inference in this model. If the spike is rotation-invariant, we show that standard spectral PCA is optimal. However, for more general priors, both PCA and the existing approximate message passing algorithm (AMP) fall short of achieving the information-theoretic limits, which we compute using the replica method from statistical mechanics. We thus propose a novel AMP, inspired by the theory of Adaptive Thouless-Anderson-Palmer equations, which saturates the theoretical limit. This AMP comes with a rigorous state evolution analysis tracking its performance. Although we focus on specific noise distributions, our methodology can be generalized to a wide class of trace matrix ensembles at the cost of more involved expressions. Finally, despite the seemingly strong assumption of rotation-invariant noise, our theory empirically predicts algorithmic performance on real data, pointing at remarkable universality properties.

cs.IT

An inference problem in a mismatched setting: a spin-glass model with Mattis interaction

The Wigner spiked model in a mismatched setting is studied with the finite temperature Statistical Mechanics approach through its representation as a Sherrington-Kirkpatrick model with added Mattis interaction. The exact solution of the model with Ising spins is rigorously proved to be given by a variational principle on two order parameters, the Parisi overlap distribution and the Mattis magnetization. The latter is identified by an ordinary variational principle and turns out to concentrate in the thermodynamic limit. The solution leads to the computation of the Mean Square Error of the mismatched reconstruction. The Gaussian signal distribution case is investigated and the corresponding phase diagram is identified.

cond-mat.dis-nn

The solution of the deep Boltzmann machine on the Nishimori line

The deep Boltzmann machine on the Nishimori line with a finite number of layers is exactly solved by a theorem that expresses its pressure through a finite dimensional variational problem of min-max type. In the absence of magnetic fields the order parameter is shown to exhibit a phase transition.

math-ph

The multi-species mean-field spin-glass on the Nishimori line

In this paper we study a multi-species disordered model on the Nishimori line. The typical properties of this line, a set of identities and inequalities among correlation functions, allow us to prove the replica symmetry i.e. the concentration of the order parameter. When the interaction structure is elliptic we rigorously compute the exact solution of the model in terms of a finite dimensional variational principle and we study its properties.

math-ph