SearcharxivSearch

arXiv subjects

Gabriele Sicuro

Publications and source records attributed to Gabriele Sicuro.

At least 19 recordsLinked to original sources

Spanning trees in the Assignment Problem: two theorems and two conjectures

The \emph{Minimum Matching Problem} consists of finding an independent edge set of minimum weight $M_{\star}(G)$ in a given edge-weighted graph $G$. When $G$ is bipartite, this reduces to the \emph{Assignment Problem}. We consider a variant of this problem defined by taking the union of optimal matchings across various slightly modified versions of the base graph: $H_{\mathcal{J}}(G)=\bigcup_{U \in \mathcal{J}} M_{\star}(G_{U})$. We establish two families of results: (1) In two distinct settings for the Assignment Problem, we prove that the resulting graphs $H_{\mathcal{J}}$, as well as certain associated graphs $\bar{H}_{\mathcal{J}}$, are spanning trees on the relevant base graphs $G$ and $\bar{G}$. (2) In these same settings, assuming the edge weights are given by the $p$-th power of Euclidean distances for point configurations in the plane, we show that for $p=1$ the tree $H_{\mathcal{J}}$ is non-crossing (i.e., its planar embedding has no crossing edges), whereas, remarkably, for $p=2$ the associated tree $\bar{H}_{\mathcal{J}}$ is non-crossing. Finally, we introduce novel conjectures in Statistical Mechanics, to be explored in future work: in the Random Euclidean Assignment Problem (where points are i.i.d.\ on a planar domain), we conjecture that for $p=2$ the trees $\bar{H}_{\mathcal{J}}$ are asymptotically distributed as Uniform Spanning Trees with free and wired boundary conditions in the two respective settings. In particular, suitable paths on the tree in the second setting, and on its planar dual in the first setting, are asymptotically distributed as $\text{SLE}_κ$ with $κ=2$.

math.CO

Sparse corruption in low-rank matrix inference: the PCA benchmark

Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations. It is known that applying PCA to a rank-one signal corrupted by a dense, homogeneous noise, in the large matrix size limit, the celebrated BBP transition occurs, where the emergence of an outlying eigenvalue and the alignment of the corresponding eigenvector occur at the same critical signal strength. Here we study the case of sparse noise corruption. The noise matrix is modelled as the adjacency matrix of a weighted undirected graph with finite average connectivity. Using the replica method, we analytically compute the typical top eigenvalue, the top eigenvector component density, and the squared overlap with the signal, through recursive distributional equations solved by population dynamics. We identify two signal-strength transitions as functions of graph connectivity: $θ_{\rm crit}$, marking signal recovery by the top eigenvector and generalising the BBP transition, and $θ_{\rm b}$, where the signal-related eigenvalue detaches from the bulk. For noise with nonzero mean, these transitions need not coincide because of a structural sparse-graph outlier, leading to a discontinuous transition in the squared overlap with the top eigenvector. The same top-eigenpair formalism also predicts the overlap of the signal with the eigenvector associated with the second largest eigenvalue when the signal eigenvalue is an outlier but remains below the structural outlier, where the transition is continuous. We specialise the equations to Poissonian and Random Regular degree distributions, recover dense-noise results in the large-connectivity limit, and validate the theory by numerical diagonalisation of large matrices.

stat.ML

Testing Transformer Learnability on the Arithmetic Sequence of Rooted Trees

We study whether a transformer network can learn the deterministic sequence of trees generated by the iterated prime factorization of the natural numbers. Each integer is mapped into a rooted planar tree and the resulting sequence $\mathbb{N}\mathscr{T}$ defines an arithmetic text with measurable statistical structure. A transformer network (the GPT-2 architecture) is trained from scratch on the first $10^{11}$ elements and evaluated on Next-Token and masked-word prediction tasks, with a Hidden Markov Model as baseline and a scaling analysis over context window, dataset size, vocabulary size and model size. The model reaches a word accuracy of about $0.4$, well above the baseline, and its performance remains stable on test blocks located at $10^{13}$--$10^{15}$, far beyond the training interval. Moreover, the likelihood assigned by the model separates the arithmetic text from two controls: synthetic sequences reproducing its word frequencies exactly but carrying no sequential organization, with a separation that widens as the evaluated context grows; and sequences containing more than three consecutive square-free integers, a configuration that arithmetic forbids. These results indicate that the transformer captures regularities of the arithmetic text that go beyond its frequency profile.

cs.AI

High-dimensional robust regression under heavy-tailed data: Asymptotics and Universality

We investigate the high-dimensional properties of robust regression estimators in the presence of heavy-tailed contamination of both the covariates and response functions. In particular, we provide a sharp asymptotic characterisation of M-estimators trained on a family of elliptical covariate and noise data distributions including cases where second and higher moments do not exist. We show that, despite being consistent, the Huber loss with optimally tuned location parameter $δ$ is suboptimal in the high-dimensional regime in the presence of heavy-tailed noise, highlighting the necessity of further regularisation to achieve optimal performance. This result also uncovers the existence of a transition in $δ$ as a function of the sample complexity and contamination. Moreover, we derive the decay rates for the excess risk of ridge regression. We show that, while it is both optimal and universal for covariate distributions with finite second moment, its decay rate can be considerably faster when the covariates' second moment does not exist. Finally, we show that our formulas readily generalise to a richer family of models and data distributions, such as generalised linear estimation with arbitrary convex regularisation trained on mixture models.

math.ST

Classification of Heavy-tailed Features in High Dimensions: a Superstatistical Approach

We characterise the learning of a mixture of two clouds of data points with generic centroids via empirical risk minimisation in the high dimensional regime, under the assumptions of generic convex loss and convex regularisation. Each cloud of data points is obtained via a double-stochastic process, where the sample is obtained from a Gaussian distribution whose variance is itself a random parameter sampled from a scalar distribution $\varrho$. As a result, our analysis covers a large family of data distributions, including the case of power-law-tailed distributions with no covariance, and allows us to test recent "Gaussian universality" claims. We study the generalisation performance of the obtained estimator, we analyse the role of regularisation, and we analytically characterise the separability transition.

stat.ML

Planted matching problems on random hypergraphs

We consider the problem of inferring a matching hidden in a weighted random $k$-hypergraph. We assume that the hyperedges' weights are random and distributed according to two different densities conditioning on the fact that they belong to the hidden matching, or not. We show that, for $k>2$ and in the large graph size limit, an algorithmic first order transition in the signal strength separates a regime in which a complete recovery of the hidden matching is feasible from a regime in which partial recovery is possible. This is in contrast to the $k=2$ case where the transition is known to be continuous. Finally, we consider the case of graphs presenting a mixture of edges and $3$-hyperedges, interpolating between the $k=2$ and the $k=3$ cases, and we study how the transition changes from continuous to first order by tuning the relative amount of edges and hyperedges.

cond-mat.dis-nn

Aligning random graphs with a sub-tree similarity message-passing algorithm

The problem of aligning Erdös-Rényi random graphs is a noisy, average-case version of the graph isomorphism problem, in which a pair of correlated random graphs is observed through a random permutation of their vertices. We study a polynomial time message-passing algorithm devised to solve the inference problem of partially recovering the hidden permutation, in the sparse regime with constant average degrees. We perform extensive numerical simulations to determine the range of parameters in which this algorithm achieves partial recovery. We also introduce a generalized ensemble of correlated random graphs with prescribed degree distributions, and extend the algorithm to this case.

cs.IT

Fluctuations, Bias, Variance & Ensemble of Learners: Exact Asymptotics for Convex Losses in High-Dimension

From the sampling of data to the initialisation of parameters, randomness is ubiquitous in modern Machine Learning practice. Understanding the statistical fluctuations engendered by the different sources of randomness in prediction is therefore key to understanding robust generalisation. In this manuscript we develop a quantitative and rigorous theory for the study of fluctuations in an ensemble of generalised linear models trained on different, but correlated, features in high-dimensions. In particular, we provide a complete description of the asymptotic joint distribution of the empirical risk minimiser for generic convex loss and regularisation in the high-dimensional limit. Our result encompasses a rich set of classification and regression tasks, such as the lazy regime of overparametrised neural networks, or equivalently the random features approximation of kernels. While allowing to study directly the mitigating effect of ensembling (or bagging) on the bias-variance decomposition of the test error, our analysis also helps disentangle the contribution of statistical fluctuations, and the singular role played by the interpolation threshold that are at the roots of the "double-descent" phenomenon.

stat.ML

Learning Gaussian Mixtures with Generalised Linear Models: Precise Asymptotics in High-dimensions

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussians with generic means and covariances via empirical risk minimisation (ERM) with any convex loss and regularisation. In particular, we prove exact asymptotics characterising the ERM estimator in high-dimensions, extending several previous results about Gaussian mixture classification in the literature. We exemplify our result in two tasks of interest in statistical learning: a) classification for a mixture with sparse means, where we study the efficiency of $\ell_1$ penalty with respect to $\ell_2$; b) max-margin multi-class classification, where we characterise the phase transition on the existence of the multi-class logistic maximum likelihood estimator for $K>2$. Finally, we discuss how our theory can be applied beyond the scope of synthetic data, showing that in different cases Gaussian mixtures capture closely the learning curve of classification tasks in real data sets.

stat.ML

Random assignment problems on ${2d}$ manifolds

We consider the assignment problem between two sets of $N$ random points on a smooth, two-dimensional manifold $Ω$ of unit area. It is known that the average cost scales as $E_Ω(N)\sim\frac{1}{2π}\ln N$ with a correction that is at most of order $\sqrt{\ln N\ln\ln N}$. In this paper, we show that, within the linearization approximation of the field-theoretical formulation of the problem, the first $Ω$-dependent correction is on the constant term, and can be exactly computed from the spectrum of the Laplace--Beltrami operator on $Ω$. We perform the explicit calculation of this constant for various families of surfaces, and compare our predictions with extensive numerics.

math-ph

Criticality and conformality in the random dimer model

In critical systems, the effect of a localized perturbation affects points that are arbitrarily far from the perturbation location. In this paper, we study the effect of localized perturbations on the solution of the random dimer problem in $2D$. By means of an accurate numerical analysis, we show that a local perturbation of the optimal covering induces an excitation whose size is extensive with finite probability. We compute the fractal dimension of the excitations and scaling exponents. In particular, excitations in random dimer problems on non-bipartite lattices have the same statistical properties of domain walls in the $2D$ spin glass. Excitations produced in bipartite lattices, instead, are compatible with a loop-erased self-avoiding random walk process. In both cases, we find evidence of conformal invariance of the excitations that is compatible with $\mathrm{SLE}_κ$ with parameter $κ$ depending on the bipartiteness of the underlying lattice only.

cond-mat.dis-nn

The planted $k$-factor problem

We consider the problem of recovering an unknown $k$-factor, hidden in a weighted random graph. For $k=1$ this is the planted matching problem, while the $k=2$ case is closely related to the planted travelling salesman problem. The inference problem is solved by exploiting the information arising from the use of two different distributions for the weights on the edges inside and outside the planted sub-graph. We argue that, in the large size limit, a phase transition can appear between a full and a partial recovery phase as function of the signal-to-noise ratio. We give a criterion for the location of the transition.

cond-mat.dis-nn

Recovery thresholds in the sparse planted matching problem

We consider the statistical inference problem of recovering an unknown perfect matching, hidden in a weighted random graph, by exploiting the information arising from the use of two different distributions for the weights on the edges inside and outside the planted matching. A recent work has demonstrated the existence of a phase transition, in the large size limit, between a full and a partial recovery phase for a specific form of the weights distribution on fully connected graphs. We generalize and extend this result in two directions: we obtain a criterion for the location of the phase transition for generic weights distributions and possibly sparse graphs, exploiting a technical connection with branching random walk processes, as well as a quantitatively more precise description of the critical regime around the phase transition.

cond-mat.dis-nn

Random-link matching problems on random regular graphs

We study the random-link matching problem on random regular graphs, alongside with two relaxed versions of the problem, namely the fractional matching and the so-called "loopy" fractional matching. We estimated the asymptotic average optimal cost using the cavity method. Moreover, we also study the finite-size corrections due to rare topological structures appearing in the graph at large sizes. We estimate these contributions using the cavity approach, and we compare our results with the output of numerical simulations. The analysis also clarifies the meaning of the finite-size contributions appearing in the fully-connected version of the problem, that has been already analyzed in the literature.

cond-mat.dis-nn

Fluctuations in the random-link matching problem

Using the replica approach and the cavity method, we study the fluctuations of the optimal cost in the random-link matching problem. By means of replica arguments, we derive the exact expression of its variance. Moreover, we study the large deviation function, deriving its expression in two different ways, namely using both the replica method and the cavity method.

cond-mat.dis-nn

Mean-field model for the density of states of jammed soft spheres

We propose a class of mean-field models for the isostatic transition of systems of soft spheres, in which the contact network is modeled as a random graph and each contact is associated to $d$ degrees of freedom. We study such models in the hypostatic, isostatic, and hyperstatic regimes. The density of states is evaluated by both the cavity method and exact diagonalization of the dynamical matrix. We show that the model correctly reproduces the main features of the density of states of real packings and, moreover, it predicts the presence of localized modes near the lower band edge. Finally, the behavior of the density of states $D(ω)\simω^α$ for $ω\to 0$ in the hyperstatic regime is studied. We find that the model predicts a nontrivial dependence of $α$ on the details of the coordination distribution.

cond-mat.dis-nn

The Random Fractional Matching Problem

We consider two formulations of the random-link fractional matching problem, a relaxed version of the more standard random-link (integer) matching problem. In one formulation, we allow each node to be linked to itself in the optimal matching configuration. In the other one, on the contrary, such a link is forbidden. Both problems have the same asymptotic average optimal cost of the random-link matching problem on the complete graph. Using a replica approach and previous results of Wästlund [Acta Mathematica 204, 91-150 (2010)], we analytically derive the finite-size corrections to the asymptotic optimal cost. We compare our results with numerical simulations and we discuss the main differences between random-link fractional matching problems and the random-link matching problem.

cond-mat.dis-nn

Anomalous scaling of the optimal cost in the one-dimensional random assignment problem

We consider the random Euclidean assignment problem on the line between two sets of $N$ random points, independently generated with the same probability density function $\varrho$. The cost of the matching is supposed to be dependent on a power $p>1$ of the Euclidean distance of the matched pairs. We discuss an integral expression for the average optimal cost for $N\gg 1$ that generalizes a previous result obtained for $p=2$. We also study the possible divergence of the given expression due to the vanishing of the probability density function. The provided regularization recipe allows us to recover the proper scaling law for the cost in the divergent cases, and possibly some of the involved coefficients. The possibility that the support of $\varrho$ is a disconnected interval is also analysed. We exemplify the proposed procedure and we compare our predictions with the results of numerical simulations.

cond-mat.dis-nn