SearcharxivSearch

arXiv subjects

Mauro Pastore

Publications and source records attributed to Mauro Pastore.

15 recordsLinked to original sources

A High-Order Cumulant Extension of Quasi-Linkage Equilibrium

A central question in evolutionary biology is how to quantitatively understand the dynamics of genetically diverse populations. Modeling the genotype distribution is challenging, as it ultimately requires tracking all correlations (or cumulants) among alleles at different loci. The quasi-linkage equilibrium (QLE) approximation simplifies this by assuming that correlations between alleles at different loci are weak -- i.e., low linkage disequilibrium -- allowing their dynamics to be modeled perturbatively. However, QLE breaks down under strong selection, significant epistatic interactions, or weak recombination. We extend the multilocus QLE framework to allow cumulants up to order $K$ to evolve dynamically, while higher-order cumulants ($>K$) are assumed to equilibrate rapidly. This extended QLE (exQLE) framework yields a general equation of motion for cumulants up to order $K$, which parallels the standard QLE dynamics (recovered when $K = 1$). In this formulation, cumulant dynamics are driven by the gradient of average fitness, mediated by a geometrically interpretable matrix that stems from competition among genotypes. Our analysis shows that the exQLE with $K=2$ accurately captures cumulant dynamics even when the fitness function includes higher-order (e.g., third- or fourth-order) epistatic interactions, capabilities that standard QLE lacks. We also applied the exQLE framework to infer fitness parameters from temporal sequence data. Overall, exQLE provides a systematic and interpretable approximation scheme, leveraging analytical cumulant dynamics and reducing complexity by progressively truncating higher-order cumulants.

q-bio.PE

Comment on "Storage properties of a quantum perceptron"

The recent paper "Storage properties of a quantum perceptron" [Phys. Rev. E 110, 024127] considers a quadratic constraint satisfaction problem, motivated by a quantum version of the perceptron. In particular, it derives its critical capacity, the density of constraints at which there is a satisfiability transition. The same problem was considered before in another context (classification of geometrically structured inputs, see [Phys. Rev. Lett. 125, 120601; Phys. Rev. E 102, 032119; J. Stat. Mech. (2021) 113301]), but the results on the critical capacity drastically differ. In this note, I substantiate the claim that the derivation performed in the quantum scenario has issues when inspected closely, I report a more principled way to perform it and I evaluate the critical capacity of an alternative constraint satisfaction problem that I consider more relevant for the quantum perceptron rule proposed by the article in question.

cond-mat.dis-nn

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size $n$ scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters $kd+k$ and of data $n$ are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate $n/d^2$, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.

stat.ML

Feature learning in finite-width Bayesian deep linear networks with multiple outputs and convolutional layers

Deep linear networks have been extensively studied, as they provide simplified models of deep learning. However, little is known in the case of finite-width architectures with multiple outputs and convolutional layers. In this manuscript, we provide rigorous results for the statistics of functions implemented by the aforementioned class of networks, thus moving closer to a complete characterization of feature learning in the Bayesian setting. Our results include: (i) an exact and elementary non-asymptotic integral representation for the joint prior distribution over the outputs, given in terms of a mixture of Gaussians; (ii) an analytical formula for the posterior distribution in the case of squared error loss function (Gaussian likelihood); (iii) a quantitative description of the feature learning infinite-width regime, using large deviation theory. From a physical perspective, deep architectures with multiple outputs or convolutional layers represent different manifestations of kernel shape renormalization, and our work provides a dictionary that translates this physics intuition and terminology into rigorous Bayesian statistics.

stat.ML

Restoring balance: principled under/oversampling of data for optimal classification

Class imbalance in real-world data poses a common bottleneck for machine learning tasks, since achieving good generalization on under-represented examples is often challenging. Mitigation strategies, such as under or oversampling the data depending on their abundances, are routinely proposed and tested empirically, but how they should adapt to the data statistics remains poorly understood. In this work, we determine exact analytical expressions of the generalization curves in the high-dimensional regime for linear classifiers (Support Vector Machines). We also provide a sharp prediction of the effects of under/oversampling strategies depending on class imbalance, first and second moments of the data, and the metrics of performance considered. We show that mixed strategies involving under and oversampling of data lead to performance improvement. Through numerical experiments, we show the relevance of our theoretical predictions on real datasets, on deeper architectures and with sampling strategies based on unsupervised probabilistic models.

cond-mat.dis-nn

Random features and polynomial rules

Random features models play a distinguished role in the theory of deep learning, describing the behavior of neural networks close to their infinite-width limit. In this work, we present a thorough analysis of the generalization performance of random features models for generic supervised learning problems with Gaussian data. Our approach, built with tools from the statistical mechanics of disordered systems, maps the random features model to an equivalent polynomial model, and allows us to plot average generalization curves as functions of the two main control parameters of the problem: the number of random features $N$ and the size $P$ of the training set, both assumed to scale as powers in the input dimension $D$. Our results extend the case of proportional scaling between $N$, $P$ and $D$. They are in accordance with rigorous bounds known for certain particular learning tasks and are in quantitative agreement with numerical experiments performed over many order of magnitudes of $N$ and $P$. We find good agreement also far from the asymptotic limits where $D\to \infty$ and at least one between $P/D^K$, $N/D^L$ remains finite.

cond-mat.dis-nn

Satisfiability transition in asymmetric neural networks

Asymmetry in the synaptic interactions between neurons plays a crucial role in determining the memory storage and retrieval properties of recurrent neural networks. In this work, we analyze the problem of storing random memories in a network of neurons connected by a synaptic matrix with a definite degree of asymmetry. We study the corresponding satisfiability and clustering transitions in the space of solutions of the constraint satisfaction problem associated with finding synaptic matrices given the memories. We find, besides the usual SAT/UNSAT transition at a critical number of memories to store in the network, an additional transition for very asymmetric matrices, where the competing constraints (definite asymmetry vs. memories storage) induce enough frustration in the problem to make it impossible to solve. This finding is particularly striking in the case of a single memory to store, where no quenched disorder is present in the system.

cond-mat.dis-nn

Critical properties of the SAT/UNSAT transitions in the classification problem of structured data

The classification problem of structured data can be solved with different strategies: a supervised learning approach, starting from a labeled training set, and an unsupervised learning one, where only the structure of the patterns in the dataset is used to find a classification compatible with it. The two strategies can be interpreted as extreme cases of a semi-supervised approach to learn multi-view data, relevant for applications. In this paper I study the critical properties of the two storage problems associated with these tasks, in the case of the linear binary classification of doublets of points sharing the same label, within replica theory. While the first approach presents a SAT/UNSAT transition in a (marginally) stable replica-symmetric phase, in the second one the satisfiability line lies in a full replica-symmetry-broken phase. A similar behavior in the problem of learning with a margin is also pointed out.

cond-mat.dis-nn

Self-induced glassy phase in multimodal cavity quantum electrodynamics

We provide strong evidence that the effective spin-spin interaction in a multimodal confocal optical cavity gives rise to a self-induced glassy phase, which emerges exclusively from the peculiar euclidean correlations and is not related to the presence of disorder as in standard spin glasses. As recently shown, this spin-spin effective interaction is both non-local and non-translational invariant, and randomness in the atoms positions produces a spin glass phase. Here we consider the simplest feasible disorder-free setting where atoms form a one-dimensional regular chain and we study the thermodynamics of the resulting effective Ising model. We present extensive results showing that the system has a low-temperature glassy phase. Notably, for rational values of the only free adimensional parameter $α=p/q$ of the interaction, the number of metastable states at low temperature grows exponentially with $q$ and the problem of finding the ground state rapidly becomes computationally intractable, suggesting that the system develops high energy barriers and ergodicity breaking occurs.

cond-mat.stat-mech

Statistical learning theory of structured data

The traditional approach of statistical physics to supervised learning routinely assumes unrealistic generative models for the data: usually inputs are independent random variables, uncorrelated with their labels. Only recently, statistical physicists started to explore more complex forms of data, such as equally-labelled points lying on (possibly low dimensional) object manifolds. Here we provide a bridge between this recently-established research area and the framework of statistical learning theory, a branch of mathematics devoted to inference in machine learning. The overarching motivation is the inadequacy of the classic rigorous results in explaining the remarkable generalization properties of deep learning. We propose a way to integrate physical models of data into statistical learning theory, and address, with both combinatorial and statistical mechanics methods, the computation of the Vapnik-Chervonenkis entropy, which counts the number of different binary classifications compatible with the loss class. As a proof of concept, we focus on kernel machines and on two simple realizations of data structure introduced in recent physics literature: $k$-dimensional simplexes with prescribed geometric relations and spherical manifolds (equivalent to margin classification). Entropy, contrary to what happens for unstructured data, is nonmonotonic in the sample size, in contrast with the rigorous bounds. Moreover, data structure induces a novel transition beyond the storage capacity, which we advocate as a proxy of the nonmonotonicity, and ultimately a cue of low generalization error. The identification of a synaptic volume vanishing at the transition allows a quantification of the impact of data structure within replica theory, applicable in cases where combinatorial methods are not available, as we demonstrate for margin learning.

cond-mat.stat-mech

Beyond the storage capacity: data driven satisfiability transition

Data structure has a dramatic impact on the properties of neural networks, yet its significance in the established theoretical frameworks is poorly understood. Here we compute the Vapnik-Chervonenkis entropy of a kernel machine operating on data grouped into equally labelled subsets. At variance with the unstructured scenario, entropy is non-monotonic in the size of the training set, and displays an additional critical point besides the storage capacity. Remarkably, the same behavior occurs in margin classifiers even with randomly labelled data, as is elucidated by identifying the synaptic volume encoding the transition. These findings reveal aspects of expressivity lying beyond the condensed description provided by the storage capacity, and they indicate the path towards more realistic bounds for the generalization error of neural networks.

cs.LG

Large deviations of the free energy in the p-spin glass spherical model

We investigate the behavior of the rare fluctuations of the free energy in the p-spin spherical model, evaluating the corresponding rate function via the Gärtner-Ellis theorem. This approach requires the knowledge of the analytic continuation of the disorder-averaged replicated partition function to arbitrary real number of replicas. In zero external magnetic field, we show via a one-step replica symmetry breaking (1RSB) calculation that the rate function is infinite for fluctuations of the free energy above its typical value, corresponding to an anomalous, super-extensive suppression of rare fluctuations. We extend this calculation to non-zero magnetic field, showing that in this case this very large deviation disappears and we try to motivate this finding in light of a geometrical interpretation of the scaled cumulant generating function.

cond-mat.dis-nn

Remarks on replica diagonal collective field condensations in SYK

In the Sachdev-Ye-Kitaev model with generic order $q \ge 4$ random couplings, we compute the critical temperature relating the Majorana fermions high temperature perturbative vacuum to the vacuum where the replica diagonal collective field $G(τ, τ')$ condenses. We study, by a finite temperature diagrammatic analysis, the effective action of an auxiliary Hubbard-Stratonovich bilocal field related to $G(τ, τ')$ in the large $N$ limit. Subtelties that arise in switching from the operatorial to the functional integral representation of the SYK thermal partition function are also discussed.

hep-th

Lattice QCD$_2$ effective action with Bogoliubov transformations

In the Wilson's lattice formulation of QCD, a fermionic Fock space of states can be explicitly built at each time slice using canonical creation and annihilation operators. The partition function $Z$ is then represented as the trace of the transfer matrix, and its usual functional representation as a path integral of $\exp(- S)$ can be recovered in a standard way. However, applying a Bogoliubov transformation on the canonical operators before passing to the functional formalism, we can isolate a vacuum contribution in the resulting action which depends only on the parameters of the transformation and fixes them via a variational principle. Then, inserting in the trace defining $Z$ an operator projecting on the mesons subspace at each time slice and making the physical assumption that the true partition function is well approximate by the projected one, we can also write an effective quadratic action for mesons. We tested the method in the renowned 't Hooft model, namely QCD in two spacetime dimensions for large number of colours, in Coulomb gauge.

hep-lat

Effective mesonic theory for the 't Hooft model on the lattice

We apply to a lattice version of the 't~Hooft model, QCD in two space-time dimensions for large number of colours, a method recently proposed to obtain an effective mesonic action starting from the fundamental, fermionic one. The idea is to pass from a canonical, operatorial representation, where the low-energy states have a direct physical interpretation in terms of a Bogoliubov vacuum and its corresponding quasiparticle excitations, to a functional, path integral representation, via the formalism of the transfer matrix. In this way we obtain a lattice effective theory for mesons in a self-consistent setting. We also verify that well-known results from other different approaches are reproduced in the continuum limit.

hep-lat