Searcharxiv⌕ Search

arXiv subjects

Chiara Cammarota

Publications and source records attributed to Chiara Cammarota.

At least 19 recordsLinked to original sources

Spectral properties and phase diagrams of sparse antagonistic random matrices with diagonal disorder and Jacobian-like structure

Complex interacting systems are often modelled by random matrices whose spectral properties dictate stability. In sparse antagonistic matrices without diagonal disorder, low connectivity gives rise to a characteristic reentrance effect in the spectral boundary near the real axis, which disappears via a continuous transition as the connectivity increases. The reentrance effect implies the presence of a complex leading eigenvalue, which suggests the existence of a phase characterized by oscillatory dynamics around equilibrium. Here, we expand the investigation to matrices featuring diagonal disorder and a Jacobian-like structure. In these settings, the spectrum also develops a segment of eigenvalues accumulating on the real axis, which can trigger a discontinuous jump of the complex leading eigenvalue to a purely real value. The interplay between connectivity and disorder produces a rich variety of spectral behaviours. Employing the cavity method and a an adaptation of the Population Dynamics algorithm, we map a phase diagram with five distinct spectral phases. Finally, we show that the algorithm underestimates the spectral support under strong disorder, motivating future technical developments to handle this limit.

cond-mat.dis-nn↗

Discontinuous BBP transitions

The Baik-Ben Arous-Peche (BBP) transition sets fundamental limits for detecting low-rank structure in noisy high-dimensional data and underlies a wide range of spectral methods in many fields from physics to statistics and data sciences. In standard settings, this transition is continuous, implying that signal recovery emerges gradually above a sharp threshold. We show that BBP transitions can instead be discontinuous in very general settings and provide a full theory of this phenomenon. When the eigenvalue density vanishes faster than linearly at the spectral edge, the overlap between the leading eigenvector and the signal jumps discontinuously at the critical point. We study this mechanism in deformed Gaussian and reweighted Wishart ensembles. We analyze in detail the finite-size effects, which play a central and qualitatively new role in the discontinuous BBP transition. Unlike the continuous BBP transition, we establish the existence of an extended pre-critical region where informative eigenvectors emerge well before the asymptotic threshold. The main consequence-and difference from the continuous BBP transition-is that signal recovery can occur at significantly lower signal-to-noise ratio and it is accompanied by strong sample-to-sample variability. Our results show the relevance and the novelty of the discontinuous BBP transition, and highlight the practical implications for signal detection.

cond-mat.dis-nn↗

Escape dynamics and implicit bias of one-pass SGD in overparameterized quadratic networks

We analyze the one-pass stochastic gradient descent dynamics of a two-layer neural network with quadratic activations in a teacher--student framework. In the high-dimensional regime, where the input dimension $N$ and the number of samples $M$ diverge at fixed ratio $α= M/N$, and for finite hidden widths $(p,p^*)$ of the student and teacher, respectively, we study the low-dimensional ordinary differential equations that govern the evolution of the student--teacher and student--student overlap matrices. We show that overparameterization ($p>p^*$) only modestly accelerates escape from a plateau of poor generalization by modifying the prefactor of the exponential decay of the loss. We then examine how unconstrained weight norms introduce a continuous rotational symmetry that results in a nontrivial manifold of zero-loss solutions for $p>1$. From this manifold the dynamics consistently selects the closest solution to the random initialization, as enforced by a conserved quantity in the ODEs governing the evolution of the overlaps. Finally, a Hessian analysis of the population-loss landscape confirms that the plateau and the solution manifold correspond to saddles with at least one negative eigenvalue and to marginal minima in the population-loss geometry, respectively.

cond-mat.dis-nn↗

Eigenvalue spectral tails and localization properties of asymmetric networks

In contrast to the neatly bounded spectra of densely populated large random matrices, sparse random matrices often exhibit unbounded eigenvalue tails on the real and imaginary axis, called Lifshitz tails. In the case of asymmetric matrices, concise mathematical results have proved elusive. In this work, we present an analytical approach to characterising these tails. We exploit the fact that eigenvalues in the tail region have corresponding eigenvectors that are exponentially localised on highly-connected hubs of the network associated to the random matrix. We approximate these eigenvectors using a series expansion in the inverse connectivity of the hub, where successive terms in the series take into account further sets of next-nearest neighbours. By considering the ensemble of such hubs, we are able to characterise the eigenvalue density and the extent of localisation in the tails of the spectrum in a general fashion. As such, we classify a number of different asymptotic behaviours in the Lifshitz tails, as well as the leading eigenvalue and the inverse participation ratio. We demonstrate how an interplay between matrix asymmetry, network structure, and the edge-weight distribution leads to the variety of observed behaviours.

cond-mat.dis-nn↗

Overparametrization bends the landscape: BBP transitions at initialization in simple Neural Networks

High-dimensional non-convex loss landscapes play a central role in the theory of Machine Learning. Gaining insight into how these landscapes interact with gradient-based optimization methods, even in relatively simple models, can shed light on this enigmatic feature of neural networks. In this work, we will focus on a prototypical simple learning problem, which generalizes the Phase Retrieval inference problem by allowing the exploration of overparametrized settings. Using techniques from field theory, we analyze the spectrum of the Hessian at initialization and identify a Baik-Ben Arous-Péché (BBP) transition in the amount of data that separates regimes where the initialization is informative or uninformative about a planted signal of a teacher-student setup. Crucially, we demonstrate how overparameterization can bend the loss landscape, shifting the transition point, even reaching the information-theoretic weak-recovery threshold in the large overparameterization limit, while also altering its qualitative nature. We distinguish between continuous and discontinuous BBP transitions and support our analytical predictions with simulations, examining how they compare to the finite-N behavior. In the case of discontinuous BBP transitions strong finite-N corrections allow the retrieval of information at a signal-to-noise ratio (SNR) smaller than the predicted BBP transition. In these cases we provide estimates for a new lower SNR threshold that marks the point at which initialization becomes entirely uninformative.

cond-mat.dis-nn↗

The Role of the Time-Dependent Hessian in High-Dimensional Optimization

Gradient descent is commonly used to find minima in rough landscapes, particularly in recent machine learning applications. However, a theoretical understanding of why good solutions are found remains elusive, especially in strongly non-convex and high-dimensional settings. Here, we focus on the phase retrieval problem as a typical example, which has received a lot of attention recently in theoretical machine learning. We analyze the Hessian during gradient descent, identify a dynamical transition in its spectral properties, and relate it to the ability of escaping rough regions in the loss landscape. When the signal-to-noise ratio (SNR) is large enough, an informative negative direction exists in the Hessian at the beginning of the descent, i.e in the initial condition. While descending, a BBP transition in the spectrum takes place in finite time: the direction is lost, and the dynamics is trapped in a rugged region filled with marginally stable bad minima. Surprisingly, for finite system sizes, this window of negative curvature allows the system to recover the signal well before the theoretical SNR found for infinite sizes, emphasizing the central role of initialization and early-time dynamics for efficiently navigating rough landscapes.

cs.LG↗

Dynamical systems on large networks with predator-prey interactions are stable and exhibit oscillations

We analyse the stability of linear dynamical systems defined on sparse, random graphs with predator-prey, competitive, and mutualistic interactions. These systems are aimed at modelling the stability of fixed points in large systems defined on complex networks, such as, ecosystems consisting of a large number of species that interact through a food-web. We develop an exact theory for the spectral distribution and the leading eigenvalue of the corresponding sparse Jacobian matrices. This theory reveals that the nature of local interactions have a strong influence on system's stability. We show that, in general, linear dynamical systems defined on random graphs with a prescribed degree distribution of unbounded support are unstable if they are large enough, implying a tradeoff between stability and diversity. Remarkably, in contrast to the generic case, antagonistic systems that only contain interactions of the predator-prey type can be stable in the infinite size limit. This qualitatively feature for antagonistic systems is accompanied by a peculiar oscillatory behaviour of the dynamical response of the system after a perturbation, when the mean degree of the graph is small enough. Moreover, for antagonistic systems we also find that there exist a dynamical phase transition and critical mean degree above which the response becomes non-oscillatory.

cond-mat.stat-mech↗

Daydreaming Hopfield Networks and their surprising effectiveness on correlated data

To improve the storage capacity of the Hopfield model, we develop a version of the dreaming algorithm that perpetually reinforces the patterns to be stored (as in the Hebb rule), and erases the spurious memories (as in dreaming algorithms). For this reason, we called it Daydreaming. Daydreaming is not destructive and it converges asymptotically to stationary retrieval maps. When trained on random uncorrelated examples, the model shows optimal performance in terms of the size of the basins of attraction of stored examples and the quality of reconstruction. We also train the Daydreaming algorithm on correlated data obtained via the random-features model and argue that it spontaneously exploits the correlations thus increasing even further the storage capacity and the size of the basins of attraction. Moreover, the Daydreaming algorithm is also able to stabilize the features hidden in the data. Finally, we test Daydreaming on the MNIST dataset and show that it still works surprisingly well, producing attractors that are close to unseen examples and class prototypes.

cond-mat.dis-nn↗

Local sign stability and its implications for spectra of sparse random graphs and stability of ecosystems

We study the spectral properties of sparse random graphs with different topologies and type of interactions, and their implications on the stability of complex systems, with particular attention to ecosystems. Specifically, we focus on the behaviour of the leading eigenvalue in different type of random matrices (including interaction matrices and Jacobian-like matrices), relevant for the assessment of different types of dynamical stability. By comparing the results on Erdos-Renyi and Husimi graphs with sign-antisymmetric interactions or mixed sign patterns, we introduce a sufficient criterion, called strong local sign stability, for stability not to be affected by system size, as traditionally implied by the complexity-stability trade-off in conventional models of random matrices. The criterion requires sign-antisymmetric or unidirectional interactions and a local structure of the graph such that the number of cycles of finite length do not increase with the system size. Note that the last requirement is stronger than the classical local tree-like condition, which we associate to the less stringent definition of local sign stability, also defined in the paper. In addition, for strong local sign stable graphs which show stability to linear perturbations irrespectively of system size, we observe that the leading eigenvalue can undergo a transition from being real to acquiring a nonnull imaginary part, which implies a dynamical transition from nonoscillatory to oscillatory linear response to perturbations. Lastly, we ascertain the discontinuous nature of this transition.

cond-mat.dis-nn↗

The Kauzmann Transition to an Ideal Glass Phase

The idea that a thermodynamic glass transition of some sort underlies the observed glass formation has been highly debated since Kauzmann first stressed the hypothetical entropy crisis that could take place if one were able to equilibrate supercooled liquids below the experimental glass transition temperature $T_g$. This a priori unreachable transition at some $T_K<T_g$ has since received a firm theoretical basis as a key feature predicted by the mean-field theory of the glass transition. In this chapter, we assess whether, and in which form, such a transition can survive in finite dimensions, and we review some of the recent computer simulation work addressing the issue in $2$- and $3$-dimensional glass-forming liquid models. We also discuss theoretical reasons to focus on an apparently inaccessible singularity.

cond-mat.dis-nn↗

Who has the last word? Understanding How to Sample Online Discussions

In online debates individual arguments support or attack each other, leading to some subset of arguments being considered more relevant than others. However, in large discussions readers are often forced to sample a subset of the arguments being put forth. Since such sampling is rarely done in a principled manner, users may not read all the relevant arguments to get a full picture of the debate. This paper is interested in answering the question of how users should sample online conversations to selectively favour the currently justified or accepted positions in the debate. We apply techniques from argumentation theory and complex networks to build a model that predicts the probabilities of the normatively justified arguments given their location in online discussions. Our model shows that the proportion of replies that are supportive, the number of replies that comments receive, and the locations of un-replied comments all determine the probability that a comment is a justified argument. We show that when the degree distribution of the number of replies is homogeneous along the discussion, for acrimonious discussions, the distribution of justified arguments depends on the parity of the graph level. In supportive discussions the probability of having justified comments increases as one moves away from the root. For discussion trees that have a non-homogeneous in-degree distribution, for supportive discussions we observe the same behaviour as before, while for acrimonious discussions we cannot observe the same parity-based distribution. This is verified with data obtained from the online debating platform Kialo. By predicting the locations of the justified arguments in reply trees, we can suggest which arguments readers should sample to grasp the currently accepted opinions in such discussions. Our models have important implications for the design of future online debating platforms.

cs.SI↗

Opinion dynamics with emergent collective memory: the impact of a long and heterogeneous news history

In modern society people are being exposed to numerous information, with some of them being frequently repeated or more disruptive than others. In this paper we use a model of opinion dynamics to study how this news impact the society. In particular, our study aims to explain how the exposure of the society to certain events deeply change people's perception of the present and future. The evolution of opinions which we consider is influenced both by external information and the pressure of the society. The latter includes imitation, differentiation, homophily and its opposite, xenophobia. The combination of these ingredients gives rise to a collective memory effect, which is triggered by external information. In this paper we focus our attention on how this memory arises when the order of appearance of external news is random. We will show which characteristics a piece of news needs to have in order to be embedded in the society's memory. We will also provide an analytical way to measure how many information a society can remember when an extensive number of news items is presented. Finally we will show that, when a certain piece of news is present in the society's history, even a distorted version of it is sufficient to trigger the memory of the originally stored information.

physics.soc-ph↗

Properties of equilibria and glassy phases of the random Lotka-Volterra model with demographic noise

In this letter we study a reference model in theoretical ecology, the disordered Lotka-Volterra model for ecological communities, in the presence of finite demographic noise. Our theoretical analysis, which takes advantage of a mapping to an equilibrium disordered system, proves that for sufficiently heterogeneous interactions and low demographic noise the system displays a multiple equilibria phase, which we fully characterize. In particular, we show that in this phase the number of stable equilibria is exponential in the number of species. Upon further decreasing the demographic noise, we unveil a "Gardner" transition to a marginally stable phase, similar to that observed in jamming of amorphous materials. We confirm and complement our analytical results by numerical simulations. Furthermore, we extend their relevance by showing that they hold for others interacting random dynamical systems, such as the Random Replicant Model. Finally, we discuss their extension to the case of asymmetric couplings.

cond-mat.dis-nn↗

Opinion dynamics with memory: how a society is shaped by its own past

In order to understand the development of common orientation of opinions in the modern world we propose a model of a society described as a large collection of agents that exchange their expressed opinions under the influence of their mutual interactions and external events. In particular we introduce an interaction bias which creates a collective memory effect such that the society is able to store and recall information coming from several external signals. Our model shows how the inner structure of the society and its future reactions can be shaped by its own history. We will provide an analytical explanation of how this might occur and we will show the emergent similarity between the reaction of a society modelled in this way and the Hopfield mechanism for information retrieval.

physics.soc-ph↗

Complex Dynamics in Simple Neural Networks: Understanding Gradient Flow in Phase Retrieval

Despite the widespread use of gradient-based algorithms for optimizing high-dimensional non-convex functions, understanding their ability of finding good minima instead of being trapped in spurious ones remains to a large extent an open problem. Here we focus on gradient flow dynamics for phase retrieval from random measurements. When the ratio of the number of measurements over the input dimension is small the dynamics remains trapped in spurious minima with large basins of attraction. We find analytically that above a critical ratio those critical points become unstable developing a negative direction toward the signal. By numerical experiments we show that in this regime the gradient flow algorithm is not trapped; it drifts away from the spurious critical points along the unstable direction and succeeds in finding the global minimum. Using tools from statistical physics we characterize this phenomenon, which is related to a BBP-type transition in the Hessian of the spurious minima.

cs.LG↗

Dynamical Mean-Field Theory and Aging Dynamics

Dynamical Mean-Field Theory (DMFT) replaces the many-body dynamical problem with one for a single degree of freedom in a thermal bath whose features are determined self-consistently. By focusing on models with soft disordered $p$-spin interactions, we show how to incorporate the mean-field theory of aging within dynamical mean-field theory. We study cases with only one slow time-scale, corresponding statically to the one-step replica symmetry breaking (1RSB) phase, and cases with an infinite number of slow time-scales, corresponding statically to the full replica symmetry breaking (FRSB) phase. For the former, we show that the effective temperature of the slow degrees of freedom is fixed by requiring critical dynamical behavior on short time-scales, i.e. marginality. For the latter, we find that aging on an infinite number of slow time-scales is governed by a stochastic equation where the clock for dynamical evolution is fixed by the change of effective temperature, hence obtaining a dynamical derivation of the stochastic equation at the basis of the FRSB phase. Our results extend the realm of the mean-field theory of aging to all situations where DMFT holds.

cond-mat.dis-nn↗

How to iron out rough landscapes and get optimal performances: Averaged Gradient Descent and its application to tensor PCA

In many high-dimensional estimation problems the main task consists in minimizing a cost function, which is often strongly non-convex when scanned in the space of parameters to be estimated. A standard solution to flatten the corresponding rough landscape consists in summing the losses associated to different data points and obtain a smoother empirical risk. Here we propose a complementary method that works for a single data point. The main idea is that a large amount of the roughness is uncorrelated in different parts of the landscape. One can then substantially reduce the noise by evaluating an empirical average of the gradient obtained as a sum over many random independent positions in the space of parameters to be optimized. We present an algorithm, called Averaged Gradient Descent, based on this idea and we apply it to tensor PCA, which is a very hard estimation problem. We show that Averaged Gradient Descent over-performs physical algorithms such as gradient descent and approximate message passing and matches the best algorithmic thresholds known so far, obtained by tensor unfolding and methods based on sum-of-squares.

stat.ML↗

Who is Afraid of Big Bad Minima? Analysis of Gradient-Flow in a Spiked Matrix-Tensor Model

Gradient-based algorithms are effective for many machine learning tasks, but despite ample recent effort and some progress, it often remains unclear why they work in practice in optimising high-dimensional non-convex functions and why they find good minima instead of being trapped in spurious ones. Here we present a quantitative theory explaining this behaviour in a spiked matrix-tensor model. Our framework is based on the Kac-Rice analysis of stationary points and a closed-form analysis of gradient-flow originating from statistical physics. We show that there is a well defined region of parameters where the gradient-flow algorithm finds a good global minimum despite the presence of exponentially many spurious local minima. We show that this is achieved by surfing on saddles that have strong negative direction towards the global minima, a phenomenon that is connected to a BBP-type threshold in the Hessian describing the critical points of the landscapes.

cs.LG↗