SearcharxivSearch

arXiv subjects

Pierfrancesco Urbani

Publications and source records attributed to Pierfrancesco Urbani.

At least 19 recordsLinked to original sources

Learning limit cycles via Hebbian synaptic plasticity

We investigate high-dimensional, non-linear dynamical systems when exposed to incoherent periodic inputs and Hebbian-like synaptic plasticity. Our findings reveal a striking phenomenon: depending on the interplay between the strength of the periodic drive and synaptic plasticity, the system's phase diagram can give rise to a region where, once both inputs are removed, the collective dynamics spontaneously settles into a limit cycle. This suggests that periodic drives can imprint lasting rhythmic patterns into the network through plasticity, effectively teaching it to oscillate on its own. Numerical simulations on finite size systems show that the limit cycle phase can be easily detected on single-sample trajectories, while averaged curves are affected by strong finite size effects due to sample-to-sample fluctuations of the period of the limit cycles.

cond-mat.dis-nn

Theory of learning of high-dimensional controlled non-linear dynamical systems (I): models and methods

Neural ordinary differential equations (neural ODEs) have rapidly gained prominence as a powerful and unifying framework for conceptualizing artificial neural networks, elegantly connecting the continuous-time modeling of dynamical systems with the discrete, data-driven paradigm of modern deep learning. Beyond their practical advantages they offer fresh theoretical insights into the training and generalization properties of neural networks. The distinctive feature of this framework is its dual dynamical nature: inference dynamics, which govern the ODE evolution during forward computation, and training dynamics, which control the optimization of model parameters. This makes neural ODEs a particularly well-suited theoretical framework for studying a large variety of settings such as multi-layer neural networks (ResNets for example), autoregressive models (with next-token generation dynamics), generative models, and recurrent neural networks in theoretical neuroscience. In this work, we introduce a theoretically grounded class of models for studying neural ODEs trained via online stochastic gradient descent. We solve the training dynamics of these models via dynamical mean field theory and derive learning curves in the high-dimensional limit.

cond-mat.dis-nn

Chaos in high-dimensional dynamical systems with tunable non-reciprocity

High-dimensional dynamical systems of interacting degrees of freedom are ubiquitous in the study of complex systems. When the directed interactions are totally uncorrelated, sufficiently strong and non-linear, many of these systems exhibit a chaotic attractor characterized by a positive maximal Lyapunov exponent (MLE). On the contrary, when the interactions are completely symmetric, the dynamics takes the form of a gradient descent on a carefully defined cost function, and it exhibits slow dynamics and aging. In this work, we consider the intermediate case in which the interactions are partially symmetric, with a parameter α tuning the degree of non-reciprocity. We show that for any value of α for which the corresponding system has non-reciprocal interactions, the dynamics lands on a chaotic attractor. Correspondingly, the MLE is a non-monotonous function of the degree of non-reciprocity. This implies that conservative forcing deriving from the gradient field of a rough energy landscape can make the system more chaotic.

cond-mat.dis-nn

High-dimensional dynamical systems: co-existence of attractors, phase transitions, maximal Lyapunov exponent and response to periodic drive

We study the dynamical properties of a broad class of high-dimensional random dynamical systems exhibiting chaotic as well as fixed point and periodic attractors. We consider cases in which attractors can co-exists in some regions of the phase diagrams and we characterize their nature by computing the maximal Lyapunov exponent. For a specific choice of the dynamical system we show that this quantity can be computed explicitly in the whole chaotic phase due to an underlying integrability of a properly defined Schrödinger problem. Furthermore, we consider the response of this dynamical systems to periodic perturbations. We show that these dynamical systems act as filters in the frequency-amplitude spectrum of the periodic forcing: only in some regions of the frequency-amplitude plane the periodic forcing leads to a synchronization of the dynamics. All in all, the results that we present mirror closely the ones observed in the past forty years in the study of standard models of random recurrent neural networks. However, the dynamical systems that we consider are easier to study and we believe that this may be an advantage if one wants to go beyond random dynamical systems and consider specific training strategies.

cond-mat.dis-nn

Non-reciprocal interactions and high-dimensional chaos: comparing dynamics and statistics of equilibria in a solvable class of models

We investigate a model of high-dimensional dynamical variables with all-to-all interactions that are random and non-reciprocal. We characterize its phase diagram and show that the model can exhibit chaotic dynamics. We show that the equations describing the system's dynamics exhibit a number of equilibria that is exponentially large in the dimensionality of the system, and these equilibria are all linearly unstable in the chaotic phase. Solving the effective equations governing the dynamics in the infinite-dimensional limit, we determine the typical properties (magnetization, overlap) of the configurations belonging to the attractor manifold. We show that these properties cannot be inferred from those of the equilibria, challenging the expectation that chaos can be understood purely in terms of the numerous unstable equilibria of the dynamical equations. We discuss the dependence of this scenario on the strength of non-reciprocity in the interactions. These results are obtained through a combination of analytical methods such as Dynamical Mean-Field Theory and the Kac-Rice formalism.

cond-mat.dis-nn

Dynamical Decoupling of Generalization and Overfitting in Large Two-Layer Networks

Understanding the inductive bias and generalization properties of large overparametrized machine learning models requires to characterize the dynamics of the training algorithm. We study the learning dynamics of large two-layer neural networks via dynamical mean field theory, a well established technique of non-equilibrium statistical physics. We show that, for large network width $m$, and large number of samples per input dimension $n/d$, the training dynamics exhibits a separation of timescales which implies: $(i)$~The emergence of a slow time scale associated with the growth in Gaussian/Rademacher complexity of the network; $(ii)$~Inductive bias towards small complexity if the initialization has small enough complexity; $(iii)$~A dynamical decoupling between feature learning and overfitting regimes; $(iv)$~A non-monotone behavior of the test error, associated `feature unlearning' regime at large times.

stat.ML

Generative modeling through internal high-dimensional chaotic activity

Generative modeling aims at producing new datapoints whose statistical properties resemble the ones in a training dataset. In recent years, there has been a burst of machine learning techniques and settings that can achieve this goal with remarkable performances. In most of these settings, one uses the training dataset in conjunction with noise, which is added as a source of statistical variability and is essential for the generative task. Here, we explore the idea of using internal chaotic dynamics in high-dimensional chaotic systems as a way to generate new datapoints from a training dataset. We show that simple learning rules can achieve this goal within a set of vanilla architectures and characterize the quality of the generated datapoints through standard accuracy measures.

cs.LG

Statistical physics of complex systems: glasses, spin glasses, continuous constraint satisfaction problems, high-dimensional inference and neural networks

The purpose of this manuscript is to review my recent activity on three main research topics. The first concerns the nature of low temperature amorphous solids and their relation with the spin glass transition in a magnetic field. This is the subject of the first chapter where I discuss a new model, the KHGPS model, which allows to make some progress. In the second chapter I review a second research line that concerns the study of the rigidity/jamming transitions in particle system models and their relation to constraint satisfaction and optimization problems in high dimension. Finally in the last chapter I review my activity on the problem of the dynamics of learning algorithms in high-dimensional inference and supervised learning problems.

cond-mat.dis-nn

Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape

Stochastic Gradient Descent (SGD) is an out-of-equilibrium algorithm used extensively to train artificial neural networks. However very little is known on to what extent SGD is crucial for to the success of this technology and, in particular, how much it is effective in optimizing high-dimensional non-convex cost functions as compared to other optimization algorithms such as Gradient Descent (GD). In this work we leverage dynamical mean field theory to benchmark its performances in the high-dimensional limit. To do that, we consider the problem of recovering a hidden high-dimensional non-linearly encrypted signal, a prototype high-dimensional non-convex hard optimization problem. We compare the performances of SGD to GD and we show that SGD largely outperforms GD for sufficiently small batch sizes. In particular, a power law fit of the relaxation time of these algorithms shows that the recovery threshold for SGD with small batch size is smaller than the corresponding one of GD.

cs.LG

Statistical physics of learning in high-dimensional chaotic systems

In many complex systems, elementary units live in a chaotic environment and need to adapt their strategies to perform a task, by extracting information from the environment and controlling the feedback loop on it. One of the main example of systems of this kind is provided by recurrent neural networks. In this case, recurrent connections between neurons drive chaotic behavior and when learning takes place, the response of the system to a perturbation should take into account also its feedback on the dynamics of the network itself. In this work, we consider an abstract model of a high-dimensional chaotic system as a paradigmatic model and study its dynamics. We study the model under two particular settings: Hebbian driving and FORCE training. In the first case, we show that Hebbian driving can be used to tune the level of chaos in the dynamics and this reproduces some results recently obtained in the study of more biologically realistic models of recurrent neural networks. In the latter case, we show that the dynamical system can be trained to reproduce simple periodic functions. To do this, we consider the FORCE algorithm -- originally developed to train recurrent neural networks -- and adapt it to our high-dimensional chaotic system. We show that this algorithm drives the dynamics close to an asymptotic attractor the larger the training time. All our results are valid in the thermodynamic limit thanks to an exact analysis of the dynamics through dynamical mean field theory.

cond-mat.dis-nn

Dynamical mean field theory for models of confluent tissues and beyond

We consider a recently proposed model to understand the rigidity transition in confluent tissues and we derive the dynamical mean field theory (DMFT) equations that describes several types of dynamics of the model in the thermodynamic limit: gradient descent, thermal Langevin noise and active drive. In particular we focus on gradient descent dynamics and we integrate numerically the corresponding DMFT equations. In this case we show that gradient descent is blind to the zero temperature replica symmetry breaking (RSB) transition point. This means that, even if the Gibbs measure in the zero temperature limit displays RSB, this algorithm is able to find its way to a zero energy configuration. We include a discussion on possible extensions of the DMFT derivation to study problems rooted in high-dimensional regression and optimization via the square loss function.

cond-mat.dis-nn

A continuous constraint satisfaction problem for the rigidity transition in confluent tissues

Models of confluent tissues are built out of tessellations of the space (both in two and three dimensions) in which the cost function is constructed in such a way that individual cells try to optimize their volume and surface in order to reach a target shape. At zero temperature, many of these models exhibit a rigidity transition that separates two phases: a liquid phase and a solid (glassy) phase. This phenomenology is now well established but the theoretical understanding is still not complete. In this work we consider an exactly soluble mean field model for the rigidity transition which is based on an abstract mapping. We replace volume and surface functions by random non-linear functions of a large number of degrees of freedom forced to be on a compact phase space. We then seek for a configuration of the degrees of freedom such that these random non-linear functions all attain the same value. This target value is a control parameter and plays the role of the target cell shape in biological tissue models. Therefore we map the microscopic models of cells to a random continuous constraint satisfaction problem (CCSP) with equality constraints. We argue that at zero temperature, the rigidity transition corresponds to the satisfiability transition of the problem. We also characterize both the satisfiable (SAT) and unsatisfiable (UNSAT) phase. In the SAT phase, before reaching the rigidity transition, the zero temperature SAT landscape undergoes an RSB/ergodicity breaking transition of the same type as the Gardner transition in amorphous solids. By solving the RSB equations we compute the SAT/UNSAT threshold and the critical behavior around it. In the UNSAT phase we also compute the average shape index as a function of the target one and we compare the thermodynamical solution of the model with the results of the numerical greedy minimization of the corresponding cost function.

cond-mat.dis-nn

Quantum exploration of high-dimensional canyon landscapes

Canyon landscapes in high dimension can be described as manifolds of small, but extensive dimension, immersed in a higher dimensional ambient space and characterized by a zero potential energy on the manifold. Here we consider the problem of a quantum particle exploring a prototype of a high-dimensional random canyon landscape. We characterize the thermal partition function and show that around the point where the classical phase space has a satisfiability transition so that zero potential energy canyons disappear, moderate quantum fluctuations have a deleterious effect and induce glassy phases at temperature where classical thermal fluctuations alone would thermalize the system. Surprisingly we show that even when, classically, diffusion is expected to be unbounded in space, the interplay between quantum fluctuations and the randomness of the canyon landscape conspire to have a confining effect.

cond-mat.dis-nn

Low temperature amorphous solids: mean field theory and beyond

Amorphous solids at low temperature display unusual features which have been escaped a clear and unified comprehension. In recent years a mean field theory of amorphous solids constructed in the limit of infinite spatial dimensions has been proposed. I will briefly review what is the outcome of this theory focusing the low temperature phase and discuss some perspectives to go beyond the mean field limit.

cond-mat.dis-nn

Field theory for zero temperature soft anharmonic spin glasses in a field

We introduce a finite dimensional anharmonic soft spin glass in a field and show how it allows the construction a field theory at zero temperature and the corresponding loop expansion. The mean field level of the model coincides with a recently introduced fully connected model, the KHGPS model, and it has a spin glass transition in a field at zero temperature driven by the appearance of pseudogapped non-linear excitations. We analyze the zero temperature limit of the theory and the behavior of the bare masses and couplings on approaching the mean field zero temperature critical point. Focusing on the so called replicon sector of the field theory, we show that the bare mass corresponding to fluctuations in this sector is strictly positive at the transition in a certain region of control parameter space. At the same time the two relevant cubic coupling constants $g_1$ and $g_2$ show a non-analytic behavior in their bare values: approaching the critical point at zero temperature, $g_1\to \infty$ while $g_2\propto T$ with a prefactor diverging at the transition. Along the same lines we also develop the field theory to study the density of states of the model in finite dimension. We show that in the mean field limit the density of states converges to the one of the KHGPS model. However the construction allows a treatment of finite dimensional effects in perturbation theory.

cond-mat.dis-nn

The effective noise of Stochastic Gradient Descent

Stochastic Gradient Descent (SGD) is the workhorse algorithm of deep learning technology. At each step of the training phase, a mini batch of samples is drawn from the training dataset and the weights of the neural network are adjusted according to the performance on this specific subset of examples. The mini-batch sampling procedure introduces a stochastic dynamics to the gradient descent, with a non-trivial state-dependent noise. We characterize the stochasticity of SGD and a recently-introduced variant, \emph{persistent} SGD, in a prototypical neural network model. In the under-parametrized regime, where the final training error is positive, the SGD dynamics reaches a stationary state and we define an effective temperature from the fluctuation-dissipation theorem, computed from dynamical mean-field theory. We use the effective temperature to quantify the magnitude of the SGD noise as a function of the problem parameters. In the over-parametrized regime, where the training error vanishes, we measure the noise magnitude of SGD by computing the average distance between two replicas of the system with the same initialization and two different realizations of SGD noise. We find that the two noise measures behave similarly as a function of the problem parameters. Moreover, we observe that noisier algorithms lead to wider decision boundaries of the corresponding constraint satisfaction problem.

cond-mat.dis-nn

High dimensional optimization under non-convex excluded volume constraints

We consider high dimensional random optimization problems where the dynamical variables are subjected to non-convex excluded volume constraints. We focus on the case in which the cost function is a simple quadratic cost and the excluded volume constraints are modeled by a perceptron constraint satisfaction problem. We show that depending on the density of constraints, one can have different situations. If the number of constraints is small, one typically has a phase where the ground state of the cost function is unique and sits on the boundary of the island of configurations allowed by the constraints. In this case, there is a hypostatic number of marginally satisfied constraints. If the number of constraints is increased one enters a glassy phase where the cost function has many local minima sitting again on the boundary of the regions of allowed configurations. At the phase transition point, the total number of marginally satisfied constraints becomes equal to the number of degrees of freedom in the problem and therefore we say that these minima are isostatic. We conjecture that by increasing further the constraints the system stays isostatic up to the point where the volume of available phase space shrinks to zero. We derive our results using the replica method and we also analyze a dynamical algorithm, the Karush-Kuhn-Tucker algorithm, through dynamical mean-field theory and we show how to recover the results of the replica approach in the replica symmetric phase.

cond-mat.dis-nn

Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification

We analyze in a closed form the learning dynamics of stochastic gradient descent (SGD) for a single-layer neural network classifying a high-dimensional Gaussian mixture where each cluster is assigned one of two labels. This problem provides a prototype of a non-convex loss landscape with interpolating regimes and a large generalization gap. We define a particular stochastic process for which SGD can be extended to a continuous-time limit that we call stochastic gradient flow. In the full-batch limit, we recover the standard gradient flow. We apply dynamical mean-field theory from statistical physics to track the dynamics of the algorithm in the high-dimensional limit via a self-consistent stochastic process. We explore the performance of the algorithm as a function of the control parameters shedding light on how it navigates the loss landscape.

cs.LG