SearcharxivSearch

arXiv subjects

John Hertz

Publications and source records attributed to John Hertz.

At least 19 recordsLinked to original sources

Glassy dynamics near the interpolation transition in deep recurrent networks

We examine learning dynamics in deep recurrent networks, focusing on the behavior near the boundary in the depth-width plane separating under- from over-parametrized networks, known as the interpolation transition. The training data are Bach chorales in 4-part harmony, and the learning is by stochastic gradient descent with a cross-entropy loss function. We find critical slowing down of the learning approaching the transition from the overparametrized side: For a given network depth, learning times to reach small training loss values appear to diverge proportional to $1/(w - w_c)$ as the width w approaches a (loss-dependent) critical value $w_c$. We identify the zero-loss limit of this value with the interpolation transition. We also study aging (the slowing down of fluctuations as the time since the beginning of learning increases). Taking a system that has been learning for a time $\tau_w$, we measure the subsequent mean-square fluctuations of the weight values at times $\tau > \tau_w$. In the underparametrized phase, we find that they are well-described by a single function of $\tau/\tau_w$. While this scaling holds approximately at short times at the transition and in the overparametrized phase, it breaks down at longer times when the training loss gets close to the lower limit imposed by the stochastic gradient descent dynamics. Both this kind of aging and the critical slowing down are also found in certain spin glass models, suggesting that those models contain the most essential features of the learning dynamics.

cond-mat.dis-nn

Stochastic activation in a genetic switch model

We study a biological autoregulation process, involving a protein that enhances its own transcription, in a parameter region where bistability would be present in the absence of fluctuations. We calculate the rate of fluctuation-induced rare transitions between locally-stable states using a path integral formulation and Master and Chapman-Kolmogorov equations. As in simpler models for rare transitions, the rate has the form of the exponential of a quantity $S_0$ (a "barrier") multiplied by a prefactor $η$. We calculate $S_0$ and $η$ first in the bursting limit (where the ratio $γ$ of the protein and mRNA lifetimes is very large). In this limit, the calculation can be done almost entirely analytically, and the results are in good agreement with simulations. For finite $γ$ numerical calculations are generally required. However, $S_0$ can be calculated analytically to first order in $1/γ$, and the result agrees well with the full numerical calculation for all $γ> 1$. Employing a method used previously on other problems, we find we can account qualitatively for the way the prefactor $η$ varies with $γ$, but its value is 15-20% higher than that inferred from simulations.

physics.bio-ph

Cumulants of Hawkes point processes

We derive explicit, closed-form expressions for the cumulant densities of a multivariate, self-exciting Hawkes point process, generalizing a result of Hawkes in his earlier work on the covariance density and Bartlett spectrum of such processes. To do this, we represent the Hawkes process in terms of a Poisson cluster process and show how the cumulant density formulas can be derived by enumerating all possible "family trees", representing complex interactions between point events. We also consider the problem of computing the integrated cumulants, characterizing the average measure of correlated activity between events of different types, and derive the relevant equations.

math.ST

Belief-Propagation and replicas for inference and learning in a kinetic Ising model with hidden spins

We propose a new algorithm for inferring the state of hidden spins and reconstructing the connections in a synchronous kinetic Ising model, given the observed history. Focusing on the case in which the hidden spins are conditionally independent of each other given the state of observable spins, we show that calculating the likelihood of the data can be simplified by introducing a set of replicated auxiliary spins. Belief Propagation (BP) and Susceptibility Propagation (SusP) can then be used to infer the states of hidden variables and learn the couplings. We study the convergence and performance of this algorithm for networks with both Gaussian-distributed and binary bonds. We also study how the algorithm behaves as the fraction of hidden nodes and the amount of data are changed, showing that it outperforms the TAP equations for reconstructing the connections.

cond-mat.dis-nn

Maximum likelihood reconstruction for Ising models with asynchronous updates

We describe how the couplings in an asynchronous kinetic Ising model can be inferred. We consider two cases, one in which we know both the spin history and the update times and one in which we only know the spin history. For the first case, we show that one can average over all possible choices of update times to obtain a learning rule that depends only on spin correlations and can also be derived from the equations of motion for the correlations. For the second case, the same rule can be derived within a further decoupling approximation. We study all methods numerically for fully asymmetric Sherrington-Kirkpatrick models, varying the data length, system size, temperature, and external field. Good convergence is observed in accordance with the theoretical expectations.

physics.data-an

The Effect of Nonstationarity on Models Inferred from Neural Data

Neurons subject to a common non-stationary input may exhibit a correlated firing behavior. Correlations in the statistics of neural spike trains also arise as the effect of interaction between neurons. Here we show that these two situations can be distinguished, with machine learning techniques, provided the data are rich enough. In order to do this, we study the problem of inferring a kinetic Ising model, stationary or nonstationary, from the available data. We apply the inference procedure to two data sets: one from salamander retinal ganglion cells and the other from a realistic computational cortical network model. We show that many aspects of the concerted activity of the salamander retinal neurons can be traced simply to the external input. A model of non-interacting neurons subject to a non-stationary external field outperforms a model with stationary input with couplings between neurons, even accounting for the differences in the number of model parameters. When couplings are added to the non-stationary model, for the retinal data, little is gained: the inferred couplings are generally not significant. Likewise, the distribution of the sizes of sets of neurons that spike simultaneously and the frequency of spike patterns as function of their rank (Zipf plots) are well-explained by an independent-neuron model with time-dependent external input, and adding connections to such a model does not offer significant improvement. For the cortical model data, robust couplings, well correlated with the real connections, can be inferred using the non-stationary model. Adding connections to this model slightly improves the agreement with the data for the probability of synchronous spikes but hardly affects the Zipf plot.

q-bio.QM

Network Inference with Hidden Units

We derive learning rules for finding the connections between units in stochastic dynamical networks from the recorded history of a ``visible'' subset of the units. We consider two models. In both of them, the visible units are binary and stochastic. In one model the ``hidden'' units are continuous-valued, with sigmoidal activation functions, and in the other they are binary and stochastic like the visible ones. We derive exact learning rules for both cases. For the stochastic case, performing the exact calculation requires, in general, repeated summations over an number of configurations that grows exponentially with the size of the system and the data length, which is not feasible for large systems. We derive a mean field theory, based on a factorized ansatz for the distribution of hidden-unit states, which offers an attractive alternative for large systems. We present the results of some numerical calculations that illustrate key features of the two models and, for the stochastic case, the exact and approximate calculations.

cond-mat.dis-nn

L$_1$ Regularization for Reconstruction of a non-equilibrium Ising Model

The couplings in a sparse asymmetric, asynchronous Ising network are reconstructed using an exact learning algorithm. L$_1$ regularization is used to remove the spurious weak connections that would otherwise be found by simply minimizing the minus likelihood of a finite data set. In order to see how L$_1$ regularization works in detail, we perform the calculation in several ways including (1) by iterative minimization of a cost function equal to minus the log likelihood of the data plus an L$_1$ penalty term, and (2) an approximate scheme based on a quadratic expansion of the cost function around its minimum. In these schemes, we track how connections are pruned as the strength of the L$_1$ penalty is increased from zero to large values. The performance of the methods for various coupling strengths is quantified using ROC curves.

stat.ME

Ising Models for Inferring Network Structure From Spike Data

Now that spike trains from many neurons can be recorded simultaneously, there is a need for methods to decode these data to learn about the networks that these neurons are part of. One approach to this problem is to adjust the parameters of a simple model network to make its spike trains resemble the data as much as possible. The connections in the model network can then give us an idea of how the real neurons that generated the data are connected and how they influence each other. In this chapter we describe how to do this for the simplest kind of model: an Ising network. We derive algorithms for finding the best model connection strengths for fitting a given data set, as well as faster approximate algorithms based on mean field theory. We test the performance of these algorithms on data from model networks and experiments.

q-bio.QM

Effect of coupling asymmetry on mean-field solutions of direct and inverse Sherrington-Kirkpatrick model

We study how the degree of symmetry in the couplings influences the performance of three mean field methods used for solving the direct and inverse problems for generalized Sherrington-Kirkpatrick models. In this context, the direct problem is predicting the potentially time-varying magnetizations. The three theories include the first and second order Plefka expansions, referred to as naive mean field (nMF) and TAP, respectively, and a mean field theory which is exact for fully asymmetric couplings. We call the last of these simply MF theory. We show that for the direct problem, nMF performs worse than the other two approximations, TAP outperforms MF when the coupling matrix is nearly symmetric, while MF works better when it is strongly asymmetric. For the inverse problem, MF performs better than both TAP and nMF, although an ad hoc adjustment of TAP can make it comparable to MF. For high temperatures the performance of TAP and MF approach each other.

cond-mat.dis-nn

Dynamical TAP equations for non-equilibrium Ising spin glasses

We derive and study dynamical TAP equations for Ising spin glasses obeying both synchronous and asynchronous dynamics using a generating functional approach. The system can have an asymmetric coupling matrix, and the external fields can be time-dependent. In the synchronously updated model, the TAP equations take the form of self consistent equations for magnetizations at time $t+1$, given the magnetizations at time $t$. In the asynchronously updated model, the TAP equations determine the time derivatives of the magnetizations at each time, again via self consistent equations, given the current values of the magnetizations. Numerical simulations suggest that the TAP equations become exact for large systems.

cond-mat.dis-nn

Statistical physics of pairwise probability models

Statistical models for describing the probability distribution over the states of biological systems are commonly used for dimensional reduction. Among these models, pairwise models are very attractive in part because they can be fit using a reasonable amount of data: knowledge of the means and correlations between pairs of elements in the system is sufficient. Not surprisingly, then, using pairwise models for studying neural data has been the focus of many studies in recent years. In this paper, we describe how tools from statistical physics can be employed for studying and using pairwise models. We build on our previous work on the subject and study the relation between different methods for fitting these models and evaluating their quality. In particular, using data from simulated cortical networks we study how the quality of various approximate methods for inferring the parameters in a pairwise model depends on the time bin chosen for binning the data. We also study the effect of the size of the time bin on the model quality itself, again using simulated data. We show that using finer time bins increases the quality of the pairwise model. We offer new ways of deriving the expressions reported in our previous work for assessing the quality of pairwise models.

q-bio.QM

The Ising Model for Neural Data: Model Quality and Approximate Methods for Extracting Functional Connectivity

We study pairwise Ising models for describing the statistics of multi-neuron spike trains, using data from a simulated cortical network. We explore efficient ways of finding the optimal couplings in these models and examine their statistical properties. To do this, we extract the optimal couplings for subsets of size up to 200 neurons, essentially exactly, using Boltzmann learning. We then study the quality of several approximate methods for finding the couplings by comparing their results with those found from Boltzmann learning. Two of these methods- inversion of the TAP equations and an approximation proposed by Sessak and Monasson- are remarkably accurate. Using these approximations for larger subsets of neurons, we find that extracting couplings using data from a subset smaller than the full network tends systematically to overestimate their magnitude. This effect is described qualitatively by infinite-range spin glass theory for the normal phase. We also show that a globally-correlated input to the neurons in the network lead to a small increase in the average coupling. However, the pair-to-pair variation of the couplings is much larger than this and reflects intrinsic properties of the network. Finally, we study the quality of these models by comparing their entropies with that of the data. We find that they perform well for small subsets of the neurons in the network, but the fit quality starts to deteriorate as the subset size grows, signalling the need to include higher order correlations to describe the statistics of large networks.

q-bio.QM

Domain wall propagation and nucleation in a metastable two-level system

We present a dynamical description and analysis of non-equilibrium transitions in the noisy one-dimensional Ginzburg-Landau equation for an extensive system based on a weak noise canonical phase space formulation of the Freidlin-Wentzel or Martin-Siggia-Rose methods. We derive propagating nonlinear domain wall or soliton solutions of the resulting canonical field equations with superimposed diffusive modes. The transition pathways are characterized by the nucleations and subsequent propagation of domain walls. We discuss the general switching scenario in terms of a dilute gas of propagating domain walls and evaluate the Arrhenius factor in terms of the associated action. We find excellent agreement with recent numerical optimization studies.

cond-mat.stat-mech

High conductance states in a mean field cortical network model

Measured responses from visual cortical neurons show that spike times tend to be correlated rather than exactly Poisson distributed. Fano factors vary and are usually greater than 1 due to the tendency of spikes being clustered into bursts. We show that this behavior emerges naturally in a balanced cortical network model with random connectivity and conductance-based synapses. We employ mean field theory with correctly colored noise to describe temporal correlations in the neuronal activity. Our results illuminate the connection between two independent experimental findings: high conductance states of cortical neurons in their natural environment, and variable non-Poissonian spike statistics with Fano factors greater than 1.

q-bio.NC

Response variability in balanced cortical networks

We study the spike statistics of neurons in a network with dynamically balanced excitation and inhibition. Our model, intended to represent a generic cortical column, comprises randomly connected excitatory and inhibitory leaky integrate-and-fire neurons, driven by excitatory input from an external population. The high connectivity permits a mean-field description in which synaptic currents can be treated as Gaussian noise, the mean and autocorrelation function of which are calculated self-consistently from the firing statistics of single model neurons. Within this description, we find that the irregularity of spike trains is controlled mainly by the strength of the synapses relative to the difference between the firing threshold and the post-firing reset level of the membrane potential. For moderately strong synapses we find spike statistics very similar to those observed in primary visual cortex.

q-bio.NC

Mean field methods for cortical network dynamics

We review the use of mean field theory for describing the dynamics of dense, randomly connected cortical circuits. For a simple network of excitatory and inhibitory leaky integrate-and-fire neurons, we can show how the firing irregularity, as measured by the Fano factor, increases with the strength of the synapses in the network and with the value to which the membrane potential is reset after a spike. Generalizing the model to include conductance-based synapses gives insight into the connection between the firing statistics and the high-conductance state observed experimentally in visual cortex. Finally, an extension of the model to describe an orientation hypercolumn provides understanding of how cortical interactions sharpen orientation tuning, in a way that is consistent with observed firing statistics.

q-bio.NC

Embedding a Native State into a Random Heteropolymer Model: The Dynamic Approach

We study a random heteropolymer model with Langevin dynamics, in the supersymmetric formulation. Employing a procedure similar to one that has been used in static calculations, we construct an ensemble in which the affinity of the system for a native state is controlled by a "selection temperature" T0. In the limit of high T0, the model reduces to a random heteropolymer, while for T0-->0 the system is forced into the native state. Within the Gaussian variational approach that we employed previously for the random heteropolymer, we explore the phases of the system for large and small T0. For large T0, the system exhibits a (dynamical) spin glass phase, like that found for the random heteropolymer, below a temperature Tg. For small T0, we find an ordered phase, characterized by a nonzero overlap with the native state, below a temperature Tn \propto 1/T0 > Tg. However, the random-globule phase remains locally stable below Tn, down to the dynamical glass transition at Tg. Thus, in this model, folding is rapid for temperatures between Tg and Tn, but below Tg the system can get trapped in conformations uncorrelated with the native state. At a lower temperature, the ordered phase can also undergo a dynamical glass transition, splitting into substates separated by large barriers.

cond-mat.stat-mech