Searcharxiv⌕ Search

arXiv subjects

Stephanie E. Palmer

Publications and source records attributed to Stephanie E. Palmer.

14 recordsLinked to original sources

Optimization and variability can coexist

Many biological systems perform close to their physical limits, but promoting this optimality to a general principle seems to require implausibly fine tuning of parameters. Using examples from a wide range of systems, we show that this intuition is wrong. Near an optimum, functional performance depends on parameters in a "sloppy'' way, with some combinations of parameters being only weakly constrained. Absent any other constraints, this predicts that we should observe widely varying parameters, and we make this precise: the entropy in parameter space can be extensive even if performance on average is very close to optimal. This removes a major objection to optimization as a general principle, and rationalizes the observed variability.

q-bio.QM↗

What makes it possible to learn probability distributions in the natural world?

Organisms and algorithms learn probability distributions from previous observations, either over evolutionary time or on the fly. In the absence of regularities, estimating the underlying distribution from data would require observing each possible outcome many times. Here we show that two conditions allow us to escape this infeasible requirement. First, the mutual information between two halves of the system should be consistently sub-extensive. Second, this shared information should be compressible, so that it can be represented by a number of bits proportional to the information rather than to the entropy. Under these conditions, a distribution can be described with a number of parameters that grows linearly with system size. These conditions are borne out in natural images and in models from statistical physics, respectively.

cond-mat.stat-mech↗

Probabilistic models, compressible interactions, and neural coding

In physics we often use very simple models to describe systems with many degrees of freedom, but it is not clear why or how this success can be transferred to the more complex biological context. We consider models for the joint distribution of many variables, as with the combinations of spiking and silence in large networks of neurons. In this probabilistic framework, we argue that simple models are possible if the mutual information between two halves of the system is consistently sub--extensive, and if this shared information is compressible. These conditions are not met generically, but they are met by real world data such as natural images and the activity in a population of retinal output neurons. We introduce compression strategies that combine the information bottleneck with an iteration scheme inspired by the renormalization group, and find that the number of parameters needed to describe the distribution of joint activity scales with the square of the number of neurons, even though the interactions are not well approximated as pairwise. Our results also show that this shared information is essentially equal to the information that individual neurons carry about natural visual inputs, which has surprising implications for the neural code.

q-bio.NC↗

Multi-Relevance: Coexisting but Distinct Notions of Scale in Large Systems

Renormalization group (RG) methods are emerging as tools in biology and computer science to support the search for simplifying structure in distributions over high-dimensional spaces. We show that mixture models can be thought of as having multiple coexisting, exactly independent RG flows, each with its own notion of scale. We define this property as ``multi-relevance''. As an example, we construct a model that has two distinct notions of scale, each corresponding to the state of an unobserved categorical variable. In the regime where this latent variable can be inferred using a linear classifier, the vertex expansion approach in non-perturbative RG can be applied successfully but will give different answers depending the choice of expansion point in state space. In the regime where linear estimation of the latent state fails, we show that the vertex expansion predicts a decrease in the total number of relevant couplings from four to three and does not admit a good polynomial truncation scheme. This indicates oversimplification. One consequence of this is that principal component analysis (PCA) may be a poor choice of coarse-graining scheme in multi-relevant systems, since it imposes a notion of scale which is incorrect from the RG perspective. Taken together, our results indicate that RG and PCA can lead to oversimplification when multi-relevance is present and not accounted for.

cond-mat.stat-mech↗

Exact minimax entropy models of large-scale neuronal activity

In the brain, fine-scale correlations combine to produce macroscopic patterns of activity. However, as experiments record from larger and larger populations, we approach a fundamental bottleneck: the number of correlations one would like to include in a model grows larger than the available data. In this undersampled regime, one must focus on a sparse subset of correlations; the optimal choice contains the maximum information about patterns of activity or, equivalently, minimizes the entropy of the inferred maximum entropy model. Applying this ``minimax entropy" principle is generally intractable, but here we present an exact and scalable solution for pairwise correlations that combine to form a tree (a network without loops). Applying our method to over one thousand neurons in the mouse hippocampus, we find that the optimal tree of correlations reduces our uncertainty about the population activity by 14% (over 50 times more than a random tree). Despite containing only 0.1% of all pairwise correlations, this minimax entropy model accurately predicts the observed large-scale synchrony in neural activity and becomes even more accurate as the population grows. The inferred Ising model is almost entirely ferromagnetic (with positive interactions) and exhibits signatures of thermodynamic criticality. These results suggest that a sparse backbone of excitatory interactions may play an important role in driving collective neuronal activity.

physics.bio-ph↗

Exactly solvable statistical physics models for large neuronal populations

Maximum entropy methods provide a principled path connecting measurements of neural activity directly to statistical physics models, and this approach has been successful for populations of $N\sim 100$ neurons. As $N$ increases in new experiments, we enter an undersampled regime where we have to choose which observables should be constrained in the maximum entropy construction. The best choice is the one that provides the greatest reduction in entropy, defining a "minimax entropy" principle. This principle becomes tractable if we restrict attention to correlations among pairs of neurons that link together into a tree; we can find the best tree efficiently, and the underlying statistical physics models are exactly solved. We use this approach to analyze experiments on $N\sim 1500$ neurons in the mouse hippocampus, and show that the resulting model captures the distribution of synchronous activity in the network.

physics.bio-ph↗

Emergent scale-free networks

Many complex systems--from social and communication networks to biological networks and the Internet--are thought to exhibit scale-free structure. However, prevailing explanations rely on the constant addition of new nodes, an assumption that fails dramatically in some real-world settings. Here, we propose a model in which nodes are allowed to die, and their connections rearrange under a mixture of preferential and random attachment. With these simple dynamics, we show that networks self-organize towards scale-free structure, with a power-law exponent $γ= 1 + \frac{1}{p}$ that depends only on the proportion $p$ of preferential (rather than random) attachment. Applying our model to several real networks, we infer $p$ directly from data, and predict the relationship between network size and degree heterogeneity. Together, these results establish that realistic scale-free structure can emerge naturally in networks of constant size and density, with broad implications for the structure and function of complex systems.

nlin.AO↗

Inferring couplings in networks across order-disorder phase transitions

Statistical inference is central to many scientific endeavors, yet how it works remains unresolved. Answering this requires a quantitative understanding of the intrinsic interplay between statistical models, inference methods and data structure. To this end, we characterize the efficacy of direct coupling analysis (DCA)--a highly successful method for analyzing amino acid sequence data--in inferring pairwise interactions from samples of ferromagnetic Ising models on random graphs. Our approach allows for physically motivated exploration of qualitatively distinct data regimes separated by phase transitions. We show that inference quality depends strongly on the nature of generative models: optimal accuracy occurs at an intermediate temperature where the detrimental effects from macroscopic order and thermal noise are minimal. Importantly our results indicate that DCA does not always outperform its local-statistics-based predecessors; while DCA excels at low temperatures, it becomes inferior to simple correlation thresholding at virtually all temperatures when data are limited. Our findings offer new insights into the regime in which DCA operates so successfully and more broadly how inference interacts with data structure.

physics.bio-ph↗

Gaussian Information Bottleneck and the Non-Perturbative Renormalization Group

The renormalization group (RG) is a class of theoretical techniques used to explain the collective physics of interacting, many-body systems. It has been suggested that the RG formalism may be useful in finding and interpreting emergent low-dimensional structure in complex systems outside of the traditional physics context, such as in biology or computer science. In such contexts, one common dimensionality-reduction framework already in use is information bottleneck (IB), in which the goal is to compress an ``input'' signal $X$ while maximizing its mutual information with some stochastic ``relevance'' variable $Y$. IB has been applied in the vertebrate and invertebrate processing systems to characterize optimal encoding of the future motion of the external world. Other recent work has shown that the RG scheme for the dimer model could be ``discovered'' by a neural network attempting to solve an IB-like problem. This manuscript explores whether IB and any existing formulation of RG are formally equivalent. A class of soft-cutoff non-perturbative RG techniques are defined by families of non-deterministic coarsening maps, and hence can be formally mapped onto IB, and vice versa. For concreteness, this discussion is limited entirely to Gaussian statistics (GIB), for which IB has exact, closed-form solutions. Under this constraint, GIB has a semigroup structure, in which successive transformations remain IB-optimal. Further, the RG cutoff scheme associated with GIB can be identified. Our results suggest that IB can be used to impose a notion of ``large scale'' structure, such as biological function, on an RG procedure.

cond-mat.stat-mech↗

State Dependence of Stimulus-Induced Variability Tuning in Macaque MT

Behavioral states marked by varying levels of arousal and attention modulate some properties of cortical responses (e.g. average firing rates or pairwise correlations), yet it is not fully understood what drives these response changes and how they might affect downstream stimulus decoding. Here we show that changes in state modulate the tuning of response variance-to-mean ratios (Fano factors) in a fashion that is neither predicted by a Poisson spiking model nor changes in the mean firing rate, with a substantial effect on stimulus discriminability. We recorded motion-sensitive neurons in middle temporal cortex (MT) in two states: alert fixation and light, opioid anesthesia. Anesthesia tended to lower average spike counts, without decreasing trial-to-trial variability compared to the alert state. Under anesthesia, within-trial fluctuations in excitability were correlated over longer time scales compared to the alert state, creating supra-Poisson Fano factors. In contrast, alert-state MT neurons have higher mean firing rates and largely sub-Poisson variability that is stimulus-dependent and cannot be explained by firing rate differences alone. The absence of such stimulus-induced variability tuning in the anesthetized state suggests different sources of variability between states. A simple model explains state-dependent shifts in the distribution of observed Fano factors via a suppression in the variance of gain fluctuations in the alert state. A population model with stimulus-induced variability tuning and behaviorally constrained information-limiting correlations explores the potential enhancement in stimulus discriminability by the cortical population in the alert state.

q-bio.NC↗

Learning to make external sensory stimulus predictions using internal correlations in populations of neurons

To compensate for sensory processing delays, the visual system must make predictions to ensure timely and appropriate behaviors. Recent work has found predictive information about the stimulus in neural populations early in vision processing, starting in the retina. However, to utilize this information, cells downstream must in turn be able to read out the predictive information from the spiking activity of retinal ganglion cells. Here we investigate whether a downstream cell could learn efficient encoding of predictive information in its inputs in the absence of other instructive signals, from the correlations in the inputs themselves. We simulate learning driven by spiking activity recorded in salamander retina. We model a downstream cell as a binary neuron receiving a small group of weighted inputs and quantify the predictive information between activity in the binary neuron and future input. Input weights change according to spike timing-dependent learning rules during a training period. We characterize the readouts learned under spike timing-dependent learning rules, finding that although the fixed points of learning dynamics are not associated with absolute optimal readouts, they convey nearly all the information conveyed by the optimal readout. Moreover, we find that learned perceptrons transmit position and velocity information of a moving bar stimulus nearly as efficiently as optimal perceptrons. We conclude that predictive information is, in principle, readable from the perspective of downstream neurons in the absence of other inputs, and consequently suggests that bottom-up prediction may play an important role in sensory processing.

q-bio.NC↗

Optimal prediction and natural scene statistics in the retina

Almost all neural computations involve making predictions. Whether an organism is trying to catch prey, avoid predators, or simply move through a complex environment, the data it collects through its senses can guide its actions only to the extent that it can extract from these data information about the future state of the world. An essential aspect of the problem in all these forms is that not all features of the past carry predictive power. Since there are costs associated with representing and transmitting information, a natural hypothesis is that sensory systems have developed coding strategies that are optimized to minimize these costs, keeping only a limited number of bits of information about the past and ensuring that these bits are maximally informative about the future. Another important feature of the prediction problem is that the physics of the world is diverse enough to contain a wide range of possible statistical ensembles, yet not all motion is probable. Thus, the brain might not be a generalized predictive machine; it might have evolved to specifically solve the prediction problems most common in the natural environment. This paper reviews recent results on predictive coding and optimal predictive information in the retina and suggests approaches for quantifying prediction in response to natural motion.

q-bio.NC↗

Temporal sequences of spikes during practice code for time in a complex motor sequence

Practice of a complex motor gesture involves exploration of motor space to attain a better match to target output, but little is known about the neural code for such exploration. Here, we examine spiking in an area of the songbird brain known to contribute to modification of song output. We find that neurons in the outflow nucleus of a specialized basal ganglia- thalamocortical circuit, the lateral magnocellular nucleus of the anterior nidopallium (LMAN), code for time in the motor gesture (song) both during singing directed to a female bird (performance) and when the bird sings alone (practice). Using mutual information to quantify the correlation between temporal sequences of spikes and time in song, we find that different symbols code for time in the two singing states. While isolated spikes code for particular parts of song during performance, extended strings of spiking and silence, particularly burst events, code for time in song during practice. This temporal coding during practice can be as precise as isolated spiking during performance to a female, supporting the hypothesis that neurons in LMAN actively sample motor space, guiding song modification at local instances in time.

q-bio.NC↗

Predictive information in a sensory population

Guiding behavior requires the brain to make predictions about future sensory inputs. Here we show that efficient predictive computation starts at the earliest stages of the visual system. We estimate how much information groups of retinal ganglion cells carry about the future state of their visual inputs, and show that every cell we can observe participates in a group of cells for which this predictive information is close to the physical limit set by the statistical structure of the inputs themselves. Groups of cells in the retina also carry information about the future state of their own activity, and we show that this information can be compressed further and encoded by downstream predictor neurons, which then exhibit interesting feature selectivity. Efficient representation of predictive information is a candidate principle that can be applied at each stage of neural computation.

q-bio.NC↗