SearcharxivSearch

arXiv subjects

Máté Lengyel

Publications and source records attributed to Máté Lengyel.

12 recordsLinked to original sources

Hierarchical Successor Representation for Robust Transfer

The successor representation (SR) provides a powerful framework for decoupling predictive dynamics from rewards, enabling rapid generalisation across reward configurations. However, the classical SR is limited by its inherent policy dependence: policies change due to ongoing learning, environmental non-stationarities, and changes in task demands, making established predictive representations obsolete. Furthermore, in topologically complex environments, SRs suffer from spectral diffusion, leading to dense and overlapping features that scale poorly. Here we propose the Hierarchical Successor Representation (HSR) for overcoming these limitations. By incorporating temporal abstractions into the construction of predictive representations, HSR learns stable state features which are robust to task-induced policy changes. Applying non-negative matrix factorisation (NMF) to the HSR yields a sparse, low-rank state representation that facilitates highly sample-efficient transfer to novel tasks in multi-compartmental environments. Further analysis reveals that HSR-NMF discovers interpretable topological structures, providing a policy-agnostic hierarchical map that effectively bridges model-free optimality and model-based flexibility. Beyond providing a useful basis for task-transfer, we show that HSR's temporally extended predictive structure can also be leveraged to drive efficient exploration, effectively scaling to large, procedurally generated environments.

cs.LG

When sufficiency is insufficient: the functional information bottleneck for identifying probabilistic neural representations

The neural basis of probabilistic computations remains elusive, even amidst growing evidence that humans and other animals track their uncertainty. Recent work has proposed that probabilistic representations arise naturally in task-optimized neural networks trained without explicitly probabilistic inductive biases. However, prior work has lacked clear criteria for distinguishing probabilistic representations, those that perform transformations characteristic of probabilistic computation, from heuristic neural codes that merely reformat inputs. We propose a novel information bottleneck framework, the functional information bottleneck (fIB), that crucially evaluates a neural representation based not only on its statistical sufficiency but also on its minimality, allowing us to disambiguate heuristic from probabilistic coding. To demonstrate the power of this framework, we study a variety of task-optimized neural networks that had been suggested to develop probabilistic representations in earlier work: networks trained to perform static inference tasks (such as cue combination and coordinate transformation) or dynamic state estimation tasks (Kalman filtering). In contrast to earlier claims, our minimality requirement reveals that probabilistic representations fail to emerge in these networks: they do not develop minimal codes of Bayesian posteriors in their hidden layer activities, and instead rely on heuristic input recoding. Therefore, it remains an open question under which conditions truly probabilistic representations emerge in neural networks. More generally, our work provides a stringent framework for identifying probabilistic neural codes. Thus, it lays the foundation for systematically examining whether, how, and which posteriors are represented in neural circuits during complex decision-making tasks.

q-bio.NC

Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors

Discovering the neural mechanisms underpinning cognition is one of the grand challenges of neuroscience. However, previous approaches for building models of RNN dynamics that explain behaviour required iterative refinement of architectures and/or optimisation objectives, resulting in a piecemeal, and mostly heuristic, human-in-the-loop process. Here, we offer an alternative approach that automates the discovery of viable RNN mechanisms by explicitly training RNNs to reproduce behaviour, including the same characteristic errors and suboptimalities, that humans and animals produce in a cognitive task. Achieving this required two main innovations. First, as the amount of behavioural data that can be collected in experiments is often too limited to train RNNs, we use a non-parametric generative model of behavioural responses to produce surrogate data for training RNNs. Second, to capture all relevant statistical aspects of the data, we developed a novel diffusion model-based approach for training RNNs. To showcase the potential of our approach, we chose a visual working memory task as our test-bed, as behaviour in this task is well known to produce response distributions that are patently multimodal (due to swap errors). The resulting network dynamics correctly qualitative features of macaque neural data. Importantly, these results were not possible to obtain with more traditional approaches, i.e., when only a limited set of behavioural signatures (rather than the full richness of behavioural response distributions) were fitted, or when RNNs were trained for task optimality (instead of reproducing behaviour). Our approach also yields novel predictions about the mechanism of swap errors, which can be readily tested in experiments. These results suggest that fitting RNNs to rich patterns of behaviour provides a powerful way to automatically discover mechanisms of important cognitive functions.

q-bio.NC

Dynamical stability for dense patterns in attractor neural networks

Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neural activity that are stored in the recurrent connections between neurons, such that they become attractors of the network's dynamics. During memory recall, network dynamics thus converge toward one of these memory patterns when started from a noisy or partial cue. Therefore, memory performance critically hinges on the dynamical stability of the stored patterns. However, previous theoretical approaches only studied dynamical stability under highly restrictive conditions that do not readily apply to biological neural circuits. Here, we develop a theory of the local stability of discrete fixed points in a broad class of networks with graded neural activities and in the presence of noise. Using methods from random matrix theory, we analyze the bulk and outliers of the eigenvalue spectra of the Jacobians that characterize network dynamics around fixed points. We show that either all fixed points are stable or all of them are unstable, depending on whether their number is below a ``critical load for stability'', which is distinct from the classical critical capacity that measures the maximal number of achievable fixed points regardless of their stability. We further analyze the dependence of this critical load for stability on experimentally measurable quantities characterizing the statistics of memory patterns and the activation functions of neurons. Our analysis highlights the computational benefits of sparse-like patterns and threshold-linear activation functions and offers testable predictions for neural circuits supporting memory.

cond-mat.dis-nn

A flexible Bayesian non-parametric mixture model reveals multiple dependencies of swap errors in visual working memory

Human behavioural data in psychophysics has been used to elucidate the underlying mechanisms of many cognitive processes, such as attention, sensorimotor integration, and perceptual decision making. Visual working memory has particularly benefited from this approach: analyses of VWM errors have proven crucial for understanding VWM capacity and coding schemes, in turn constraining neural models of both. One poorly understood class of VWM errors are swap errors, whereby participants recall an uncued item from memory. Swap errors could arise from erroneous memory encoding, noisy storage, or errors at retrieval time - previous research has mostly implicated the latter two. However, these studies made strong a priori assumptions on the detailed mechanisms and/or parametric form of errors contributed by these sources. Here, we pursue a data-driven approach instead, introducing a Bayesian non-parametric mixture model of swap errors (BNS) which provides a flexible descriptive model of swapping behaviour, such that swaps are allowed to depend on both the probed and reported features of every stimulus item. We fit BNS to the trial-by-trial behaviour of human participants and show that it recapitulates the strong dependence of swaps on cue similarity in multiple datasets. Critically, BNS reveals that this dependence coexists with a non-monotonic modulation in the report feature dimension for a random dot motion direction-cued, location-reported dataset. The form of the modulation inferred by BNS opens new questions about the importance of memory encoding in causing swap errors in VWM, a distinct source to the previously suggested binding and cueing errors. Our analyses, combining qualitative comparisons of the highly interpretable BNS parameter structure with rigorous quantitative model comparison and recovery methods, show that previous interpretations of swap errors may have been incomplete.

q-bio.NC

Asymptotic scaling properties of the posterior mean and variance in the Gaussian scale mixture model

The Gaussian scale mixture model (GSM) is a simple yet powerful probabilistic generative model of natural image patches. In line with the well-established idea that sensory processing is adapted to the statistics of the natural environment, the GSM has also been considered a model of the early visual system, as a reasonable "first-order" approximation of the internal model that the primary visual cortex (V1) implements. According to this view, neural activities in V1 represent the posterior distribution under the GSM given a particular visual stimulus. Indeed, (approximate) inference under the GSM has successfully accounted for various nonlinearities in the mean (trial-average) responses of V1 neurons, as well as the dependence of (across-trial) response variability with stimulus contrast found in V1 recordings. However, previous work almost exclusively relied on numerical simulations to obtain these results. Thus, for a deeper insight into the realm of possible behaviours the GSM can (and cannot) exhibit and predict, here we present analytical derivations for the limiting behaviour of the mean and (co)variance of the GSM posterior at very low and very high contrast levels. These results should guide future work exploring neural circuit dynamics appropriate for implementing inference under the GSM.

q-bio.QM

Characterizing variability in nonlinear recurrent neuronal networks

In this note, we develop semi-analytical techniques to obtain the full correlational structure of a stochastic network of nonlinear neurons described by rate variables. Under the assumption that pairs of membrane potentials are jointly Gaussian -- which they tend to be in large networks -- we obtain deterministic equations for the temporal evolution of the mean firing rates and the noise covariance matrix that can be solved straightforwardly given the network connectivity. We also obtain spike count statistics such as Fano factors and pairwise correlations, assuming doubly-stochastic action potential firing. Importantly, our theory does not require fluctuations to be small, and works for several biologically motivated, convex single-neuron nonlinearities.

q-bio.NC

On the role of time in perceptual decision making

According to the dominant view, time in perceptual decision making is used for integrating new sensory evidence. Based on a probabilistic framework, we investigated the alternative hypothesis that time is used for gradually refining an internal estimate of uncertainty, that is to obtain an increasingly accurate approximation of the posterior distribution through collecting samples from it. In the context of a simple orientation estimation task, we analytically derived predictions of how humans should behave under the two hypotheses, and identified the across-trial correlation between error and subjective uncertainty as a proper assay to distinguish between them. Next, we developed a novel experimental paradigm that could be used to reliably measure these quantities, and tested the predictions derived from the two hypotheses. We found that in our task, humans show clear evidence that they use time mostly for probabilistic sampling and not for evidence integration. These results provide the first empirical support for iteratively improving probabilistic representations in perceptual decision making, and open the way to reinterpret the role of time in the cortical processing of complex sensory information.

q-bio.NC

The Hamiltonian brain: efficient probabilistic inference with excitatory-inhibitory neural circuit dynamics

Probabilistic inference offers a principled framework for understanding both behaviour and cortical computation. However, two basic and ubiquitous properties of cortical responses seem difficult to reconcile with probabilistic inference: neural activity displays prominent oscillations in response to constant input, and large transient changes in response to stimulus onset. Here we show that these dynamical behaviours may in fact be understood as hallmarks of the specific representation and algorithm that the cortex employs to perform probabilistic inference. We demonstrate that a particular family of probabilistic inference algorithms, Hamiltonian Monte Carlo (HMC), naturally maps onto the dynamics of excitatory-inhibitory neural networks. Specifically, we constructed a model of an excitatory-inhibitory circuit in primary visual cortex that performed HMC inference, and thus inherently gave rise to oscillations and transients. These oscillations were not mere epiphenomena but served an important functional role: speeding up inference by rapidly spanning a large volume of state space. Inference thus became an order of magnitude more efficient than in a non-oscillatory variant of the model. In addition, the network matched two specific properties of observed neural dynamics that would otherwise be difficult to account for in the context of probabilistic inference. First, the frequency of oscillations as well as the magnitude of transients increased with the contrast of the image stimulus. Second, excitation and inhibition were balanced, and inhibition lagged excitation. These results suggest a new functional role for the separation of cortical populations into excitatory and inhibitory neurons, and for the neural oscillations that emerge in such excitatory-inhibitory networks: enhancing the efficiency of cortical computations.

q-bio.NC

Fast sampling for Bayesian inference in neural circuits

Time is at a premium for recurrent network dynamics, and particularly so when they are stochastic and correlated: the quality of inference from such dynamics fundamentally depends on how fast the neural circuit generates new samples from its stationary distribution. Indeed, behavioral decisions can occur on fast time scales (~100 ms), but it is unclear what neural circuit dynamics afford sampling at such high rates. We analyzed a stochastic form of rate-based linear neuronal network dynamics with synaptic weight matrix $W$, and the dependence on $W$ of the covariance of the stationary distribution of joint firing rates. This covariance $\Sigma$ can be actively used to represent posterior uncertainty via sampling under a linear-Gaussian latent variable model. The key insight is that the mapping between $W$ and $\Sigma$ is degenerate: there are infinitely many $W$'s that lead to sampling from the same $\Sigma$ but differ greatly in the speed at which they sample. We were able to explicitly separate these extra degrees of freedom in a parametric form and thus study their effects on sampling speed. We show that previous proposals for probabilistic sampling in neural circuits correspond to using a symmetric $W$ which violates Dale's law and results in critically slow sampling, even for moderate stationary correlations. In contrast, optimizing network dynamics for speed consistently yielded asymmetric $W$'s and dynamics characterized by fast transients, such that samples of network activity became fully decorrelated over ~10 ms. Importantly, networks with separate excitatory/inhibitory populations proved to be particularly efficient samplers, and were in the balanced regime. Thus, plausible neural circuit dynamics can perform fast sampling for efficient decoding and inference.

q-bio.NC

Bayesian Active Learning for Classification and Preference Learning

Information theoretic active learning has been widely studied for probabilistic models. For simple regression an optimal myopic policy is easily tractable. However, for other tasks and with more complex models, such as classification with nonparametric models, the optimal solution is harder to compute. Current approaches make approximations to achieve tractability. We propose an approach that expresses information gain in terms of predictive entropies, and apply this method to the Gaussian Process Classifier (GPC). Our approach makes minimal approximations to the full information theoretic objective. Our experimental performance compares favourably to many popular active learning algorithms, and has equal or lower computational complexity. We compare well to decision theoretic approaches also, which are privy to more information and require much more computational time. Secondly, by developing further a reformulation of binary preference learning to a classification problem, we extend our algorithm to Gaussian Process preference learning.

stat.ML