SearcharxivSearch

arXiv subjects

Anirvan M. Sengupta

Publications and source records attributed to Anirvan M. Sengupta.

At least 19 recordsLinked to original sources

Flexible Online Representation Learning Based on Similarity Matching

Sparse high-dimensional representations are conducive to uncovering nontrivial structures in unsupervised exploration of data. Such a representation can deal with the dense connectivity in graphs relevant to community detection problems. However, sparse high-dimensional representations are capable of doing more, including manifold tiling and feature learning. Conventional algorithms optimize in the space of computationally intractable completely positive matrices or relax the problem to the space of doubly nonnegative matrices that scale with sample size in a way rendering them impractical for large data sets. Some of these methods also impose a row sum constraint, such as double stochasticity. Row sum constraints have the added advantage of being shift-invariant, in the context of manifold tiling. Constraints on the row sum of output similarity matrices require nontrivial online learning rules. Addressing these needs, we propose a versatile online biologically plausible learning algorithm capable of learning sparse shift-invariant representations, useful for clustering, manifold tiling, or sparse coding, depending on the data structure.

cs.LG

Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to the empirical distribution defined by the context. Depth refines this distribution through particle dynamics (Stage 1), while a long-range skip-connection carries the noisy input as a query for posterior inference (Stage 2), revealing distinct statistical roles for depth and attention residuals. The framework isolates a minimal setting in which the context itself induces a depth-dependent energy landscape governing in-context inference. We show that effective denoising can emerge without an explicit noise schedule: a fixed kernel bandwidth and finite integration horizon suffice, yielding a principled depth-noise relationship. We further establish a posterior-mean recovery guarantee for a class of well-behaved priors, where the empirical estimator converges to the Bayes-optimal predictor under asymptotic conditions. Connecting these dynamics to reverse-diffusion limits, our results provide a statistical interpretation of attention as in-context inference via sample-based posterior estimation, without explicit density modeling.

cs.LG

Double descent: When do neural quantum states generalize?

Neural quantum states (NQS) provide flexible and compact wavefunction parameterizations for numerical studies of quantum many-body physics. In particular, NQS aim to circumvent the exponential scaling of the Hilbert space by compressing quantum many-body wavefunctions with a tractable amount of parameters. While inspired by deep learning, it remains unclear to what extent NQS share characteristics with neural networks used for standard machine learning tasks. We demonstrate that, in a simplified supervised setting, NQS exhibit the double descent phenomenon, a key feature of modern deep learning, where generalization worsens as network size increases before improving again in an overparameterized regime. Notably, we find the second descent to occur only for network sizes much larger than the Hilbert space dimension, i.e. network sizes that are out of reach for problems of practical interest. Within our setting, this observation places typical NQS in the underparameterized regime. We also observe that the optimal network size in the underparameterized regime depends on the number of unique training samples. While the double descent phenomenon does indeed translate to the NQS setting, potential practical consequences of our findings point more towards the need for symmetry-aware, physics-informed architecture design, rather than directly adopting machine learning heuristics.

cond-mat.dis-nn

A Network of Biologically Inspired Rectified Spectral Units (ReSUs) Learns Hierarchical Features Without Error Backpropagation

We introduce a biologically inspired, multilayer neural architecture composed of Rectified Spectral Units (ReSUs). Each ReSU projects a recent window of its input history onto a canonical direction obtained via canonical correlation analysis (CCA) of previously observed past-future input pairs, and then rectifies either its positive or negative component. By encoding canonical directions in synaptic weights and temporal filters, ReSUs implement a local, self-supervised algorithm for progressively constructing increasingly complex features. To evaluate both computational power and biological fidelity, we trained a two-layer ReSU network in a self-supervised regime on translating natural scenes. First-layer units, each driven by a single pixel, developed temporal filters resembling those of Drosophila post-photoreceptor neurons (L1/L2 and L3), including their empirically observed adaptation to signal-to-noise ratio (SNR). Second-layer units, which pooled spatially over the first layer, became direction-selective -- analogous to T4 motion-detecting cells -- with learned synaptic weight patterns approximating those derived from connectomic reconstructions. Together, these results suggest that ReSUs offer (i) a principled framework for modeling sensory circuits and (ii) a biologically grounded, backpropagation-free paradigm for constructing deep self-supervised neural networks.

q-bio.NC

Neurons as Detectors of Coherent Sets in Sensory Dynamics

We model sensory streams as observations from high-dimensional stochastic dynamical systems and conceptualize sensory neurons as self-supervised learners of compact representations of such dynamics. From prior experience, neurons learn coherent sets-regions of stimulus state space whose trajectories evolve cohesively over finite times-and assign membership indices to new stimuli. Coherent sets are identified via spectral clustering of the stochastic Koopman operator (SKO), where the sign pattern of a subdominant singular function partitions the state space into minimally coupled regions. For multivariate Ornstein-Uhlenbeck processes, this singular function reduces to a linear projection onto the dominant singular vector of the whitened state-transition matrix. Encoding this singular vector as a receptive field enables neurons to compute membership indices via the projection sign in a biologically plausible manner. Each neuron detects either a predictive coherent set (stimuli with common futures) or a retrospective coherent set (stimuli with common pasts), suggesting a functional dichotomy among neurons. Since neurons lack access to explicit dynamical equations, the requisite singular vectors must be estimated directly from data, for example, via past-future canonical correlation analysis on lag-vector representations-an approach that naturally extends to nonlinear dynamics. This framework provides a novel account of neuronal temporal filtering, the ubiquity of rectification in neural responses, and known functional dichotomies. Coherent-set clustering thus emerges as a fundamental computation underlying sensory processing and transferable to bio-inspired artificial systems.

q-bio.NC

Learning interactions between Rydberg atoms

Quantum simulators have the potential to solve quantum many-body problems that are beyond the reach of classical computers, especially when they feature long-range entanglement. To fulfill their prospects, quantum simulators must be fully controllable, allowing for precise tuning of the microscopic physical parameters that define their implementation. We consider Rydberg-atom arrays, a promising platform for quantum simulations. Experimental control of such arrays is limited by the imprecision on the optical tweezers positions when assembling the array, hence introducing uncertainties in the simulated Hamiltonian. In this work, we introduce a scalable approach to Hamiltonian learning using graph neural networks (GNNs). We employ the Density Matrix Renormalization Group (DMRG) to generate ground-state snapshots of the transverse field Ising model realized by the array, for many realizations of the Hamiltonian parameters. Correlation functions reconstructed from these snapshots serve as input data to carry out the training. We demonstrate that our GNN model has a remarkable capacity to extrapolate beyond its training domain, both regarding the size and the shape of the system, yielding an accurate determination of the Hamiltonian parameters with a minimal set of measurements. We prove a theorem establishing a bijective correspondence between the correlation functions and the interaction parameters in the Hamiltonian, which provides a theoretical foundation to our learning algorithm. Our work could open the road to feedback control of the positions of the optical tweezers, hence providing a decisive improvement of analog quantum simulators.

quant-ph

In-context denoising with one-layer transformers: connections between attention and associative memory retrieval

We introduce in-context denoising, a task that refines the connection between attention-based architectures and dense associative memory (DAM) networks, also known as modern Hopfield networks. Using a Bayesian framework, we show theoretically and empirically that certain restricted denoising problems can be solved optimally even by a single-layer transformer. We demonstrate that a trained attention layer processes each denoising prompt by performing a single gradient descent update on a context-aware DAM energy landscape, where context tokens serve as associative memories and the query token acts as an initial state. This one-step update yields better solutions than exact retrieval of either a context token or a spurious local minimum, providing a concrete example of DAM networks extending beyond the standard retrieval paradigm. Overall, this work solidifies the link between associative memory and attention mechanisms first identified by Ramsauer et al., and demonstrates the relevance of associative memory models in the study of in-context learning.

cs.LG

Deep Learning Based Superconductivity: Prediction and Experimental Tests

The discovery of novel superconducting materials is a longstanding challenge in materials science, with a wealth of potential for applications in energy, transportation, and computing. Recent advances in artificial intelligence (AI) have enabled expediting the search for new materials by efficiently utilizing vast materials databases. In this study, we developed an approach based on deep learning (DL) to predict new superconducting materials. We have synthesized a compound derived from our DL network and confirmed its superconducting properties in agreement with our prediction. Our approach is also compared to previous work based on random forests (RFs). In particular, RFs require knowledge of the chemical properties of the compound, while our neural net inputs depend solely on the chemical composition. With the help of hints from our network, we discover a new ternary compound $\textrm{Mo}_{20} \textrm{Re}_{6} \textrm{Si}_{4}$, which becomes superconducting below 5.4 K. We further discuss the existing limitations and challenges associated with using AI to predict and, along with potential future research directions.

cs.LG

Compressing the two-particle Green's function using wavelets: Theory and application to the Hubbard atom

Precise algorithms capable of providing controlled solutions in the presence of strong interactions are transforming the landscape of quantum many-body physics. Particularly exciting breakthroughs are enabling the computation of non-zero temperature correlation functions. However, computational challenges arise due to constraints in resources and memory limitations, especially in scenarios involving complex Green's functions and lattice effects. Leveraging the principles of signal processing and data compression, this paper explores the wavelet decomposition as a versatile and efficient method for obtaining compact and resource-efficient representations of the many-body theory of interacting systems. The effectiveness of the wavelet decomposition is illustrated through its application to the representation of generalized susceptibilities and self-energies in a prototypical interacting fermionic system, namely the Hubbard model at half-filling in its atomic limit. These results are the first proof-of-principle application of the wavelet compression within the realm of many-body physics and demonstrate the potential of this wavelet-based compression scheme for understanding the physics of correlated electron systems.

cond-mat.str-el

Fast Scrambling at the Boundary

Many-body systems which saturate the quantum bound on chaos are attracting interest across a wide range of fields. Notable examples include the Sachdev-Ye-Kitaev model and its variations, all characterised by some form or randomness and all to all couplings. Here we study many-body quantum chaos in a quantum impurity model showing Non-Fermi-Liquid physics, the overscreened multichannel $SU(N)$ Kondo model. We compute exactly the low-temperature behavior of the out-of time order correlator in the limit of large $N$ and large number of channels $K$, at fixed ratio $γ=K/N$. Due to strong correlations at the impurity site the spin fractionalizes in auxiliary fermions and bosons. We show that all the degrees of freedom of our theory acquire a Lyapunov exponent which is linear in temperature as $T\rightarrow 0$, with a prefactor that depends on $γ$. Remarkably, for $N=K$ the impurity spin displays maximal chaos, while bosons and fermions only get up to half of the maximal Lyapunov exponent. Our results highlights two new features: a non-disordered model which is maximally chaotic due to strong correlations at its boundary and a fractionalization of quantum chaos.

cond-mat.str-el

Machine learning-based compression of quantum many body physics: PCA and autoencoder representation of the vertex function

Characterizing complex many-body phases of matter has been a central question in quantum physics for decades. Numerical methods built around approximations of the renormalization group (RG) flow equations have offered reliable and systematically improvable answers to the initial question -- what simple physics drives quantum order and disorder? The flow equations are a very high dimensional set of coupled nonlinear equations whose solution is the two particle vertex function, a function of three continuous momenta that describes particle-particle scattering and encodes much of the low energy physics including whether the system exhibits various forms of long ranged order. In this work, we take a simple and interpretable data-driven approach to the open question of compressing the two-particle vertex. We use PCA and an autoencoder neural network to derive compact, low-dimensional representations of underlying physics for the case of interacting fermions on a lattice. We quantify errors in the representations by multiple metrics and show that a simple linear PCA offers more physical insight and better out-of-distribution (zero-shot) generalization than the nominally more expressive nonlinear models. Even with a modest number of principal components (10 - 20), we find excellent reconstruction of vertex functions across the phase diagram. This result suggests that many other many-body functions may be similarly compressible, potentially allowing for efficient computation of observables. Finally, we identify principal component subspaces that are shared between known phases, offering new physical insight.

cond-mat.str-el

Learning the eigenstructure of quantum dynamics using classical shadows

Learning dynamics from repeated observation of the time evolution of an open quantum system, namely, the problem of quantum process tomography is an important task. This task is difficult in general, but, with some additional constraints could be tractable. This motivates us to look at the problem of Lindblad operator discovery from observations. We point out that for moderate size Hilbert spaces, low Kraus rank of the channel, and short time steps, the eigenvalues of the Choi matrix corresponding to the channel have a special structure. We use the least-square method for the estimation of a channel where, for fixed inputs, we estimate the outputs by classical shadows. The resultant noisy estimate of the channel can then be denoised by diagonalizing the nominal Choi matrix, truncating some eigenvalues, and altering it to a genuine Choi matrix. This processed Choi matrix is then compared to the original one. We see that as the number of samples increases, our reconstruction becomes more accurate. We also use tools from random matrix theory to understand the effect of estimation noise in the eigenspectrum of the estimated Choi matrix.

quant-ph

Normative framework for deriving neural networks with multi-compartmental neurons and non-Hebbian plasticity

An established normative approach for understanding the algorithmic basis of neural computation is to derive online algorithms from principled computational objectives and evaluate their compatibility with anatomical and physiological observations. Similarity matching objectives have served as successful starting points for deriving online algorithms that map onto neural networks (NNs) with point neurons and Hebbian/anti-Hebbian plasticity. These NN models account for many anatomical and physiological observations; however, the objectives have limited computational power and the derived NNs do not explain multi-compartmental neuronal structures and non-Hebbian forms of plasticity that are prevalent throughout the brain. In this article, we unify and generalize recent extensions of the similarity matching approach to address more complex objectives, including a large class of unsupervised and self-supervised learning tasks that can be formulated as symmetric generalized eigenvalue problems or nonnegative matrix factorization problems. Interestingly, the online algorithms derived from these objectives naturally map onto NNs with multi-compartmental neurons and local, non-Hebbian learning rules. Therefore, this unified extension of the similarity matching approach provides a normative framework that facilitates understanding multi-compartmental neuronal structures and non-Hebbian plasticity found throughout the brain.

q-bio.NC

Unlocking the Potential of Similarity Matching: Scalability, Supervision and Pre-training

While effective, the backpropagation (BP) algorithm exhibits limitations in terms of biological plausibility, computational cost, and suitability for online learning. As a result, there has been a growing interest in developing alternative biologically plausible learning approaches that rely on local learning rules. This study focuses on the primarily unsupervised similarity matching (SM) framework, which aligns with observed mechanisms in biological systems and offers online, localized, and biologically plausible algorithms. i) To scale SM to large datasets, we propose an implementation of Convolutional Nonnegative SM using PyTorch. ii) We introduce a localized supervised SM objective reminiscent of canonical correlation analysis, facilitating stacking SM layers. iii) We leverage the PyTorch implementation for pre-training architectures such as LeNet and compare the evaluation of features against BP-trained models. This work combines biologically plausible algorithms with computational efficiency opening multiple avenues for further explorations.

cs.NE

Duality Principle and Biologically Plausible Learning: Connecting the Representer Theorem and Hebbian Learning

A normative approach called Similarity Matching was recently introduced for deriving and understanding the algorithmic basis of neural computation focused on unsupervised problems. It involves deriving algorithms from computational objectives and evaluating their compatibility with anatomical and physiological observations. In particular, it introduces neural architectures by considering dual alternatives instead of primal formulations of popular models such as PCA. However, its connection to the Representer theorem remains unexplored. In this work, we propose to use teachings from this approach to explore supervised learning algorithms and clarify the notion of Hebbian learning. We examine regularized supervised learning and elucidate the emergence of neural architecture and additive versus multiplicative update rules. In this work, we focus not on developing new algorithms but on showing that the Representer theorem offers the perfect lens to study biologically plausible learning algorithms. We argue that many past and current advancements in the field rely on some form of dual formulation to introduce biological plausibility. In short, as long as a dual formulation exists, it is possible to derive biologically plausible algorithms. Our work sheds light on the pivotal role of the Representer theorem in advancing our comprehension of neural computation.

cs.NE

Deep Learning the Functional Renormalization Group

We perform a data-driven dimensionality reduction of the scale-dependent 4-point vertex function characterizing the functional Renormalization Group (fRG) flow for the widely studied two-dimensional $t - t'$ Hubbard model on the square lattice. We demonstrate that a deep learning architecture based on a Neural Ordinary Differential Equation solver in a low-dimensional latent space efficiently learns the fRG dynamics that delineates the various magnetic and $d$-wave superconducting regimes of the Hubbard model. We further present a Dynamic Mode Decomposition analysis that confirms that a small number of modes are indeed sufficient to capture the fRG dynamics. Our work demonstrates the possibility of using artificial intelligence to extract compact representations of the 4-point vertex functions for correlated electrons, a goal of utmost importance for the success of cutting-edge quantum field theoretical methods for tackling the many-electron problem.

cond-mat.str-el

Constrained Predictive Coding as a Biologically Plausible Model of the Cortical Hierarchy

Predictive coding has emerged as an influential normative model of neural computation, with numerous extensions and applications. As such, much effort has been put into mapping PC faithfully onto the cortex, but there are issues that remain unresolved or controversial. In particular, current implementations often involve separate value and error neurons and require symmetric forward and backward weights across different brain regions. These features have not been experimentally confirmed. In this work, we show that the PC framework in the linear regime can be modified to map faithfully onto the cortical hierarchy in a manner compatible with empirical observations. By employing a disentangling-inspired constraint on hidden-layer neural activities, we derive an upper bound for the PC objective. Optimization of this upper bound leads to an algorithm that shows the same performance as the original objective and maps onto a biologically plausible network. The units of this network can be interpreted as multi-compartmental neurons with non-Hebbian learning rules, with a remarkable resemblance to recent experimental findings. There exist prior models which also capture these features, but they are phenomenological, while our work is a normative derivation. The network we derive does not involve one-to-one connectivity or signal multiplexing, which the phenomenological models required, indicating that these features are not necessary for learning in the cortex. The normative nature of our algorithm in the simplified linear case allows us to prove interesting properties of the framework and analytically understand the computational role of our network's components. The parameters of our network have natural interpretations as physiological quantities in a multi-compartmental model of pyramidal neurons, providing a concrete link between PC and experimental measurements carried out in the cortex.

q-bio.NC

Neural optimal feedback control with local learning rules

A major problem in motor control is understanding how the brain plans and executes proper movements in the face of delayed and noisy stimuli. A prominent framework for addressing such control problems is Optimal Feedback Control (OFC). OFC generates control actions that optimize behaviorally relevant criteria by integrating noisy sensory stimuli and the predictions of an internal model using the Kalman filter or its extensions. However, a satisfactory neural model of Kalman filtering and control is lacking because existing proposals have the following limitations: not considering the delay of sensory feedback, training in alternating phases, and requiring knowledge of the noise covariance matrices, as well as that of systems dynamics. Moreover, the majority of these studies considered Kalman filtering in isolation, and not jointly with control. To address these shortcomings, we introduce a novel online algorithm which combines adaptive Kalman filtering with a model free control approach (i.e., policy gradient algorithm). We implement this algorithm in a biologically plausible neural network with local synaptic plasticity rules. This network performs system identification and Kalman filtering, without the need for multiple phases with distinct update rules or the knowledge of the noise covariances. It can perform state estimation with delayed sensory feedback, with the help of an internal model. It learns the control policy without requiring any knowledge of the dynamics, thus avoiding the need for weight transport. In this way, our implementation of OFC solves the credit assignment problem needed to produce the appropriate sensory-motor control in the presence of stimulus delay.

q-bio.NC