Searcharxiv⌕ Search

arXiv subjects

Pankaj Mehta

Publications and source records attributed to Pankaj Mehta.

At least 37 records · Page 2Linked to original sources

Bias-variance decomposition of overparameterized regression with random linear features

In classical statistics, the bias-variance trade-off describes how varying a model's complexity (e.g., number of fit parameters) affects its ability to make accurate predictions. According to this trade-off, optimal performance is achieved when a model is expressive enough to capture trends in the data, yet not so complex that it overfits idiosyncratic features of the training data. Recently, it has become clear that this classic understanding of the bias-variance must be fundamentally revisited in light of the incredible predictive performance of "overparameterized models" -- models that avoid overfitting even when the number of fit parameters is large enough to perfectly fit the training data. Here, we present results for one of the simplest examples of an overparameterized model: regression with random linear features (i.e. a two-layer neural network with a linear activation function). Using the zero-temperature cavity method, we derive analytic expressions for the training error, test error, bias, and variance. We show that the linear random features model exhibits three phase transitions: two different transitions to an interpolation regime where the training error is zero, along with an additional transition between regimes with large bias and minimal bias. Using random matrix theory, we show how each transition arises due to small nonzero eigenvalues in the Hessian matrix. Finally, we compare and contrast the phase diagram of the random linear features model to the random nonlinear features model and ordinary regression, highlighting the new phase transitions that result from the use of linear basis functions.

stat.ML↗

Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models

The bias-variance trade-off is a central concept in supervised learning. In classical statistics, increasing the complexity of a model (e.g., number of parameters) reduces bias but also increases variance. Until recently, it was commonly believed that optimal performance is achieved at intermediate model complexities which strike a balance between bias and variance. Modern Deep Learning methods flout this dogma, achieving state-of-the-art performance using "over-parameterized models" where the number of fit parameters is large enough to perfectly fit the training data. As a result, understanding bias and variance in over-parameterized models has emerged as a fundamental problem in machine learning. Here, we use methods from statistical physics to derive analytic expressions for bias and variance in two minimal models of over-parameterization (linear regression and two-layer neural networks with nonlinear data distributions), allowing us to disentangle properties stemming from the model architecture and random sampling of data. In both models, increasing the number of fit parameters leads to a phase transition where the training error goes to zero and the test error diverges as a result of the variance (while the bias remains finite). Beyond this threshold, the test error of the two-layer neural network decreases due to a monotonic decrease in \emph{both} the bias and variance in contrast with the classical bias-variance trade-off. We also show that in contrast with classical intuition, over-parameterized models can overfit even in the absence of noise and exhibit bias even if the student and teacher models match. We synthesize these results to construct a holistic understanding of generalization error and the bias-variance trade-off in over-parameterized models and relate our results to random matrix theory.

stat.ML↗

Cross-feeding shapes both competition and cooperation in microbial ecosystems

Recent work suggests that cross-feeding -- the secretion and consumption of metabolic biproducts by microbes -- is essential for understanding microbial ecology. Yet how cross-feeding and competition combine to give rise to ecosystem-level properties remains poorly understood. To address this question, we analytically analyze the Microbial Consumer Resource Model (MiCRM), a prominent ecological model commonly used to study microbial communities. Our mean-field solution exploits the fact that unlike replicas, the cavity method does not require the existence of a Lyapunov function. We use our solution to derive new species-packing bounds for diverse ecosystems in the presence of cross-feeding, as well as simple expressions for species richness and the abundance of secreted resources as a function of cross-feeding (metabolic leakage) and competition. Our results show how a complex interplay between competition for resources and cooperation resulting from metabolic exchange combine to shape the properties of microbial ecosystems.

q-bio.PE↗

Diverse communities behave like typical random ecosystems

In 1972, Robert May triggered a worldwide research program studying ecological communities using random matrix theory. Yet, it remains unclear if and when we can treat real communities as random ecosystems. Here, we draw on recent progress in random matrix theory and statistical physics to extend May's approach to generalized consumer-resource models. We show that in diverse ecosystems adding even modest amounts of noise to consumer preferences results in a transition to "typicality" where macroscopic ecological properties of communities are indistinguishable from those of random ecosystems, even when resource preferences have prominent designed structures. We test these ideas using numerical simulations on a wide variety of ecological models. Our work offers an explanation for the success of random consumer-resource models in reproducing experimentally observed ecological patterns in microbial communities and highlights the difficulty of scaling up bottom-up approaches in synthetic ecology to diverse communities.

q-bio.PE↗

The Geometry of Over-parameterized Regression and Adversarial Perturbations

Classical regression has a simple geometric description in terms of a projection of the training labels onto the column space of the design matrix. However, for over-parameterized models -- where the number of fit parameters is large enough to perfectly fit the training data -- this picture becomes uninformative. Here, we present an alternative geometric interpretation of regression that applies to both under- and over-parameterized models. Unlike the classical picture which takes place in the space of training labels, our new picture resides in the space of input features. This new feature-based perspective provides a natural geometric interpretation of the double-descent phenomenon in the context of bias and variance, explaining why it can occur even in the absence of label noise. Furthermore, we show that adversarial perturbations -- small perturbations to the input features that result in large changes in label values -- are a generic feature of biased models, arising from the underlying geometry. We demonstrate these ideas by analyzing three minimal models for over-parameterized linear least squares regression: without basis functions (input features equal model features) and with linear or nonlinear basis functions (two-layer neural networks with linear or nonlinear activation functions, respectively).

stat.ML↗

Understanding Species Abundance Distributions in Complex Ecosystems of Interacting Species

Niche and neutral theory are two prevailing, yet much debated, ideas in ecology proposed to explain the patterns of biodiversity. Whereas niche theory emphasizes selective differences between species and interspecific interactions in shaping the community, neutral theory supposes functional equivalence between species and points to stochasticity as the primary driver of ecological dynamics. In this work, we draw a bridge between these two opposing theories. Starting from a Lotka-Volterra (LV) model with demographic noise and random symmetric interactions, we analytically derive the stationary population statistics and species abundance distribution (SAD). Using these results, we demonstrate that the model can exhibit three classes of SADs commonly found in niche and neutral theories and found conditions that allow an ecosystem to transition between these various regimes. Thus, we reconcile how neutral-like statistics may arise from a diverse community with niche differentiation.

q-bio.PE↗

Arnol'd Tongues in Oscillator Systems with Nonuniform Spatial Driving

Nonlinear oscillator systems are ubiquitous in biology and physics, and their control is a practical problem in many experimental systems. Here we study this problem in the context of the two models of spatially-coupled oscillators: the complex Ginzburg-Landau equation (CGLE) and a generalization of the CGLE in which oscillators are coupled through an external medium (emCGLE). We focus on external control drives that vary in both space and time. We find that the spatial distribution of the drive signal controls the frequency ranges over which oscillators synchronize to the drive and that boundary conditions strongly influence synchronization to external drives for the CGLE. Our calculations also show that the emCGLE has a low density regime in which a broad range of frequencies can be synchronized for low drive amplitudes. We study the bifurcation structure of these models and find that they are very similar to results for the driven Kuramoto model, a system with no spatial structure. We conclude by discussing the implications of our results for controlling coupled oscillator systems such as the social amoebae \emph{Dictyostelium} and populations of BZ catalytic particles using spatially structured external drives.

nlin.PS↗

Tregs self-organize into a "computing ecosystem" and implement a sophisticated optimization algorithm for mediating immune response

Regulatory T cells (Tregs) play a crucial role in mediating immune response. Yet an algorithmic understanding of the role of Tregs in adaptive immunity remains lacking. Here, we present a biophysically realistic model of Treg mediated self-tolerance in which Tregs bind to self-antigens and locally inhibit the proliferation of nearby activated T cells. By exploiting a duality between ecological dynamics and constrained optimization, we show that Tregs tile the potential antigen space while simultaneously minimizing the overlap between Treg activation profiles. We find that for sufficiently high Treg diversity, Treg mediated self-tolerance is robust to fluctuations in self-antigen concentrations but lowering the Treg diversity results in a sharp transition -- related to the Gardner transition in perceptrons -- to a regime where changes in self-antigen concentrations can result in an auto-immune response. We propose a novel experimental test of this transition in immune-deficient mice and discuss potential implications for autoimmune diseases.

q-bio.PE↗

The Perturbative Resolvent Method: spectral densities of random matrix ensembles via perturbation theory

We present a simple, perturbative approach for calculating spectral densities for random matrix ensembles in the thermodynamic limit we call the Perturbative Resolvent Method (PRM). The PRM is based on constructing a linear system of equations and calculating how the solutions to these equation change in response to a small perturbation using the zero-temperature cavity method. We illustrate the power of the method by providing simple analytic derivations of the Wigner Semi-circle Law for symmetric matrices, the Marchenko-Pastur Law for Wishart matrices, the spectral density for a product Wishart matrix composed of two square matrices, and the Circle and elliptic laws for real random matrices.

cond-mat.dis-nn↗

Effect of resource dynamics on species packing in diverse ecosystems

The competitive exclusion principle asserts that coexisting species must occupy distinct ecological niches (i.e. the number of surviving species can not exceed the number of resources). An open question is to understand if and how different resource dynamics affect this bound. Here, we analyze a generalized consumer resource model with externally supplied resources and show that -- in contrast to self-renewing resources -- species can occupy only half of all available environmental niches. This motivates us to construct a new schema for classifying ecosystems based on species packing properties.

q-bio.PE↗

Data-driven modeling reveals a universal dynamic underlying the COVID-19 pandemic under social distancing

We show that the COVID-19 pandemic under social distancing exhibits universal dynamics. The cumulative numbers of both infections and deaths quickly cross over from exponential growth at early times to a longer period of power law growth, before eventually slowing. In agreement with a recent statistical forecasting model by the IHME, we show that this dynamics is well described by the erf function. Using this functional form, we perform a data collapse across countries and US states with very different population characteristics and social distancing policies, confirming the universal behavior of the COVID-19 outbreak. We show that the predictive power of statistical models is limited until a few days before curves flatten, forecast deaths and infections assuming current policies continue and compare our predictions to the IHME models. We present simulations showing this universal dynamics is consistent with disease transmission on scale-free networks and random networks with non-Markovian transmission dynamics.

q-bio.PE↗

The Community Simulator: A Python package for microbial ecology

Natural microbial communities contain hundreds to thousands of interacting species. For this reason, computational simulations are playing an increasingly important role in microbial ecology. In this manuscript, we present a new open-source, freely available Python package called Community Simulator for simulating microbial population dynamics in a reproducible, transparent and scalable way. The Community Simulator includes five major elements: tools for preparing the initial states and environmental conditions for a set of samples, automatic generation of dynamical equations based on a dictionary of modeling assumptions, random parameter sampling with tunable levels of metabolic and taxonomic structure, parallel integration of the dynamical equations, and support for metacommunity dynamics with migration between samples. To significantly speed up simulations using Community Simulator, our Python package implements a new Expectation-Maximization (EM) algorithm for finding equilibrium states of community dynamics that exploits a recently discovered duality between ecological dynamics and convex optimization. We present data showing that this EM algorithm improves performance by between one and two orders compared to direct numerical integration of the corresponding ordinary differential equations. We conclude by listing several recent applications of the Community Simulator to problems in microbial ecology, and discussing possible extensions of the package for directly analyzing microbiome compositional data.

q-bio.PE↗

A minimal model for microbial biodiversity can reproduce experimentally observed ecological patterns

Surveys of microbial biodiversity such as the Earth Microbiome Project (EMP) and the Human Microbiome Project (HMP) have revealed robust ecological patterns across different environments. A major goal in ecology is to leverage these patterns to identify the ecological processes shaping microbial ecosystems. One promising approach is to use minimal models that can relate mechanistic assumptions at the microbe scale to community-level patterns. Here, we demonstrate the utility of this approach by showing that the Microbial Consumer Resource Model (MiCRM) -- a minimal model for microbial communities with resource competition, metabolic crossfeeding and stochastic colonization -- can qualitatively reproduce patterns found in survey data including compositional gradients, dissimilarity/overlap correlations, richness/harshness correlations, and nestedness of community composition. By using the MiCRM to generate synthetic data with different environmental and taxonomical structure, we show that large scale patterns in the EMP can be reproduced by considering the energetic cost of surviving in harsh environments and HMP patterns may reflect the importance of environmental filtering in shaping competition. We also show that recently discovered dissimilarity-overlap correlations in the HMP likely arise from communities that share similar environments rather than reflecting universal dynamics. We identify ecologically meaningful changes in parameters that alter or destroy each one of these patterns, suggesting new mechanistic hypotheses for further investigation. These findings highlight the promise of minimal models for microbial ecology.

q-bio.PE↗

The Minimum Environmental Perturbation Principle: A New Perspective on Niche Theory

Fifty years ago, Robert MacArthur showed that stable equilibria optimize quadratic functions of the population sizes in several important ecological models. Here, we generalize this finding to a broader class of systems within the framework of contemporary niche theory, and precisely state the conditions under which an optimization principle (not necessarily quadratic) can be obtained. We show that conducting the optimization in the space of environmental states instead of population sizes leads to a universal and transparent physical interpretation of the objective function. Specifically, the equilibrium state minimizes the perturbation of the environment induced by the presence of the competing species, subject to the constraint that no species has a positive net growth rate. We use this "minimum environmental perturbation principle" to make new predictions for eco-evolution and community assembly, and describe a simple experimental setting where its conditions of validity have been empirically tested.

q-bio.PE↗

Machine Learning as Ecology

Machine learning methods have had spectacular success on numerous problems. Here we show that a prominent class of learning algorithms - including Support Vector Machines (SVMs) -- have a natural interpretation in terms of ecological dynamics. We use these ideas to design new online SVM algorithms that exploit ecological invasions, and benchmark performance using the MNIST dataset. Our work provides a new ecological lens through which we can view statistical learning and opens the possibility of designing ecosystems for machine learning. Supplemental code is found at https://github.com/owenhowell20/EcoSVM.

cs.LG↗

A high-bias, low-variance introduction to Machine Learning for physicists

Machine Learning (ML) is one of the most exciting and dynamic areas of modern research and application. The purpose of this review is to provide an introduction to the core concepts and tools of machine learning in a manner easily understood and intuitive to physicists. The review begins by covering fundamental concepts in ML and modern statistics such as the bias-variance tradeoff, overfitting, regularization, generalization, and gradient descent before moving on to more advanced topics in both supervised and unsupervised learning. Topics covered in the review include ensemble models, deep learning and neural networks, clustering and data visualization, energy-based models (including MaxEnt models and Restricted Boltzmann Machines), and variational methods. Throughout, we emphasize the many natural connections between ML and statistical physics. A notable aspect of the review is the use of Python Jupyter notebooks to introduce modern ML/statistical packages to readers using physics-inspired datasets (the Ising Model and Monte-Carlo simulations of supersymmetric decays of proton-proton collisions). We conclude with an extended outlook discussing possible uses of machine learning for furthering our understanding of the physical world as well as open problems in ML where physicists may be able to contribute. (Notebooks are available at https://physics.bu.edu/~pankajm/MLnotebooks.html )

physics.comp-ph↗

Available energy fluxes drive a transition in the diversity, stability, and functional structure of microbial communities

A fundamental goal of microbial ecology is to understand what determines the diversity, stability, and structure of microbial ecosystems. The microbial context poses special conceptual challenges because of the strong mutual influences between the microbes and their chemical environment through the consumption and production of metabolites. By analyzing a generalized consumer resource model that explicitly includes cross-feeding, stochastic colonization, and thermodynamics, we show that complex microbial communities generically exhibit a transition as a function of available energy fluxes from a "resource-limited" regime where community structure and stability is shaped by energetic and metabolic considerations to a diverse regime where the dominant force shaping microbial communities is the overlap between species' consumption preferences. These two regimes have distinct species abundance patterns, different functional profiles, and respond differently to environmental perturbations. Our model reproduces large-scale ecological patterns observed across multiple experimental settings such as nestedness and differential beta diversity patterns along energy gradients. We discuss the experimental implications of our results and possible connections with disorder-induced phase transitions in statistical physics.

physics.bio-ph↗

Glassy Phase of Optimal Quantum Control

We study the problem of preparing a quantum many-body system from an initial to a target state by optimizing the fidelity over the family of bang-bang protocols. We present compelling numerical evidence for a universal spin-glass-like transition controlled by the protocol time duration. The glassy critical point is marked by a proliferation of protocols with close-to-optimal fidelity and with a true optimum that appears exponentially difficult to locate. Using a machine learning (ML) inspired framework based on the manifold learning algorithm t-SNE, we are able to visualize the geometry of the high-dimensional control landscape in an effective low-dimensional representation. Across the transition, the control landscape features an exponential number of clusters separated by extensive barriers, which bears a strong resemblance with replica symmetry breaking in spin glasses and random satisfiability problems. We further show that the quantum control landscape maps onto a disorder-free classical Ising model with frustrated nonlocal, multibody interactions. Our work highlights an intricate but unexpected connection between optimal quantum control and spin glass physics, and shows how tools from ML can be used to visualize and understand glassy optimization landscapes.

quant-ph↗