Searcharxiv⌕ Search

arXiv subjects

Matt Thomson

Publications and source records attributed to Matt Thomson.

At least 37 records · Page 2Linked to original sources

Engineering flexible machine learning systems by traversing functionally-invariant paths

Transformers have emerged as the state of the art neural network architecture for natural language processing and computer vision. In the foundation model paradigm, large transformer models (BERT, GPT3/4, Bloom, ViT) are pre-trained on self-supervised tasks such as word or image masking, and then, adapted through fine-tuning for downstream user applications including instruction following and Question Answering. While many approaches have been developed for model fine-tuning including low-rank weight update strategies (eg. LoRA), underlying mathematical principles that enable network adaptation without knowledge loss remain poorly understood. Here, we introduce a differential geometry framework, functionally invariant paths (FIP), that provides flexible and continuous adaptation of neural networks for a range of machine learning goals and network sparsification objectives. We conceptualize the weight space of a neural network as a curved Riemannian manifold equipped with a metric tensor whose spectrum defines low rank subspaces in weight space that accommodate network adaptation without loss of prior knowledge. We formalize adaptation as movement along a geodesic path in weight space while searching for networks that accommodate secondary objectives. With modest computational resources, the FIP algorithm achieves comparable to state of the art performance on continual learning and sparsification tasks for language models (BERT), vision transformers (ViT, DeIT), and the CNNs. Broadly, we conceptualize a neural network as a mathematical object that can be iteratively transformed into distinct configurations by the path-sampling algorithm to define a sub-manifold of weight space that can be harnessed to achieve user goals.

cs.LG↗

Tryage: Real-time, intelligent Routing of User Prompts to Large Language Models

The introduction of the transformer architecture and the self-attention mechanism has led to an explosive production of language models trained on specific downstream tasks and data domains. With over 200, 000 models in the Hugging Face ecosystem, users grapple with selecting and optimizing models to suit multifaceted workflows and data domains while addressing computational, security, and recency concerns. There is an urgent need for machine learning frameworks that can eliminate the burden of model selection and customization and unleash the incredible power of the vast emerging model library for end users. Here, we propose a context-aware routing system, Tryage, that leverages a language model router for optimal selection of expert models from a model library based on analysis of individual input prompts. Inspired by the thalamic router in the brain, Tryage employs a perceptive router to predict down-stream model performance on prompts and, then, makes a routing decision using an objective function that integrates performance predictions with user goals and constraints that are incorporated through flags (e.g., model size, model recency). Tryage allows users to explore a Pareto front and automatically trade-off between task accuracy and secondary goals including minimization of model size, recency, security, verbosity, and readability. Across heterogeneous data sets that include code, text, clinical data, and patents, the Tryage framework surpasses Gorilla and GPT3.5 turbo in dynamic model selection identifying the optimal model with an accuracy of 50.9% , compared to 23.6% by GPT 3.5 Turbo and 10.8% by Gorilla. Conceptually, Tryage demonstrates how routing models can be applied to program and control the behavior of multi-model LLM systems to maximize efficient use of the expanding and evolving language model ecosystem.

cs.LG↗

Spatiotemporal patterning of extensile active stresses in microtubule-based active fluids

Active stresses, which are collectively generated by the motion of energy-consuming rod-like constituents, generate chaotic autonomous flows. Controlling active stresses in space and time is an essential prerequisite for controlling the intrinsically chaotic dynamics of extensile active fluids. We design single-headed kinesin molecular motors that exhibit optically enhanced clustering, and thus enable precise and repeatable spatial and temporal control of extensile active stresses. Such motors enable rapid, reversible switching between flowing and quiescent states. In turn, spatio-temporal patterning of the active stress controls the evolution of the ubiquitous bend-instability of extensile active fluids and determines its critical length dependence. Combining optically controlled clusters with conventional kinesin motors enables one-time switching from contractile to extensile active stresses. These results open a path towards real-time control of the autonomous flows generated by active fluids.

cond-mat.soft↗

Therapeutic algebra of immunomodulatory drug responses at single-cell resolution

Therapeutic modulation of immune states is central to the treatment of human disease. However, how drugs and drug combinations impact the diverse cell types in the human immune system remains poorly understood at the transcriptome scale. Here, we apply single-cell mRNA-seq to profile the response of human immune cells to 502 immunomodulatory drugs alone and in combination. We develop a unified mathematical model that quantitatively describes the transcriptome scale response of myeloid and lymphoid cell types to individual drugs and drug combinations through a single inferred regulatory network. The mathematical model reveals how drug combinations generate novel, macrophage and T-cell states by recruiting combinations of gene expression programs through both additive and non-additive drug interactions. A simplified drug response algebra allows us to predict the continuous modulation of immune cell populations between activated, resting and hyper-inhibited states through combinatorial drug dose titrations. Our results suggest that transcriptome-scale mathematical models could enable the design of therapeutic strategies for programming the human immune system using combinations of therapeutics.

q-bio.GN↗

Active feature selection discovers minimal gene sets for classifying cell types and disease states with single-cell mRNA-seq data

Sequencing costs currently prohibit the application of single-cell mRNA-seq to many biological and clinical analyses. Targeted single-cell mRNA-sequencing reduces sequencing costs by profiling reduced gene sets that capture biological information with a minimal number of genes. Here, we introduce an active learning method (ActiveSVM) that identifies minimal but highly-informative gene sets that enable the identification of cell-types, physiological states, and genetic perturbations in single-cell data using a small number of genes. Our active feature selection procedure generates minimal gene sets from single-cell data through an iterative cell-type classification task where misclassified cells are examined at each round of analysis to identify maximally informative genes through an `active' support vector machine (ActiveSVM) classifier. By focusing computational resources on misclassified cells, ActiveSVM scales to analyze data sets with over a million single cells. We demonstrate that ActiveSVM feature selection identifies gene sets that enable ~90% cell-type classification accuracy across a variety of data sets including cell atlas and disease characterization data sets. The method generalizes to reveal genes that respond to genetic perturbations and to identify region specific gene expression patterns in spatial transcriptomics data. The discovery of small but highly informative gene sets should enable substantial reductions in the number of measurements necessary for application of single-cell mRNA-seq to clinical tests, therapeutic discovery, and genetic screens.

q-bio.GN↗

Cell density controls signal propagation waves in a multicellular synthetic gene circuit

During organismal development, biochemical reaction networks sense and respond to mechanical forces to coordinate embryonic patterning with embryo morphogenesis. Factors such as cortical tension, cell density, and matrix mechanical properties influence differentiation and cell fate decisions by modulating gene regulatory signaling networks. A major goal in synthetic development is to construct gene regulatory circuits that program the patterning and morphogenesis of synthetic multicellular structures. However, in the synthetic context, little is known regarding how the physical properties of the growth environment impact the behavior of synthetic gene circuits. Here, we exploit physical-chemical coupling observed in a synthetic patterning circuit in order to control the size and spatial distribution of patterned synthetic cell sheets. We show that cell density attenuates the propagation of signal between neighboring cells in a multicellular sheet containing a contact-dependent patterning circuit based on the synNotch signaling system. Density-dependent attenuation leads to a signal propagation wave that exhibits distinct qualitative phases of persistent propagation, transient propagation, and no propagation. Through computational modeling, we demonstrate that cell growth parameters determine the phase of propagation observed within a growing cell sheet. Using growth-modulating drugs and spatial density gradients, we control the size of synNotch-activated cell populations and generate tissue-scale activation gradients and kinematic waves. Our study reveals that density-dependent synNotch activity can be exploited to control a synthetic multicellular patterning circuit. More broadly, we show that synthetic gene circuits can be critically impacted by their physical context, providing an alternate means for programming circuit behavior.

q-bio.CB↗

Signaling receptor localization maximizes cellular information acquisition in spatially-structured, natural environments

Cells in natural environments like tissue or soil sense and respond to extracellular ligands with intricately structured and non-monotonic spatial distributions that are sculpted by processes such as fluid flow and substrate adhesion. Nevertheless, traditional approaches to studying cell sensing assume signals are either uniform or monotonic, neglecting spatial structures of natural environments. In this work, we show that spatial sensing and navigation can be optimized by adapting the spatial organization of signaling pathways to the spatial structure of the environment. By viewing cell surface receptors as a sensor network, we develop an information theoretic framework for computing the optimal spatial organization of a sensing system for a given spatial signaling environment. Applying the framework to simulated environments, we find that spatial receptor localization maximizes information acquisition in many natural contexts, including tissue and soil. Receptor localization extends naturally to produce a dynamic protocol for redistributing signaling receptors during cell navigation and can be implemented in a cell using a feedback scheme. In a simulated tissue environment, dynamic receptor localization boosts navigation efficiency by 30-fold. Broadly, our framework readily adapts to studying how the spatial organization of signaling components other than receptors can be modulated to improve cellular information processing.

q-bio.CB↗

Phenomenological model of motility by spatiotemporal modulation of active interactions

Transport at microscopic length scales is essential in biological systems and various technologies, including microfluidics. Recent experiments achieved self-organized transport phenomena in microtubule active matter using light to modulate motor-protein activity in time and space. Here, we introduce a novel phenomenological model to explain such experiments. Our model, based on spatially modulated particle interactions, reveals a possible mechanism for emergent transport phenomena in light-controlled active matter, including motility and contraction. In particular, the model's analytic treatment elucidates the conservation of the center of mass of activated particles as a fundamental mechanism of material transport and demonstrates the necessity of memory for sustained motility. Furthermore, we generalize the model to explain other phenomena, like microtubule aster-aster interactions induced by more complicated activation geometries. Our results demonstrate that the model provides a possible foundation for the phenomenological understanding of light-controlled active matter, and it will enable the design and optimization of transport protocols for active matter devices.

cond-mat.soft↗

Solving hybrid machine learning tasks by traversing weight space geodesics

Machine learning problems have an intrinsic geometric structure as central objects including a neural network's weight space and the loss function associated with a particular task can be viewed as encoding the intrinsic geometry of a given machine learning problem. Therefore, geometric concepts can be applied to analyze and understand theoretical properties of machine learning strategies as well as to develop new algorithms. In this paper, we address three seemingly unrelated open questions in machine learning by viewing them through a unified framework grounded in differential geometry. Specifically, we view the weight space of a neural network as a manifold endowed with a Riemannian metric that encodes performance on specific tasks. By defining a metric, we can construct geodesic, minimum length, paths in weight space that represent sets of networks of equivalent or near equivalent functional performance on a specific task. We, then, traverse geodesic paths while identifying networks that satisfy a second objective. Inspired by the geometric insight, we apply our geodesic framework to 3 major applications: (i) Network sparsification (ii) Mitigating catastrophic forgetting by constructing networks with high performance on a series of objectives and (iii) Finding high-accuracy paths connecting distinct local optima of deep networks in the non-convex loss landscape. Our results are obtained on a wide range of network architectures (MLP, VGG11/16) trained on MNIST, CIFAR-10/100. Broadly, we introduce a geometric framework that unifies a range of machine learning objectives and that can be applied to multiple classes of neural network architectures.

cs.LG↗

Reinforcement Learning reveals fundamental limits on the mixing of active particles

The control of far-from-equilibrium physical systems, including active materials, has emerged as an important area for the application of reinforcement learning (RL) strategies to derive control policies for physical systems. In active materials, non-linear dynamics and long-range interactions between particles prohibit closed-form descriptions of the system's dynamics and prevent explicit solutions to optimal control problems. Due to fundamental challenges in solving for explicit control strategies, RL has emerged as an approach to derive control strategies for far-from-equilibrium active matter systems. However, an important open question is how the mathematical structure and the physical properties of the active matter systems determine the tractability of RL for learning control policies. In this work, we show that RL can only find good strategies to the canonical active matter task of mixing for systems that combine attractive and repulsive particle interactions. Using mathematical results from dynamical systems theory, we relate the availability of both interaction types with the existence of hyperbolic dynamics and the ability of RL to find homogeneous mixing strategies. In particular, we show that for drag-dominated translational-invariant particle systems, hyperbolic dynamics and, therefore, mixing requires combining attractive and repulsive interactions. Broadly, our work demonstrates how fundamental physical and mathematical properties of dynamical systems can enable or constrain reinforcement learning-based control.

cs.LG↗

Programming Boundary Deformation Patterns in Active Networks

Active materials take advantage of their internal sources of energy to self-organize in an automated manner. This feature provides a novel opportunity to design micron-scale machines with minimal required control. However, self-organization goes hand in hand with predetermined dynamics that are hardly susceptible to environmental perturbations. Therefore utilizing this feature of active systems requires harnessing and directing the macroscopic dynamics to achieve specific functions; which in turn necessitates understanding the underlying mechanisms of active forces. Here we devise an optical control protocol to engineer the dynamics of active networks composed of microtubules and light-activatable motor proteins. The protocol enables carving activated networks of different shapes, and isolating them from the embedding solution. Studying a large set of shapes, we observe that the active networks contract in a shape-preserving manner that persists over the course of contraction. We formulate a coarse-grained theory and demonstrate that self-similarity of contraction is associated with viscous-like active stresses. These findings help us program the dynamics of the network through manipulating the light intensity in space and time, and maneuver the network into bending in specific directions, as well as temporally alternating directions. Our work improves understanding the active dynamics in contractile networks, and paves a new path towards engineering the dynamics of a large class of active materials.

cond-mat.soft↗

Sparsifying networks by traversing Geodesics

The geometry of weight spaces and functional manifolds of neural networks play an important role towards 'understanding' the intricacies of ML. In this paper, we attempt to solve certain open questions in ML, by viewing them through the lens of geometry, ultimately relating it to the discovery of points or paths of equivalent function in these spaces. We propose a mathematical framework to evaluate geodesics in the functional space, to find high-performance paths from a dense network to its sparser counterpart. Our results are obtained on VGG-11 trained on CIFAR-10 and MLP's trained on MNIST. Broadly, we demonstrate that the framework is general, and can be applied to a wide variety of problems, ranging from sparsification to alleviating catastrophic forgetting.

cs.LG↗

Persistent fluid flows defined by active matter boundaries

Biological systems achieve precise control over ambient fluids through the self-organization of active protein structures including flagella, cilia, and cytoskeletal networks. In active structures individual proteins consume chemical energy to generate force and motion at molecular length scales. Self-organization of protein components enables the control and modulation of fluid flow fields on micron scales. The physical principles underlying the organization and control of active-matter driven fluid flows are poorly understood. Here, we apply an optically-controlled active-matter system composed of microtubule filaments and light-switchable kinesin motor proteins to analyze the emergence of persistent flow fields in a model active matter system. Using light, we form contractile microtubule networks of varying shape. We analyze the fluid flow fields generated by a wide range of microtubule network geometries and explain the resulting flow fields within a unified theoretical framework. We specifically demonstrate that the geometry of microtubule flux at the boundary of contracting microtubule networks predicts the steady-state fluid flow fields across polygonal network geometries through finite-element simulations. Our work provides a foundation for programming microscopic fluid-flows with controllable active matter and could enable the engineering of versatile and dynamic microfluidic devices.

cond-mat.soft↗

Self-organization of multi-layer spiking neural networks

Living neural networks in our brains autonomously self-organize into large, complex architectures during early development to result in an organized and functional organic computational device. A key mechanism that enables the formation of complex architecture in the developing brain is the emergence of traveling spatio-temporal waves of neuronal activity across the growing brain. Inspired by this strategy, we attempt to efficiently self-organize large neural networks with an arbitrary number of layers into a wide variety of architectures. To achieve this, we propose a modular tool-kit in the form of a dynamical system that can be seamlessly stacked to assemble multi-layer neural networks. The dynamical system encapsulates the dynamics of spiking units, their inter/intra layer interactions as well as the plasticity rules that control the flow of information between layers. The key features of our tool-kit are (1) autonomous spatio-temporal waves across multiple layers triggered by activity in the preceding layer and (2) Spike-timing dependent plasticity (STDP) learning rules that update the inter-layer connectivity based on wave activity in the connecting layers. Our framework leads to the self-organization of a wide variety of architectures, ranging from multi-layer perceptrons to autoencoders. We also demonstrate that emergent waves can self-organize spiking network architecture to perform unsupervised learning, and networks can be coupled with a linear classifier to perform classification on classic image datasets like MNIST. Broadly, our work shows that a dynamical systems framework for learning can be used to self-organize large computational devices.

cs.NE↗

Geometric algorithms for predicting resilience and recovering damage in neural networks

Biological neural networks have evolved to maintain performance despite significant circuit damage. To survive damage, biological network architectures have both intrinsic resilience to component loss and also activate recovery programs that adjust network weights through plasticity to stabilize performance. Despite the importance of resilience in technology applications, the resilience of artificial neural networks is poorly understood, and autonomous recovery algorithms have yet to be developed. In this paper, we establish a mathematical framework to analyze the resilience of artificial neural networks through the lens of differential geometry. Our geometric language provides natural algorithms that identify local vulnerabilities in trained networks as well as recovery algorithms that dynamically adjust networks to compensate for damage. We reveal striking vulnerabilities in commonly used image analysis networks, like MLP's and CNN's trained on MNIST and CIFAR10 respectively. We also uncover high-performance recovery paths that enable the same networks to dynamically re-adjust their parameters to compensate for damage. Broadly, our work provides procedures that endow artificial systems with resilience and rapid-recovery routines to enhance their integration with IoT devices as well as enable their deployment for critical applications.

cs.NE↗

Active Learning of Spin Network Models

The inverse statistical problem of finding direct interactions in complex networks is difficult. In the natural sciences, well-controlled perturbation experiments are widely used to probe the structure of complex networks. However, our understanding of how and why perturbations aid inference remains heuristic, and we lack automated procedures that determine network structure by combining inference and perturbation. Therefore, we propose a general mathematical framework to study inference with iteratively applied perturbations. Using the formulation of information geometry, our framework quantifies the difficulty of inference and the information gain from perturbations through the curvature of the underlying parameter manifold, measured by Fisher information. We apply the framework to the inference of spin network models and find that designed perturbations can reduce the sampling complexity by $10^6$-fold across a variety of network architectures. Physically, our framework reveals that perturbations boost inference by causing a network to explore previously inaccessible states. Optimal perturbations break spin-spin correlations within a network, increasing the information available for inference and thus reducing sampling complexity by orders of magnitude. Our active learning framework could be powerful in the analysis of complex networks as well as in the rational design of experiments.

cond-mat.dis-nn↗

Neural networks grown and self-organized by noise

Living neural networks emerge through a process of growth and self-organization that begins with a single cell and results in a brain, an organized and functional computational device. Artificial neural networks, however, rely on human-designed, hand-programmed architectures for their remarkable performance. Can we develop artificial computational devices that can grow and self-organize without human intervention? In this paper, we propose a biologically inspired developmental algorithm that can 'grow' a functional, layered neural network from a single initial cell. The algorithm organizes inter-layer connections to construct a convolutional pooling layer, a key constituent of convolutional neural networks (CNN's). Our approach is inspired by the mechanisms employed by the early visual system to wire the retina to the lateral geniculate nucleus (LGN), days before animals open their eyes. The key ingredients for robust self-organization are an emergent spontaneous spatiotemporal activity wave in the first layer and a local learning rule in the second layer that 'learns' the underlying activity pattern in the first layer. The algorithm is adaptable to a wide-range of input-layer geometries, robust to malfunctioning units in the first layer, and so can be used to successfully grow and self-organize pooling architectures of different pool-sizes and shapes. The algorithm provides a primitive procedure for constructing layered neural networks through growth and self-organization. Broadly, our work shows that biologically inspired developmental algorithms can be applied to autonomously grow functional 'brains' in-silico.

cs.NE↗

Controlling Organization and Forces in Active Matter Through Optically-Defined Boundaries

Living systems are capable of locomotion, reconfiguration, and replication. To perform these tasks, cells spatiotemporally coordinate the interactions of force-generating, "active" molecules that create and manipulate non-equilibrium structures and force fields that span up to millimeter length scales [1-3]. Experimental active matter systems of biological or synthetic molecules are capable of spontaneously organizing into structures [4,5] and generating global flows [6-9]. However, these experimental systems lack the spatiotemporal control found in cells, limiting their utility for studying non-equilibrium phenomena and bioinspired engineering. Here, we uncover non-equilibrium phenomena and principles by optically controlling structures and fluid flow in an engineered system of active biomolecules. Our engineered system consists of purified microtubules and light-activatable motor proteins that crosslink and organize microtubules into distinct structures upon illumination. We develop basic operations, defined as sets of light patterns, to create, move, and merge microtubule structures. By composing these basic operations, we are able to create microtubule networks that span several hundred microns in length and contract at speeds up to an order of magnitude faster than the speed of an individual motor. We manipulate these contractile networks to generate and sculpt persistent fluid flows. The principles of boundary-mediated control we uncover may be used to study emergent cellular structures and forces and to develop programmable active matter devices.

cond-mat.soft↗