SearcharxivSearch

arXiv subjects

Michael R. DeWeese

Publications and source records attributed to Michael R. DeWeese.

At least 19 recordsLinked to original sources

The Entropy of Floating-Point Numbers

Here we present an analytic approximation for the entropy of floating-point numbers, along with bounds on the error of this approximation. It is well-known that the differential entropy is tightly linked to the discrete entropy of a uniformly quantized random variable. Our approximation uncovers a different quantity that provides this link for floating-point quantization. Additionally, we prove that the entropy of a floating-point quantized random variable is approximately unchanged under scaling. Closed-form expressions for the floating-point entropy of common distributions are provided and compared to exact results.

cs.IT

A Theory of Saddle Escape in Deep Nonlinear Networks

In deep networks with small initialization, training exhibits long plateaus separated by sharp feature-acquisition transitions. Whereas shallow nonlinear networks and deep linear networks are well studied, extending these analyses to deep nonlinear networks remains challenging. We derive an exact identity for the imbalance of Frobenius norms of layer weight matrices that holds for any smooth activation and any differentiable loss and use this to classify activation functions into four universality classes. On the permutation-symmetric submanifold, the identity combines with an approximate balance law to reduce the full matrix flow to a scalar ODE, giving a critical-depth escape time law $τ_\star = Θ(\varepsilon^{-(r-2)})$ governed by the number $r$ of layers at the bottleneck scale rather than the total depth $L$. We find that this same $r-2$ exponent is recovered under He-normal initialization with $r$ bottleneck layers rescaled by $\varepsilon$, where the symmetry manifold is preserved by the flow but not attracting. We find close agreement between our theory and numerical simulations.

cs.LG

Optimal active engines obey the thermodynamic Lorentz force law

What are the fundamental limitations for finite-time engines that extract work from active nonequilibrium systems, and what are the optimal protocols that approach them? We show that the finite-time work extraction for nonconservative overdamped Langevin systems may be rewritten as a Lorentz force Lagrangian action, with the kinetic term corresponding to a thermodynamic metric term that is an $L_2$-optimal transport cost for the time-dependent probability density, and the magnetic field coupling term corresponding to an effective quasistatic work extraction, proving that optimal protocols counterdiabatically steer the thermodynamic state trajectory to satisfy a Lorentz force law defined on thermodynamic state space. We utilize and reinterpret classic concepts from electromagnetism in the setting of cyclical nonequilibrium processes. We show that the housekeeping heat can be controlled to be arbitrarily close to zero by minimizing nonequilibrium fluctuations. It immediately follows from our results that the constant-velocity angle clamp protocol applied to the $F_1$ molecular motor in a recent experiment [Mishima et at, 2025] is in fact the globally optimal protocol: it produces zero housekeeping heat while simultaneously minimizing dissipation and maximizing work transduction.

cond-mat.stat-mech

The Thermodynamic Costs of Simple Linear Regression

The construction of models from data is a significant contributor to the energetic costs of computation. Because of this, understanding how foundational thermodynamic bounds apply to modeling algorithms will be increasingly important. Here, we study the thermodynamic costs of a basic and fundamental modeling algorithm: simple linear regression. Following Landauer, we approximate the thermodynamic lower bound on irreversibly performing both exact linear regression and linear regression via stochastic gradient descent as implemented on floating-point numbers. From this, we derive energycost aware scaling laws for the optimal dataset size for training a linear regression model given a generalization error dependent demand for inference. Additionally, we discuss a method to lower bound the entropy production from the mismatch cost for algorithms with continuous input variables.

cond-mat.stat-mech

Higher-order response theory in optimal stochastic thermodynamics

Linear response theory has found many applications in statistical physics. One of these is to compute minimal-work protocols that drive nonequilibrium systems between different thermodynamic states, which are useful for designing engineered nanoscale systems and understanding biomolecular machines. We compare and explore the relationships between linear-response-based approximations used to study optimal protocols in different driving regimes by showing that they arise as controlled truncations of a general causal response (Volterra) expansion. We then construct higher-order response terms and discuss the drawbacks and utility of their inclusion. We illustrate our results for an overdamped particle in a harmonic trap, ultimately showing that the inclusion of higher-order response in calculating optimal protocols provides marginal improvement in effectiveness despite incurring a significant computational expense, while introducing the possibility of predicting arbitrarily low and unphysical negative excess work.

cond-mat.stat-mech

Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks

What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dynamics of feature learning in two-layer networks trained from small initialization. Prior works have shown that gradient flow in this regime exhibits a staircase-like loss curve, alternating between plateaus where neurons slowly align to useful directions and sharp drops where neurons rapidly grow in norm. AGF approximates this behavior as an alternating two-step process: maximizing a utility function over dormant neurons and minimizing a cost function over active ones. AGF begins with all neurons dormant. At each iteration, a dormant neuron activates, triggering the acquisition of a feature and a drop in the loss. AGF quantifies the order, timing, and magnitude of these drops, matching experiments across several commonly studied architectures. We show that AGF unifies and extends existing saddle-to-saddle analyses in fully connected linear networks and attention-only linear transformers, where the learned features are singular modes and principal components, respectively. In diagonal linear networks, we prove AGF converges to gradient flow in the limit of vanishing initialization. Applying AGF to quadratic networks trained to perform modular addition, we give the first complete characterization of the training dynamics, revealing that networks learn Fourier features in decreasing order of coefficient magnitude. Altogether, AGF offers a promising step towards understanding feature learning in neural networks.

cs.LG

Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models

Self-supervised word embedding algorithms such as word2vec provide a minimal setting for studying representation learning in language modeling. We examine the quartic Taylor approximation of the word2vec loss around the origin, and we show that both the resulting training dynamics and the final performance on downstream tasks are empirically very similar to those of word2vec. Our main contribution is to analytically solve for both the gradient flow training dynamics and the final word embeddings in terms of only the corpus statistics and training hyperparameters. The solutions reveal that these models learn orthogonal linear subspaces one at a time, each one incrementing the effective rank of the embeddings until model capacity is saturated. Training on Wikipedia, we find that each of the top linear subspaces represents an interpretable topic-level concept. Finally, we apply our theory to describe how linear representations of more abstract semantic concepts emerge during training; these can be used to complete analogies via vector addition.

cs.LG

Beyond Linear Response: Equivalence between Thermodynamic Geometry and Optimal Transport

A fundamental result of thermodynamic geometry is that the optimal, minimal-work protocol that drives a nonequilibrium system between two thermodynamic states in the slow-driving limit is given by a geodesic of the friction tensor, a Riemannian metric defined on control space. For overdamped dynamics in arbitrary dimensions, we demonstrate that thermodynamic geometry is equivalent to $L^2$ optimal transport geometry defined on the space of equilibrium distributions corresponding to the control parameters. We show that obtaining optimal protocols past the slow-driving or linear response regime is computationally tractable as the sum of a friction tensor geodesic and a counterdiabatic term related to the Fisher information metric. These geodesic-counterdiabatic optimal protocols are exact for parameteric harmonic potentials, reproduce the surprising non-monotonic behavior recently discovered in linearly-biased double well optimal protocols, and explain the ubiquitous discontinuous jumps observed at the beginning and end times.

cond-mat.stat-mech

Time-Asymmetric Fluctuation Theorem and Efficient Free Energy Estimation

The free-energy difference $ΔF$ between two high-dimensional systems is notoriously difficult to compute, but very important for many applications, such as drug discovery. We demonstrate that an unconventional definition of work introduced by Vaikuntanathan and Jarzynski (2008) satisfies a microscopic fluctuation theorem that relates path ensembles that are driven by protocols unequal under time-reversal. It has been shown before that counterdiabatic protocols -- those having additional forcing that enforces the system to remain in instantaneous equilibrium, also known as escorted dynamics or engineered swift equilibration -- yield zero-variance work measurements for this definition. We show that this time-asymmetric microscopic fluctuation theorem can be exploited for efficient free energy estimation by developing a simple (i.e., neural-network free) and efficient adaptive time-asymmetric protocol optimization algorithm that yields $ΔF$ estimates that are orders of magnitude lower in mean squared error than the generic linear interpolation protocol with which it is initialized.

cond-mat.soft

The Eigenlearning Framework: A Conservation Law Perspective on Kernel Regression and Wide Neural Networks

We derive simple closed-form estimates for the test risk and other generalization metrics of kernel ridge regression (KRR). Relative to prior work, our derivations are greatly simplified and our final expressions are more readily interpreted. These improvements are enabled by our identification of a sharp conservation law which limits the ability of KRR to learn any orthonormal basis of functions. Test risk and other objects of interest are expressed transparently in terms of our conserved quantity evaluated in the kernel eigenbasis. We use our improved framework to: i) provide a theoretical explanation for the "deep bootstrap" of Nakkiran et al (2020), ii) generalize a previous result regarding the hardness of the classic parity problem, iii) fashion a theoretical tool for the study of adversarial robustness, and iv) draw a tight analogy between KRR and a well-studied system in statistical physics.

cs.LG

Shortcut engineering of active matter: run-and-tumble particles

Shortcut engineering consists of a class of approaches to rapidly manipulate physical systems by means of specially designed external controls. In this Letter, we apply these approaches to run-and-tumble particles, which are designed to mimic the chemotactic behavior of bacteria and therefore exhibit complex dynamics due to their self-propulsion and random reorientation, making them difficult to control. Following a recent successful application to active Brownian particles, we find a general solution for the rapid control of 1D run-and-tumble particles in a harmonic potential. We demonstrate the effectiveness of our approach using numerical simulations and show that it can lead to a significant speedup compared to simple quenched protocols. Our results extend shortcut engineering to a wider class of active systems and demonstrate that it is a promising tool for controlling the dynamics of active matter, which has implications for a wide range of applications in fields such as materials science and biophysics.

cond-mat.stat-mech

Reverse Engineering the Neural Tangent Kernel

The development of methods to guide the design of neural networks is an important open challenge for deep learning theory. As a paradigm for principled neural architecture design, we propose the translation of high-performing kernels, which are better-understood and amenable to first-principles design, into equivalent network architectures, which have superior efficiency, flexibility, and feature learning. To this end, we constructively prove that, with just an appropriate choice of activation function, any positive-semidefinite dot-product kernel can be realized as either the NNGP or neural tangent kernel of a fully-connected neural network with only one hidden layer. We verify our construction numerically and demonstrate its utility as a design tool for finite fully-connected networks in several experiments.

cs.LG

Limited-control optimal protocols arbitrarily far from equilibrium

Recent studies have explored finite-time dissipation-minimizing protocols for stochastic thermodynamic systems driven arbitrarily far from equilibrium, when granted full external control to drive the system. However, in both simulation and experimental contexts, systems often may only be controlled with a limited set of degrees of freedom. Here, going beyond slow- and fast-driving approximations employed in previous studies, we obtain exact finite-time optimal protocols for this unexplored limited-control setting. By working with deterministic Fokker-Planck probability density time evolution, we can frame the work-minimizing protocol problem in the standard form of an optimal control theory problem. We demonstrate that finding the exact optimal protocol is equivalent to solving a system of Hamiltonian partial differential equations, which in many cases admit efficiently calculatable numerical solutions. Within this framework, we reproduce analytical results for the optimal control of harmonic potentials, and numerically devise novel optimal protocols for two anharmonic examples: varying the stiffness of a quartic potential, and linearly biasing a double-well potential. We confirm that these optimal protocols outperform other protocols produced through previous methods, in some cases by a substantial amount. We find that for the linearly biased double-well problem, the mean position under the optimal protocol travels at a near-constant velocity. Surprisingly, for a certain timescale and barrier height regime, the optimal protocol is also non-monotonic in time.

cond-mat.stat-mech

Optimal finite-time Brownian Carnot engine

Recent advances in experimental control of colloidal systems have spurred a revolution in the production of mesoscale thermodynamic devices. Functional "textbook" engines, such as the Stirling and Carnot cycles, have been produced in colloidal systems where they operate far from equilibrium. Simultaneously, significant theoretical advances have been made in the design and analysis of such devices. Here, we use methods from thermodynamic geometry to characterize the optimal finite-time, nonequilibrium cyclic operation of the parametric harmonic oscillator contact with a time-varying heat bath, with particular focus on the Brownian Carnot cycle. We derive the optimally parametrized Carnot cycle, along with two other new cycles and compare their dissipated energy, efficiency, and steady-state power production against each other and a previously tested experimental protocol for the Carnot cycle. We demonstrate a 20\% improvement in dissipated energy over previous experimentally tested protocols and a $\sim$50\% improvement under other conditions for one of our engines, while our final engine is more efficient and powerful than the others we considered. Our results provide the means for experimentally realizing optimal mesoscale heat engines.

cond-mat.stat-mech

A geometric bound on the efficiency of irreversible thermodynamic cycles

Stochastic thermodynamics has revolutionized our understanding of heat engines operating in finite time. Recently, numerous studies have considered the optimal operation of thermodynamic cycles acting as heat engines with a given profile in thermodynamic space (e.g. $P-V$ space in classical thermodynamics), with a particular focus on the Carnot engine. In this work, we use the lens of thermodynamic geometry to explore the full space of thermodynamic cycles with continuously-varying bath temperature in search of optimally shaped cycles acting in the slow-driving regime. We apply classical isoperimetric inequalities to derive a universal geometric bound on the efficiency of any irreversible thermodynamic cycle and explicitly construct efficient heat engines operating in finite time that nearly saturate this bound for a specific model system. Given the bound, these optimal cycles perform more efficiently than all other thermodynamic cycles operating as heat engines in finite time, including notable cycles, such as those of Carnot, Stirling, and Otto. For example, in comparison to recent experiments, this corresponds to orders of magnitude improvement in the efficiency of engines operating in certain time regimes. Our results suggest novel design principles for future mesoscopic heat engines and are ripe for experimental investigation.

cond-mat.stat-mech

Solution to the Fokker-Planck equation for slowly driven Brownian motion: Emergent geometry and a formula for the corresponding thermodynamic metric

Considerable progress has recently been made with geometrical approaches to understanding and controlling small out-of-equilibrium systems, but a mathematically rigorous foundation for these methods has been lacking. Towards this end, we develop a perturbative solution to the Fokker-Planck equation for one-dimensional driven Brownian motion in the overdamped limit enabled by the spectral properties of the corresponding single-particle Schrödinger operator. The perturbation theory is in powers of the inverse characteristic timescale of variation of the fastest varying control parameter, measured in units of the system timescale, which is set by the smallest eigenvalue of the corresponding Schrödinger operator. It applies to any Brownian system for which the Schrödinger operator has a confining potential. We use the theory to rigorously derive an exact formula for a Riemannian "thermodynamic" metric in the space of control parameters of the system. We show that up to second-order terms in the perturbation theory, optimal dissipation-minimizing driving protocols minimize the length defined by this metric. We also show that a previously proposed metric is calculable from our exact formula with corrections that are exponentially suppressed in a characteristic length scale. We illustrate our formula using the two-dimensional example of a harmonic oscillator with time-dependent spring constant in a time-dependent electric field. Lastly, we demonstrate that the Riemannian geometric structure of the optimal control problem is emergent; it derives from the form of the perturbative expansion for the probability density and persists to all orders of the expansion.

cond-mat.stat-mech

Engineered swift equilibration for arbitrary geometries

Engineered swift equilibration (ESE) is a class of driving protocols that enforce an equilibrium distribution with respect to external control parameters at the beginning and end of rapid state transformations of open, classical non-equilibrium systems. ESE protocols have previously been derived and experimentally realized for Brownian particles in simple, one-dimensional, time-varying trapping potentials; one recent study considered ESE in two-dimensional Euclidean configuration space. Here we extend the ESE framework to generic, overdamped Brownian systems in arbitrary curved configuration space and illustrate our results with specific examples not amenable to previous techniques. Our approach may be used to impose the necessary dynamics to control the full temporal configurational distribution in a wide variety of experimentally realizable settings.

cond-mat.stat-mech

Stochastic optimization for learning quantum state feedback control

High fidelity state preparation represents a fundamental challenge in the application of quantum technology. While the majority of optimal control approaches use feedback to improve the controller, the controller itself often does not incorporate explicit state dependence. Here, we present a general framework for training deep feedback networks for open quantum systems with quantum nondemolition measurement that allows a variety of system and control structures that are prohibitive by many other techniques and can in effect react to unmodeled effects through nonlinear filtering. We demonstrate that this method is efficient due to inherent parallelizability, robust to open system interactions, and outperforms landmark state feedback control results in simulation.

quant-ph