SearcharxivSearch

arXiv subjects

Yuhai Tu

Publications and source records attributed to Yuhai Tu.

At least 19 recordsLinked to original sources

Stochastic Thermodynamics of Score Matching in Diffusion Models

Score-based diffusion models are a powerful class of generative AI systems capable of sampling from complex, high-dimensional probability distributions. Their dynamics consist of a forward diffusion process that transforms data into noise and a learned reverse process that reconstructs data by reversing the probability flow. Here, we develop a stochastic thermodynamic framework for diffusion models and their score-matching objective. We introduce a trajectory-dependent quantity, time-asymmetry entropy production (TAEP), defined from the forward and reverse diffusion dynamics, and show that it obeys exact fluctuation theorems. Remarkably, Hyv\"{a}rinen's implicit score-matching kernel emerges naturally as a fluctuating component of TAEP, while the average TAEP is exactly proportional to the score-matching objective. We further show that fluctuations of TAEP quantify sampling unevenness and provide a thermodynamic measure of data-manifold coverage. These results yield a quantitative explanation for the superior sampling diversity of diffusion models and reveal a thermodynamic mechanism by which stochastic gradient descent favors flatter, more generalizable solutions. By uncovering the entropic nature of score matching, our work establishes fundamental statistical-mechanical principles underlying diffusion-based generative AI.

cond-mat.dis-nn

Contact-Dependent Ion Gating Explains Directional Asymmetry in the Bacterial Flagellar Motor

The bacterial flagellar motor (BFM) is a rotary molecular machine driven by the ion electrochemical potential across the cell membrane. Recent cryo-EM structures reveal a cogwheel-like architecture in which multiple stators engage a large rotor. A longstanding puzzle is the directional asymmetry of its torque-speed relation: concave in counterclockwise (CCW) rotation but nearly linear in clockwise (CW) rotation. Here, we develop a stochastic mechanochemical model that explicitly incorporates rotor-stator coupling and detailed ion translocation kinetics. By integrating physiological torque-speed data with recent measurements of rotor-stator relative motion, we show that under physiological conditions the motor operates in a tight engagement regime, rendering the torque-speed relation largely insensitive to the specific form of mechanical interactions. This finding rules out differences in rotor-stator mechanics as the origin of CW-CCW asymmetry. Guided by cryo-EM structures, we propose a contact-dependent gating mechanism in which the MotA-FliG interaction modulates the ion release rate of the MotB subunit proximal to the FliG ring. Molecular dynamics simulations indicate tighter MotA-FliG contact in the CW motor, implying a reduced ion release rate compared to CCW. Our model demonstrates that differential gating strength accounts for the observed asymmetry: stronger gating in CCW shortens torque-free waiting phases, enhances torque generation, and produces a concave torque-speed curve, whereas weaker gating in CW yields lower torque and a linear relation. This structure-based framework quantitatively links molecular asymmetry to motor function and identifies specific interfaces for targeted perturbation and mutational studies.

physics.bio-ph

On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature

Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima. Prior work often assumes an equivalence between the Fisher Information Matrix and the Hessian for negative log-likelihood losses, leading to the claim that the SGD noise covariance $\mathbf{C}$ is proportional to the Hessian $\mathbf{H}$. We show that this assumption holds only under restrictive conditions that are typically violated in deep neural networks. Using the recently discovered Activity--Weight Duality, we find a more general relationship agnostic to the specific loss formulation, showing that $\mathbf{C} \propto \mathbb{E}_p[\mathbf{h}_p^2]$, where $\mathbf{h}_p$ denotes the per-sample Hessian with $\mathbf{H} = \mathbb{E}_p[\mathbf{h}_p]$. As a consequence, $\mathbf{C}$ and $\mathbf{H}$ commute approximately rather than coincide exactly. We further find that, within the analyzed fully connected layers, their diagonal elements follow per-layer empirical power laws $C_{ii} \propto H_{ii}^{\gamma}$, with layer-dependent fitted exponents bounded by $1 \leq \gamma \leq 2$. Experiments across datasets, architectures, and loss functions support the resulting layerwise bounds, providing a unified characterization of the noise-curvature relationship in deep learning.

cs.LG

Noise-Driven Exploration and Transient Freezing Select Flat Minima in Stochastic Gradient Descent

Stochastic gradient descent (SGD) is central to deep learning, yet the dynamical origin of its preference for flatter, more generalizable solutions remains unclear. Here, by analyzing SGD learning dynamics, we identify a nonequilibrium mechanism that governs solution selection during training. Numerical experiments reveal a transient exploratory phase in which SGD trajectories repeatedly escape sharp valleys and migrate toward flatter regions of the loss landscape before becoming confined to a final basin. Using a tractable physical model, we show that SGD noise reshapes the loss landscape into an effective potential that preferentially stabilizes flat solutions. We further uncover a transient freezing mechanism: as training progresses, the flattening landscape suppresses transitions between competing valleys. Stronger SGD noise delays this freezing transition, prolonging the exploratory phase and thereby increasing the probability of convergence to flatter minima. Together, these results provide a unified physical framework connecting learning dynamics, loss-landscape geometry, and generalization, and suggest guiding principles for the design of more effective optimization algorithms.

cs.LG

Ultrasensitivity without conformational spread: A mechanical origin for non-equilibrium cooperativity in the bacterial flagellar motor

Flagellar motors enable bacteria to navigate their environments by switching rotation direction in response to external cues with high sensitivity. Previous work suggested that ultrasensitivity of the flagellar motor originates from conformational spread, in which subunits of the switching complex are strongly coupled to their neighbors as in an equilibrium Ising model. However, dynamic single-motor measurements indicated that rotation switching is driven out of equilibrium, and the mechanism for this dissipative driving remains unknown. Here, based on recent cryo-EM structures, we propose that local mechanical torques on motor subunits can affect their conformation dynamics. This gives rise to a tug of war between stator-associated subunits, which produces cooperative, non-equilibrium switching responses without requiring nearest-neighbor interactions. Since subunits are effectively coupled at a distance, we call this mechanism ``Global Mechanical Coupling." Our model makes a qualitatively new prediction that the motor response cooperativity grows with the number of stators driving rotation. Re-analyzing published motor dose-response curves in varying load conditions, we find tentative experimental evidence for this prediction. Finally, we show that operating out of equilibrium enables motors to achieve high cooperativity with faster responses compared to equilibrium motors. Our results suggest a general role for mechanics in sensitive chemical regulation.

physics.bio-ph

An altruistic resource-sharing mechanism for synchronization: The energy-speed-accuracy tradeoff

Synchronization among a group of active agents is ubiquitous in nature. Although synchronization based on direct interactions between agents described by the Kuramoto model is well understood, the other general mechanism based on indirect interactions among agents sharing limited resources are less known. Here, we propose a minimal thermodynamically consistent model for the altruistic resource-sharing (ARS) mechanism wherein resources are needed for individual agent to advance but a more advanced agent has a lower competence to obtain resources. We show that while differential competence in ARS mechanism provides a negative feedback leading to synchronization it also breaks detailed balance and thus requires additional energy dissipation besides the cost of driving individual agents. By solving the model analytically, our study reveals a general tradeoff relation between the total energy dissipation rate and the two key performance measures of the system: average speed and synchronization accuracy. For a fixed dissipation rate, there is a distinct speed-accuracy Pareto front traversed by the scarcity of resources: scarcer resources lead to slower speed but more accurate synchronization. Increasing energy dissipation eases this tradeoff by pushing the speed-accuracy Pareto front outwards. The connections of our work to realistic biological systems such as the KaiABC system in cyanobacterial circadian clock and other theoretical results based on thermodynamic uncertainty relation are also discussed.

cond-mat.stat-mech

Representational Drift and Learning-Induced Stabilization in the Olfactory Cortex

The brain encodes external stimuli through patterns of neural activity, forming internal representations of the world. Recent experiments show that neural representations for a given stimulus change over time. However, the mechanistic origin for the observed "representational drift" (RD) remains unclear. Here, we propose a biologically-realistic computational model of the piriform cortex to study RD in the mammalian olfactory system by combining two mechanisms for the dynamics of synaptic weights at two separate timescales: spontaneous fluctuations on a scale of days and spike-time dependent plasticity (STDP) on a scale of seconds. Our study shows that, while spontaneous fluctuations in synaptic weights induce RD, STDP-based learning during repeated stimulus presentations can reduce it. Our model quantitatively explains recent experiments on RD in the olfactory system and offers a mechanistic explanation for the emergence of drift and its relation to learning, which may be useful to study RD in other brain regions.

q-bio.NC

Solution landscape of reaction-diffusion systems reveals a nonlinear mechanism and spatial robustness of pattern formation

Spontaneous pattern formation in homogeneous systems is ubiquitous in nature. Although Turing demonstrated that spatial patterns can emerge in reaction-diffusion (RD) systems when the homogeneous state becomes linearly unstable, it remains unclear whether the Turing mechanism is the only route for pattern formation. Here, we develop an efficient algorithm to systematically map the solution landscape to find all steady-state solutions. By applying our method to generic RD models, we find that stable spatial patterns can emerge via saddle-node bifurcations before the onset of Turing instability. Furthermore, by using a generalized action in functional space based on large deviation theory, our method is extended to evaluate stability of spatial patterns against noise. Applying this general approach in a three-species RD model, we show that though formation of Turing patterns only requires two chemical species, the third species is critical for stabilizing patterns against strong intrinsic noise in small biochemical systems.

physics.bio-ph

Geometry of optimal control in chemical reaction networks

Although optimal control (OC) has been studied in stochastic thermodynamics for systems with continuous state variables, less is known in systems with discrete state variables, such as Chemical Reaction Networks (CRNs). Here, we develop a general theoretical framework to study OC of CRNs for changing the system from an initial distribution of states to a final distribution with minimum dissipation. We derive a ``Kirchhoff's law" for the probability current in the adiabatic limit, from which the optimal kinetic rates are determined analytically for any given probability trajectory. By using the optimal rates, we show that the total dissipation is determined by a $L_2$-distance measure in the probability space and derive an analytical expression for the metric tensor that depends on the probability distribution, network topology, and capacity of each link. Minimizing the total dissipation leads to the geodesic trajectory in the probability space and the corresponding OC protocol is determined by the Kirchhoff's law. To demonstrate our general approach, we use it to find a lower bound for the minimum dissipation that is tighter than existing bounds obtained with only global constraints. We also apply it to simple networks, e.g., fully connected 3-state CRNs with different local constraints and show that indirect pathway and non-functional transient state can play a crucial role in switching between different probability distributions efficiently. Future directions in studying OC in CRNs by using our general framework are discussed.

cond-mat.stat-mech

One nose but two nostrils: Learn to align with sparse connections between two olfactory cortices

The integration of neural representations in the two hemispheres is an important problem in neuroscience. Recent experiments revealed that odor responses in cortical neurons driven by separate stimulation of the two nostrils are highly correlated. This bilateral alignment points to structured inter-hemispheric connections, but detailed mechanism remains unclear. Here, we hypothesized that continuous exposure to environmental odors shapes these projections and modeled it as online learning with local Hebbian rule. We found that Hebbian learning with sparse connections achieves bilateral alignment, exhibiting a linear trade-off between speed and accuracy. We identified an inverse scaling relationship between the number of cortical neurons and the inter-hemispheric projection density required for desired alignment accuracy, i.e., more cortical neurons allow sparser inter-hemispheric projections. We next compared the alignment performance of local Hebbian rule and the global stochastic-gradient-descent (SGD) learning for artificial neural networks. We found that although SGD leads to the same alignment accuracy with modestly sparser connectivity, the same inverse scaling relation holds. We showed that their similar performance originates from the fact that the update vectors of the two learning rules align significantly throughout the learning process. This insight may inspire efficient sparse local learning algorithms for more complex problems.

q-bio.NC

Temperature Compensation through Kinetic Regulation in Biochemical Oscillators

Nearly all circadian clocks maintain a period that is insensitive to temperature changes, a phenomenon known as temperature compensation (TC). Yet, it is unclear whether there is any common feature among different systems that exhibit TC. From a general timescale invariance, we show that TC relies on existence of certain period-lengthening reactions wherein the period of the system increases strongly with the rates in these reactions. By studying several generic oscillator models, we show that this counter-intuitive dependence is nonetheless a common feature of oscillators in the nonlinear (far-from-onset) regime where the oscillation can be separated into fast and slow phases. The increase of the period with the period-lengthening reaction rates occurs when the amplitude of the slow phase in the oscillation increases with these rates while the progression-speed in the slow phase is controlled by other rates of the system. The positive dependence of the period on the period-lengthening rates balances its inverse dependence on other kinetic rates in the system, which gives rise to robust TC in a wide range of parameters. We demonstrate the existence of such period-lengthening reactions and their relevance for TC in all four model systems we considered. Theoretical results for a model of the Kai system are supported by experimental data. A study of the energy dissipation also shows that better TC performance requires higher energy consumption. Our study unveils a general mechanism by which a biochemical oscillator achieves TC by operating at regimes far from the onset where period-lengthening reactions exist.

q-bio.MN

Time-reversal symmetry breaking in the chemosensory array reveals mechanisms for dissipation-enhanced cooperative sensing

The Escherichia coli chemoreceptors form an extensive array that achieves cooperative and adaptive sensing of extracellular signals. The receptors control the activity of histidine kinase CheA, which drives a nonequilibrium phosphorylation-dephosphorylation reaction cycle for response regulator CheY. Cooperativity and dissipation are both important aspects of chemotaxis signaling, yet their consequences have only been studied separately. Recent single-cell FRET measurements revealed that kinase activity of the array spontaneously switches between active and inactive states, with asymmetric switching times that signify time-reversal symmetry breaking in the underlying dynamics. Here, we present a nonequilibrium lattice model of the chemosensory array, which demonstrates that the observed asymmetric switching dynamics can only be explained by an interplay between the dissipative reactions within individual core units and the cooperative coupling between neighboring units. Microscopically, the switching time asymmetry originates from irreversible transition paths. The model shows that strong dissipation enables sensitive and rapid signaling response by relieving the speed-sensitivity trade-off, which can be tested by future single-cell experiments. Overall, our model provides a general framework for studying biological complexes composed of coupled subunits that are individually driven by dissipative cycles and the rich nonequilibrium physics within.

physics.bio-ph

Stochastic dynamics of granular hopper flows: a configurational mode controls the stability of clogs

Granular flows in small-outlet hoppers exhibit several characteristic but poorly understood behaviors: temporary clogs (pauses) where flow stops before later spontaneously restarting, permanent clogs that last indefinitely, and non-Gaussian, non-monotonic flow-rate statistics. These aspects have been studied independently, but a model of hopper flow that includes all three has not been formulated. Here, we introduce a phenomenological model that provides a unifying dynamical explanation of all three behaviors: coupling between the flow rate and a hidden mode that controls the stability of clogs. In the theory, flow rate evolves according to Langevin dynamics with multiplicative noise and an absorbing state at zero flow, conditional on the hidden mode. The model fully reproduces the statistics of pause and clog events of a large ($>40,000$ flows) experimental dataset, including non-exponentially distributed clogging times and non-Gaussian flow rate distribution, and explains the stretched-exponential growth of the average clogging time with outlet size. Further, we identify the physical nature of the hidden mode in microscopic configurational features, including size and smoothness of the static arch structure formed during pauses and clogs. Our work provides a unifying framework for several poorly understood clogging phenomena, and suggests numerous new paths toward further understanding of this complex system.

cond-mat.soft

The activity-weight duality in feed forward neural networks: The geometric determinants of generalization

One of the fundamental problems in machine learning is generalization. In neural network models with a large number of weights (parameters), many solutions can be found to fit the training data equally well. The key question is which solution can describe testing data not in the training set. Here, we report the discovery of an exact duality (equivalence) between changes in activities in a given layer of neurons and changes in weights that connect to the next layer of neurons in a densely connected layer in any feed forward neural network. The activity-weight (A-W) duality allows us to map variations in inputs (data) to variations of the corresponding dual weights. By using this mapping, we show that the generalization loss can be decomposed into a sum of contributions from different eigen-directions of the Hessian matrix of the loss function at the solution in weight space. The contribution from a given eigen-direction is the product of two geometric factors (determinants): the sharpness of the loss landscape and the standard deviation of the dual weights, which is found to scale with the weight norm of the solution. Our results provide an unified framework, which we used to reveal how different regularization schemes (weight decay, stochastic gradient descent with different batch sizes and learning rates, dropout), training data size, and labeling noise affect generalization performance by controlling either one or both of these two geometric determinants for generalization. These insights can be used to guide development of algorithms for finding more generalizable solutions in overparametrized neural networks.

cs.LG

Maximizing Information in Domain-Invariant Representation Improves Transfer Learning

We propose MaxDIRep, a domain adaptation method that improves the decomposition of data representations into domain-independent and domain-dependent components. Existing methods, such as Domain-Separation Networks (DSN), use a weak orthogonality constraint between these components, which can lead to label-relevant features being partially encoded in the domain-dependent representation (DDRep) rather than the domain-independent representation (DIRep). As a result, information crucial for target-domain classification may be missing from the DIRep. MaxDIRep addresses this issue by applying a Kullback-Leibler (KL) divergence constraint to minimize the information content of the DDRep, thereby encouraging the DIRep to retain features that are both domain-invariant and predictive of target labels. Through geometric analysis and an ablation study on synthetic datasets, we show why DSN's weaker constraint can lead to suboptimal adaptation. Experiments on standard image benchmarks and a network intrusion detection task demonstrate that MaxDIRep achieves strong performance, works with pretrained models, and generalizes to non-image classification tasks.

cs.CV

Resolving the binding-kinase discrepancy in bacterial chemotaxis: A nonequilibrium allosteric model and the role of energy dissipation

The Escherichia coli chemotaxis signaling pathway has served as a model system for studying the adaptive sensing of environmental signals by large protein complexes. The chemoreceptors control the kinase activity of CheA in response to the extracellular ligand concentration and adapt across a wide concentration range by undergoing methylation and demethylation. Methylation shifts the kinase response curve by orders of magnitude in ligand concentration while incurring a much smaller change in the ligand binding curve. Here, we show that this asymmetric shift in binding and kinase response is inconsistent with equilibrium allosteric models regardless of parameter choices. To resolve this inconsistency, we present a nonequilibrium allosteric model that explicitly includes the dissipative reaction cycles driven by ATP hydrolysis. The model successfully explains all existing measurements for both aspartate and serine receptors. Our results suggest that while ligand binding controls the equilibrium balance between the ON and OFF states of the kinase, receptor methylation modulates the kinetic properties (e.g., the phosphorylation rate) of the ON state. Furthermore, sufficient energy dissipation is necessary for maintaining and enhancing the sensitivity range and amplitude of the kinase response. We demonstrate that the nonequilibrium allosteric model is broadly applicable to other sensor-kinase systems by successfully fitting previously unexplained data from the DosP bacterial oxygen-sensing system. Overall, this work provides a new perspective on cooperative sensing by large protein complexes and opens up new research directions for understanding their microscopic mechanisms through simultaneous measurements and modeling of ligand binding and downstream responses.

physics.bio-ph

The energy cost for flocking of active spins: the cusped dissipation maximum at the flocking transition

We study the energy cost of flocking in the active Ising model (AIM) and show that besides the energy cost for self-propelled motion, an additional energy dissipation is required to power the alignment of spins. We find that this additional alignment dissipation reaches its maximum at the flocking transition point in the form of a cusp with a discontinuous first derivative with respect to the control parameter. To understand this singular behavior, we analytically solve the two- and three-site AIM models and obtain the exact dependence of the alignment dissipation on the flocking order parameter and control parameter, which explains the cusped dissipation maximum at the flocking transition. Our results reveal a trade-off between the energy cost of the system and its performance measured by the flocking speed and sensitivity to external perturbations. This tradeoff relationship provides a new perspective for understanding the dynamics of natural flocks and designing optimal artificial flocking systems.

cond-mat.stat-mech

Effective Dynamics of Generative Adversarial Networks

Generative adversarial networks (GANs) are a class of machine-learning models that use adversarial training to generate new samples with the same (potentially very complex) statistics as the training samples. One major form of training failure, known as mode collapse, involves the generator failing to reproduce the full diversity of modes in the target probability distribution. Here, we present an effective model of GAN training, which captures the learning dynamics by replacing the generator neural network with a collection of particles in the output space; particles are coupled by a universal kernel valid for certain wide neural networks and high-dimensional inputs. The generality of our simplified model allows us to study the conditions under which mode collapse occurs. Indeed, experiments which vary the effective kernel of the generator reveal a mode collapse transition, the shape of which can be related to the type of discriminator through the frequency principle. Further, we find that gradient regularizers of intermediate strengths can optimally yield convergence through critical damping of the generator dynamics. Our effective GAN model thus provides an interpretable physical framework for understanding and improving adversarial training.

cond-mat.dis-nn