Searcharxiv⌕ Search

arXiv subjects

Grant M. Rotskoff

Publications and source records attributed to Grant M. Rotskoff.

At least 19 recordsLinked to original sources

Morphology and depletion force-based large-scale self-assembly of nanocubes on surface

Self-assembling nanoparticles is a highly efficient and facile way to form functional nano-, micro- and macrostructures. However, currently available methods lack precision and controllability in size, shape and composition, suffer from poor reproducibility and scalability, and require complex steps and expensive materials. Here, we present the uniform morphology-induced and depletion force-directed nanoparticle assembly on surface (MIDAS) method with gold nanocubes (AuNCs). Using this approach, morphology-sensitive depletion forces trigger shape-selective flocculation and assembly of the AuNCs with uniform size and shape. Importantly, the surface roughness-controlled substrate drives the large-scale formation of AuNC-assembled monolayers (2D AuNAMs) or three-dimensional AuNC-assembled multilayers (3D AuNAMs) in a highly specific manner without any complex ligand modification or preparation steps. Lattice-gas modeling and kinetic Monte Carlo simulations show the depletion force is the key parameter determining supercrystal morphology and corroborate the experimental findings. This work establishes a generalizable nanoparticle assembly mechanism and demonstrates the broad applicability of the MIDAS strategy for scalable fabrication and patterning of 2D and 3D nanoparticle architectures.

cond-mat.mtrl-sci↗

Optimal parameterization of nonequilibrium generalized master equations from discrete-time experimental data

Kinetic analyses of experiments often require coarse-grained descriptions, but complex systems rarely conform to the widely used modeling assumptions of Markovianity and thermodynamic equilibrium. Memory is indeed a general and often inevitable consequence of coarse-graining. Markov state models (MSMs) are a popular choice of coarse-grained description, but require microstate assignments -- which are rarely experimentally tunable -- to macrostates that minimize memory. Generalized master equations (GMEs) circumvent this limitation of MSMs by explicitly capturing memory. However, GMEs are difficult to parameterize and usually formally approximate in the experimentally relevant discrete-time setting. Here we introduce a maximum-likelihood-based procedure to parameterize formally exact, physically feasible, discrete-time generalized master equations from experiments and simulations in and out of equilibrium. By adapting algorithms typically used in optimal transport, we construct physical-constraint-satisfying conditional-maximum-likelihood estimators of both exact Nakajima-Zwanzig memory kernels and time-convolutionless GME propagators in discrete time. Applying these estimators to three examples -- experimental recordings of Förster-resonance energy-transfer in an ion channel, experimental nanoparticle tracking of a processive molecular motor, and simulated folding of a benchmark protein domain -- we recover kinetic parameters including relaxation rates, irreversibilities, dwell times, and first-passage times. These results establish discrete-time GMEs as a physically and statistically principled alternative to MSMs for kinetic analyses of experimental and simulated biomolecular systems.

cond-mat.stat-mech↗

Pushing the limits of one-dimensional NMR spectroscopy for automated structure elucidation using artificial intelligence

One-dimensional NMR spectroscopy is one of the most widely used techniques for the characterization of organic compounds and natural products. For molecules with up to 36 non-hydrogen atoms, the number of possible structures has been estimated to range from $10^{20} - 10^{60}$. The task of determining the structure (formula and connectivity) of a molecule of this size using only its one-dimensional $^1$H and/or $^{13}$C NMR spectrum, i.e. de novo structure generation, thus appears completely intractable. Here we show how it is possible to achieve this task for systems with up to 40 non-hydrogen atoms across the full elemental coverage typically encountered in organic chemistry (C, N, O, H, P, S, Si, B, and the halogens) using a deep learning framework, thus covering a vast portion of the drug-like chemical space. Leveraging insights from natural language processing, we show that our transformer-based architecture predicts the correct molecule with 60.4% accuracy within the first 15 predictions using only the $^1$H and $^{13}$C NMR spectra, thus overcoming the combinatorial growth of the chemical space while also being extensible to experimental data via fine-tuning.

physics.chem-ph↗

A Unified Approach to Analysis and Design of Denoising Markov Models

Probabilistic generative models based on measure transport, such as diffusion and flow-based models, are often formulated in the language of Markovian stochastic dynamics, where the choice of the underlying process impacts both algorithmic design choices and theoretical analysis. In this paper, we aim to establish a rigorous mathematical foundation for denoising Markov models, a broad class of generative models that postulate a forward process transitioning from the target distribution to a simple, easy-to-sample distribution, alongside a backward process particularly constructed to enable efficient sampling in the reverse direction. Leveraging deep connections with nonequilibrium statistical mechanics and generalized Doob's $h$-transform, we propose a minimal set of assumptions that ensure: (1) explicit construction of the backward generator, (2) a unified variational objective directly minimizing the measure transport discrepancy, and (3) adaptations of the classical score-matching approach across diverse dynamics. Our framework unifies existing formulations of continuous and discrete diffusion models, identifies the most general form of denoising Markov models under certain regularity assumptions on forward generators, and provides a systematic recipe for designing denoising Markov models driven by arbitrary Lévy-type processes. We illustrate the versatility and practical effectiveness of our approach through novel denoising Markov models employing geometric Brownian motion and jump processes as forward dynamics, highlighting the framework's potential flexibility and capability in modeling complex distributions.

cs.LG↗

DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models

We study inference-time scaling for diffusion models, where the goal is to adapt a pre-trained model to new target distributions without retraining. Existing guidance-based methods are simple but introduce bias, while particle-based corrections suffer from weight degeneracy and high computational cost. We introduce DriftLite, a lightweight, training-free particle-based approach that steers the inference dynamics on the fly with provably optimal stability control. DriftLite exploits a previously unexplored degree of freedom in the Fokker-Planck equation between the drift and particle potential, and yields two practical instantiations: Variance- and Energy-Controlling Guidance (VCG/ECG) for approximating the optimal drift with minimal overhead. Across Gaussian mixture models, particle systems, and large-scale protein-ligand co-folding problems, DriftLite consistently reduces variance and improves sample quality over pure guidance and sequential Monte Carlo baselines. These results highlight a principled, efficient route toward scalable inference-time adaptation of diffusion models. Our source code is publicly available at https://github.com/yinuoren/DriftLite.

cs.LG↗

Scaling Transferable Coarse-graining with Mean Force Matching

Coarse-grained molecular dynamics often sacrifices accuracy and transferability for computational efficiency, but the use of machine learned potentials is helping coarse-grained models attain performance on par with atomistic molecular dynamics. Nevertheless, developing representations of the coarse-grained potential energy surface faces severe scaling challenges due to the extreme data demands of widely used "bottom-up" coarse-graining objectives. In this work, we show that mean force matching, a strategy for training thermodynamically consistent coarse-grained models, requires 50x fewer training samples and 87% less total atomistic simulation time, while obtaining better accuracy on the potential of mean force for unseen proteins compared to other commonly used objectives. By systematically removing noise from the objective function, we demonstrate that it is possible to scale machine learning architectures for coarse-graining, enabling highly accurate and transferable models. We show the advantages of mean force matching both theoretically and through exhaustive benchmarking using thermodynamic consistency as the primary metric of accuracy.

physics.chem-ph↗

The cost of quantum algorithms for biochemistry: A case study in metaphosphate hydrolysis

We evaluate the quantum resource requirements for ATP/metaphosphate hydrolysis, one of the most important reactions in all of biology with implications for metabolism, cellular signaling, and cancer therapeutics. In particular, we consider three algorithms for solving the ground state energy estimation problem: the variational quantum eigensolver, quantum Krylov, and quantum phase estimation. By utilizing exact classical simulation, numerical estimation, and analytical bounds, we provide a current and future outlook for using quantum computers to solve impactful biochemical and biological problems. Our results show that variational methods, while being the most heuristic, still require substantially fewer overall resources on quantum hardware, and could feasibly address such problems on current or near-future devices. We include our complete dataset of biomolecular Hamiltonians and code as benchmarks to improve upon with future techniques.

quant-ph↗

Accelerating Diffusion Models with Parallel Sampling: Inference at Sub-Linear Time Complexity

Diffusion models have become a leading method for generative modeling of both image and scientific data. As these models are costly to train and \emph{evaluate}, reducing the inference cost for diffusion models remains a major goal. Inspired by the recent empirical success in accelerating diffusion models via the parallel sampling technique~\cite{shih2024parallel}, we propose to divide the sampling process into $\mathcal{O}(1)$ blocks with parallelizable Picard iterations within each block. Rigorous theoretical analysis reveals that our algorithm achieves $\widetilde{\mathcal{O}}(\mathrm{poly} \log d)$ overall time complexity, marking \emph{the first implementation with provable sub-linear complexity w.r.t. the data dimension $d$}. Our analysis is based on a generalized version of Girsanov's theorem and is compatible with both the SDE and probability flow ODE implementations. Our results shed light on the potential of fast and efficient sampling of high-dimensional data on fast-evolving modern large-memory GPU clusters.

cs.LG↗

Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order Algorithms

Discrete diffusion models have emerged as a powerful generative modeling framework for discrete data with successful applications spanning from text generation to image synthesis. However, their deployment faces challenges due to the high dimensionality of the state space, necessitating the development of efficient inference algorithms. Current inference approaches mainly fall into two categories: exact simulation and approximate methods such as $τ$-leaping. While exact methods suffer from unpredictable inference time and redundant function evaluations, $τ$-leaping is limited by its first-order accuracy. In this work, we advance the latter category by tailoring the first extension of high-order numerical inference schemes to discrete diffusion models, enabling larger step sizes while reducing error. We rigorously analyze the proposed schemes and establish the second-order accuracy of the $θ$-Trapezoidal method in KL divergence. Empirical evaluations on GSM8K-level math-reasoning, GPT-2-level text, and ImageNet-level image generation tasks demonstrate that our method achieves superior sample quality compared to existing approaches under equivalent computational constraints, with consistent performance gains across models ranging from 200M to 8B. Our code is available at https://github.com/yuchen-zhu-zyc/DiscreteFastSolver.

cs.LG↗

Quantifying and minimizing dissipation in a non-equilibrium phase transition

In a finite-time continuous phase transition, topological defects emerge as the system undergoes spontaneous symmetry breaking. The Kibble-Zurek mechanism predicts how the defect density scales with the quench rate. During such processes, dissipation also arises as the system fails to adiabatically follow the control protocol near the critical point. Quantifying and minimizing this dissipation is fundamentally relevant to nonequilibrium thermodynamics and practically important for energy-efficient computing and devices. However, there are no prior experimental measurements of dissipation, or the optimization of control protocols to reduce it in many-body systems. In addition, it is an open question to what extent dissipation is correlated with the formation of defects. Here, we directly measure the dissipation generated during the voltage-driven Freedericksz transition of a liquid crystal with a sensitivity equivalent to a ~10 nanokelvin temperature rise. We observe Kibble-Zurek scaling of dissipation and its breakdown, both in quantitative agreement with existing theoretical works. We further implement a fully automated in-situ optimization approach that discovers more optimal driving protocols, reducing dissipation by a factor of three relative to a simple linear protocol.

cond-mat.stat-mech↗

Aligning Transformers with Continuous Feedback via Energy Rank Alignment

Searching through chemical space is an exceptionally challenging problem because the number of possible molecules grows combinatorially with the number of atoms. Large, autoregressive models trained on databases of chemical compounds have yielded powerful generators, but we still lack robust strategies for generating molecules with desired properties. This molecular search problem closely resembles the "alignment" problem for large language models, though for many chemical tasks we have a specific and easily evaluable reward function. Here, we introduce an algorithm called energy rank alignment (ERA) that leverages an explicit reward function to produce a gradient-based objective that we use to optimize autoregressive policies. We show theoretically that this algorithm is closely related to proximal policy optimization (PPO) and direct preference optimization (DPO), but has a minimizer that converges to an ideal Gibbs-Boltzmann distribution with the reward playing the role of an energy function. Furthermore, this algorithm is highly scalable, does not require reinforcement learning, and performs well relative to DPO when the number of preference observations per pairing is small. We deploy this approach to align molecular transformers and protein language models to generate molecules and protein sequences, respectively, with externally specified properties and find that it does so robustly, searching through diverse parts of chemical space.

cs.LG↗

Features are fate: a theory of transfer learning in high-dimensional regression

With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of "task similarity" is still lacking. While conventional wisdom suggests that simple measures of similarity between source and target distributions, such as $ϕ$-divergences or integral probability metrics, can directly predict the success of transfer, we prove the surprising fact that, in general, this is not the case. We adopt, instead, a feature-centric viewpoint on transfer learning and establish a number of theoretical results that demonstrate that when the target task is well represented by the feature space of the pre-trained model, transfer learning outperforms training from scratch. We study deep linear networks as a minimal model of transfer learning in which we can analytically characterize the transferability phase diagram as a function of the target dataset size and the feature space overlap. For this model, we establish rigorously that when the feature space overlap between the source and target tasks is sufficiently strong, both linear transfer and fine-tuning improve performance, especially in the low data limit. These results build on an emerging understanding of feature learning dynamics in deep linear networks, and we demonstrate numerically that the rigorous results we derive for the linear case also apply to nonlinear networks.

stat.ML↗

Non-Clifford diagonalization for measurement shot reduction in quantum expectation value estimation

Estimating expectation values on near-term quantum computers often requires a prohibitively large number of measurements. One widely-used strategy to mitigate this problem has been to partition an operator's Pauli terms into sets of mutually commuting operators. Here, we introduce a method that relaxes this constraint of commutativity, instead allowing for entirely arbitrary terms to be grouped together, save a locality constraint. The key idea is that we decompose the operator into arbitrary tensor products with bounded tensor size, ignoring Pauli commuting relations. This method -- named $k$-NoCliD ($k$-local non-Clifford diagonalization) -- allows one to measure in far fewer bases in most cases, often (though not always) at the cost of increasing the circuit depth. We introduce several partitioning algorithms tailored to different Hamiltonian classes. For electronic structure, we numerically demonstrate the existence of threshold values of $k$ for which $k$-NoCliD leads to the lowest shot counts, though we leave improved partitioning algorithms to future work. We focus primarily on three Hamiltonian classes -- molecular vibrational structure, Fermi-Hubbard, and Bose-Hubbard -- and show that $k$-NoCliD reduces the number of circuit shots, often by a very large margin, and often even for $k$ as small as 2.

quant-ph↗

Minimally dissipative multi-bit logical operations

Modern computing architectures are vastly more energy-dissipative than fundamental thermodynamic limits suggest, motivating the search for principled approaches to low-dissipation logical operations. We formulate multi-bit logical gates (bit erasure, NAND) as optimal transport problems, extending beyond classical one-dimensional bit erasure to scenarios where existing methods fail. Using entropically regularized unbalanced optimal transport, we derive tractable solutions and establish general energy-speed-accuracy trade-offs that demonstrate that faster, more accurate operations necessarily dissipate more energy. Furthermore, we demonstrate that the Landauer limits cannot be trivially overcome in higher dimensional geometries. We develop practical algorithms combining optimal transport with generative modeling techniques to construct dynamical controllers that follow Wasserstein geodesics. These protocols achieve near-optimal dissipation and can, in principle, be implemented in realistic experimentally set-ups. The framework bridges fundamental thermodynamic limits with scalable computational design for energy-efficient information processing.

cond-mat.stat-mech↗

How Discrete and Continuous Diffusion Meet: Comprehensive Analysis of Discrete Diffusion Models via a Stochastic Integral Framework

Discrete diffusion models have gained increasing attention for their ability to model complex distributions with tractable sampling and inference. However, the error analysis for discrete diffusion models remains less well-understood. In this work, we propose a comprehensive framework for the error analysis of discrete diffusion models based on Lévy-type stochastic integrals. By generalizing the Poisson random measure to that with a time-independent and state-dependent intensity, we rigorously establish a stochastic integral formulation of discrete diffusion models and provide the corresponding change of measure theorems that are intriguingly analogous to Itô integrals and Girsanov's theorem for their continuous counterparts. Our framework unifies and strengthens the current theoretical results on discrete diffusion models and obtains the first error bound for the $τ$-leaping scheme in KL divergence. With error sources clearly identified, our analysis gives new insight into the mathematical properties of discrete diffusion models and offers guidance for the design of efficient and accurate algorithms for real-world discrete diffusion model applications.

cs.LG↗

Universal energy-speed-accuracy trade-offs in driven nonequilibrium systems

Physical systems driven away from equilibrium by an external controller dissipate heat to the environment; the excess entropy production in the thermal reservoir can be interpreted as a "cost" to transform the system in a finite time. The connection between measure theoretic optimal transport and dissipative nonequilibrium dynamics provides a language for quantifying this cost and has resulted in a collection of "thermodynamic speed limits", which argue that the minimum dissipation of a transformation between two probability distributions is directly proportional to the rate of driving. Thermodynamic speed limits rely on the assumption that the target probability distribution is perfectly realized, which is almost never the case in experiments or numerical simulations. Here, we address the ubiquitous situation in which the external controller is imperfect. As a consequence, we obtain a lower bound for the dissipated work in generic nonequilibrium control problems that 1) is asymptotically tight and 2) matches the thermodynamic speed limit in the case of optimal driving. We illustrate these bounds on analytically solvable examples and also develop a strategy for optimizing minimally dissipative protocols based on optimal transport flow matching, a generative machine learning technique. This latter approach ensures the scalability of both the theoretical and computational framework we put forth. Crucially, we demonstrate that we can compute the terms in our bound numerically using efficient algorithms from the computational optimal transport literature and that the protocols that we learn saturate the bound.

cond-mat.stat-mech↗

Scaling field-theoretic simulation for multi-component mixtures with neural operators

Multi-component polymer mixtures are ubiquitous in biological self-organization but are notoriously difficult to study computationally. Plagued by both slow single molecule relaxation times and slow equilibration within dense mixtures, molecular dynamics simulations are typically infeasible at the spatial scales required to study the stability of mesophase structure. Polymer field theories offer an attractive alternative, but analytical calculations are only tractable for mean-field theories and nearby perturbations, constraints that become especially problematic for fluctuation-induced effects such as coacervation. Here, we show that a recently developed technique for obtaining numerical solutions to partial differential equations based on operator learning, *neural operators*, lends itself to a highly scalable training strategy by parallelizing per-species operator maps. We illustrate the efficacy of our approach on six-component mixtures with randomly selected compositions and that it significantly outperforms the state-of-the-art pseudospectral integrators for field-theoretic simulations, especially as polymer lengths become long.

cond-mat.soft↗

Coacervation drives morphological diversity of mRNA encapsulating nanoparticles

The spatial arrangement of components within an mRNA encapsulating nanoparticle has consequences for its thermal stability, which is a key parameter for therapeutic utility. The mesostructure of mRNA nanoparticles formed with cationic polymers have several distinct putative structures: here, we develop a field theoretic simulation model to compute the phase diagram for amphiphilic block copolymers that balance coacervation and hydrophobicity as driving forces for assembly. We predict several distinct morphologies for the mesostructure of these nanoparticles, depending on salt conditions and hydrophobicity. We compare our predictions with cryogenic-electron microscopy images of mRNA encapsulated by charge altering releasable transporters. In addition, we provide a GPU-accelerated, open-source codebase for general purpose field theoretic simulations, which we anticipate will be a useful tool for the community.

cond-mat.soft↗