Searcharxiv⌕ Search

arXiv subjects

Thomas F. Miller III

Publications and source records attributed to Thomas F. Miller III.

At least 19 recordsLinked to original sources

Continuum Limit of Dendritic Deposition

Continuum models are commonly used to study dendritic deposition in fields ranging from nonequilibrium statistical mechanics to battery research. However, the continuum approximation underlying these models is poorly understood, even in the simplified case of Brownian particles depositing onto a small, reactive cluster. Specifically, this system transitions from a compact to a dendritic morphology at a critical radius that depends on the particle size. But in simulations of the continuum (small-particle) limit, the critical radius does not reproduce the scaling predicted by a purely continuum analysis. This discrepancy suggests that continuum models may not be able to capture the microscopic physics of dendrite formation, raising doubts about their experimental relevance. To clarify the continuum limit of dendritic deposition, here, we reexamine the critical radius scaling of the Brownian particle system using Brownian dynamics simulations. Compared to past studies, we probe larger system sizes, up to hundreds of millions of particles in some cases, and adopt an improved paradigm for the surface reaction. This paradigm allows us to converge our simulations and to work with well-defined physical parameters. Our results show that the critical radius scaling is, in fact, consistent with the continuum analysis, validating the continuum approach to modeling dendritic deposition. Nonetheless, the Brownian particle system converges to its continuum limit slowly. As a result, when applying continuum models to more complex deposition processes, the continuum approximation itself may be a significant source of error.

cond-mat.stat-mech↗

NeuralPLexer3: Accurate Biomolecular Complex Structure Prediction with Flow Models

Structure determination is essential to a mechanistic understanding of diseases and the development of novel therapeutics. Machine-learning-based structure prediction methods have made significant advancements by computationally predicting protein and bioassembly structures from sequences and molecular topology alone. Despite substantial progress in the field, challenges remain to deliver structure prediction models to real-world drug discovery. Here, we present NeuralPLexer3 -- a physics-inspired flow-based generative model that achieves state-of-the-art prediction accuracy on key biomolecular interaction types and improves training and sampling efficiency compared to its predecessors and alternative methodologies. Examined through newly developed benchmarking strategies, NeuralPLexer3 excels in vital areas that are crucial to structure-based drug design, such as physical validity and ligand-induced conformational changes.

cs.LG↗

State-specific protein-ligand complex structure prediction with a multi-scale deep generative model

The binding complexes formed by proteins and small molecule ligands are ubiquitous and critical to life. Despite recent advancements in protein structure prediction, existing algorithms are so far unable to systematically predict the binding ligand structures along with their regulatory effects on protein folding. To address this discrepancy, we present NeuralPLexer, a computational approach that can directly predict protein-ligand complex structures solely using protein sequence and ligand molecular graph inputs. NeuralPLexer adopts a deep generative model to sample the 3D structures of the binding complex and their conformational changes at an atomistic resolution. The model is based on a diffusion process that incorporates essential biophysical constraints and a multi-scale geometric deep learning system to iteratively sample residue-level contact maps and all heavy-atom coordinates in a hierarchical manner. NeuralPLexer achieves state-of-the-art performance compared to all existing methods on benchmarks for both protein-ligand blind docking and flexible binding site structure recovery. Moreover, owing to its specificity in sampling both ligand-free-state and ligand-bound-state ensembles, NeuralPLexer consistently outperforms AlphaFold2 in terms of global protein structure accuracy on both representative structure pairs with large conformational changes (average TM-score=0.93) and recently determined ligand-binding proteins (average TM-score=0.89). Case studies reveal that the predicted conformational variations are consistent with structure determination experiments for important targets, including human KRAS$^\textrm{G12C}$, ketol-acid reductoisomerase, and purine GPCRs. Our study suggests that a data-driven approach can capture the structural cooperativity between proteins and small molecules, showing promise in accelerating the design of enzymes, drug molecules, and beyond.

q-bio.QM↗

Exploring PROTAC cooperativity with coarse-grained alchemical methods

Proteolysis targeting chimera (PROTAC) is a novel drug modality that facilitates the degradation of a target protein by inducing proximity with an E3 ligase. In this work, we present a new computational framework to model the cooperativity between PROTAC-E3 binding and PROTAC-target binding principally through protein-protein interactions (PPIs) induced by the PROTAC. Due to the scarcity and low resolution of experimental measurements, the physical and chemical drivers of these non-native PPIs remain to be elucidated. We develop a coarse-grained (CG) approach to model interactions in the target-PROTAC-E3 complexes, which enables converged thermodynamic estimations using alchemical free energy calculation methods despite an unconventional scale of perturbations. With minimal parameterization, we successfully capture fundamental principles of cooperativity, including the optimality of intermediate PROTAC linker lengths that originates from configurational entropy. We qualitatively characterize the dependency of cooperativity on PROTAC linker lengths and protein charges and shapes. Minimal inclusion of sequence- and conformation-specific features in our current forcefield, however, limits quantitative modeling to reproduce experimental measurements, but further development of the CG model may allow for efficient computational screening to optimize PROTAC cooperativity.

physics.bio-ph↗

Molecular-orbital-based Machine Learning for Open-shell and Multi-reference Systems with Kernel Addition Gaussian Process Regression

We introduce a novel machine learning strategy, kernel addition Gaussian process regression (KA-GPR), in molecular-orbital-based machine learning (MOB-ML) to learn the total correlation energies of general electronic structure theories for closed- and open-shell systems by introducing a machine learning strategy. The learning efficiency of MOB-ML (KA-GPR) is the same as the original MOB-ML method for the smallest criegee molecule, which is a closed-shell molecule with multi-reference characters. In addition, the prediction accuracies of different small free radicals could reach the chemical accuracy of 1 kcal/mol by training on one example structure. Accurate potential energy surfaces for the H10 chain (closed-shell) and water OH bond dissociation (open-shell) could also be generated by MOB-ML (KA-GPR). To explore the breadth of chemical systems that KA-GPR can describe, we further apply MOB-ML to accurately predict the large benchmark datasets for closed- (QM9, QM7b-T, GDB-13-T) and open-shell (QMSpin) molecules.

physics.chem-ph↗

Molecular Dipole Moment Learning via Rotationally Equivariant Gaussian Process Regression with Derivatives in Molecular-orbital-based Machine Learning

This study extends the accurate and transferable molecular-orbital-based machine learning (MOB-ML) approach to modeling the contribution of electron correlation to dipole moments at the cost of Hartree-Fock computations. A molecular-orbital-based (MOB) pairwise decomposition of the correlation part of the dipole moment is applied, and these pair dipole moments could be further regressed as a universal function of molecular orbitals (MOs). The dipole MOB features consist of the energy MOB features and their responses to electric fields. An interpretable and rotationally equivariant Gaussian process regression (GPR) with derivatives algorithm is introduced to learn the dipole moment more efficiently. The proposed problem setup, feature design, and ML algorithm are shown to provide highly-accurate models for both dipole moment and energies on water and fourteen small molecules. To demonstrate the ability of MOB-ML to function as generalized density-matrix functionals for molecular dipole moments and energies of organic molecules, we further apply the proposed MOB-ML approach to train and test the molecules from the QM9 dataset. The application of local scalable GPR with Gaussian mixture model unsupervised clustering (GMM/GPR) scales up MOB-ML to a large-data regime while retaining the prediction accuracy. In addition, compared with literature results, MOB-ML provides the best test MAEs of 4.21 mDebye and 0.045 kcal/mol for dipole moment and energy models, respectively, when training on 110000 QM9 molecules. The excellent transferability of the resulting QM9 models is also illustrated by the accurate predictions for four different series of peptides.

physics.chem-ph↗

Accurate Molecular-Orbital-Based Machine Learning Energies via Unsupervised Clustering of Chemical Space

We introduce an unsupervised clustering algorithm to improve training efficiency and accuracy in predicting energies using molecular-orbital-based machine learning (MOB-ML). This work determines clusters via the Gaussian mixture model (GMM) in an entirely automatic manner and simplifies an earlier supervised clustering approach [J. Chem. Theory Comput., 15, 6668 (2019)] by eliminating both the necessity for user-specified parameters and the training of an additional classifier. Unsupervised clustering results from GMM have the advantage of accurately reproducing chemically intuitive groupings of frontier molecular orbitals and having improved performance with an increasing number of training examples. The resulting clusters from supervised or unsupervised clustering is further combined with scalable Gaussian process regression (GPR) or linear regression (LR) to learn molecular energies accurately by generating a local regression model in each cluster. Among all four combinations of regressors and clustering methods, GMM combined with scalable exact Gaussian process regression (GMM/GPR) is the most efficient training protocol for MOB-ML. The numerical tests of molecular energy learning on thermalized datasets of drug-like molecules demonstrate the improved accuracy, transferability, and learning efficiency of GMM/GPR over not only other training protocols for MOB-ML, i.e., supervised regression-clustering combined with GPR(RC/GPR) and GPR without clustering. GMM/GPR also provide the best molecular energy predictions compared with the ones from literature on the same benchmark datasets. With a lower scaling, GMM/GPR has a 10.4-fold speedup in wall-clock training time compared with scalable exact GPR with a training size of 6500 QM7b-T molecules.

physics.chem-ph↗

Informing Geometric Deep Learning with Electronic Interactions to Accelerate Quantum Chemistry

Predicting electronic energies, densities, and related chemical properties can facilitate the discovery of novel catalysts, medicines, and battery materials. By developing a physics-inspired equivariant neural network, we introduce a method to learn molecular representations based on the electronic interactions among atomic orbitals. Our method, OrbNet-Equi, leverages efficient tight-binding simulations and learned mappings to recover high fidelity quantum chemical properties. OrbNet-Equi models a wide spectrum of target properties with an accuracy consistently better than standard machine learning methods and a speed orders of magnitude greater than density functional theory. Despite only using training samples collected from readily available small-molecule libraries, OrbNet-Equi outperforms traditional methods on comprehensive downstream benchmarks that encompass diverse main-group chemical processes. Our method also describes interactions in challenging charge-transfer complexes and open-shell systems. We anticipate that the strategy presented here will help to expand opportunities for studies in chemistry and materials science, where the acquisition of experimental or reference training data is costly.

cs.LG↗

Equilibrium-nonequilibrium ring-polymer molecular dynamics for nonlinear spectroscopy

Two-dimensional Raman and hybrid terahertz/Raman spectroscopic techniques provide invaluable insight into molecular structure and dynamics of condensed-phase systems. However, corroborating experimental results with theory is difficult due to the high computational cost of incorporating quantum-mechanical effects in the simulations. Here, we present the equilibrium-nonequilibrium ring-polymer molecular dynamics (RPMD), a practical computational method that can account for nuclear quantum effects on the two-time response function of nonlinear optical spectroscopy. Unlike a recently developed approach based on the double Kubo transformed (DKT) correlation function, our method is exact in the classical limit, where it reduces to the established equilibrium-nonequilibrium classical molecular dynamics method. Using benchmark model calculations, we demonstrate the advantages of the equilibrium-nonequilibrium RPMD over classical and DKT-based approaches. Importantly, its derivation, which is based on the nonequilibrium RPMD, obviates the need for identifying an appropriate Kubo transformed correlation function and paves the way for applying real-time path-integral techniques to multidimensional spectroscopy.

physics.chem-ph↗

OrbNet: Deep Learning for Quantum Chemistry Using Symmetry-Adapted Atomic-Orbital Features

We introduce a machine learning method in which energy solutions from the Schrodinger equation are predicted using symmetry adapted atomic orbitals features and a graph neural-network architecture. \textsc{OrbNet} is shown to outperform existing methods in terms of learning efficiency and transferability for the prediction of density functional theory results while employing low-cost features that are obtained from semi-empirical electronic structure calculations. For applications to datasets of drug-like molecules, including QM7b-T, QM9, GDB-13-T, DrugBank, and the conformer benchmark dataset of Folmsbee and Hutchison, \textsc{OrbNet} predicts energies within chemical accuracy of DFT at a computational cost that is thousand-fold or more reduced.

physics.chem-ph↗

Molecular Energy Learning Using Alternative Blackbox Matrix-Matrix Multiplication Algorithm for Exact Gaussian Process

We present an application of the blackbox matrix-matrix multiplication (BBMM) algorithm to scale up the Gaussian Process (GP) training of molecular energies in the molecular-orbital based machine learning (MOB-ML) framework. An alternative implementation of BBMM (AltBBMM) is also proposed to train more efficiently (over four-fold speedup) with the same accuracy and transferability as the original BBMM implementation. The training of MOB-ML was limited to 220 molecules, and BBMM and AltBBMM scale the training of MOB-ML up by over 30 times to 6500 molecules (more than a million pair energies). The accuracy and transferability of both algorithms are examined on the benchmark datasets of organic molecules with 7 and 13 heavy atoms. These lower-scaling implementations of the GP preserve the state-of-the-art learning efficiency in the low-data regime while extending it to the large-data regime with better accuracy than other available machine learning works on molecular energies.

physics.chem-ph↗

OrbNet Denali: A machine learning potential for biological and organic chemistry with semi-empirical cost and DFT accuracy

We present OrbNet Denali, a machine learning model for electronic structure that is designed as a drop-in replacement for ground-state density functional theory (DFT) energy calculations. The model is a message-passing neural network that uses symmetry-adapted atomic orbital features from a low-cost quantum calculation to predict the energy of a molecule. OrbNet Denali is trained on a vast dataset of 2.3 million DFT calculations on molecules and geometries. This dataset covers the most common elements in bio- and organic chemistry (H, Li, B, C, N, O, F, Na, Mg, Si, P, S, Cl, K, Ca, Br, I) as well as charged molecules. OrbNet Denali is demonstrated on several well-established benchmark datasets, and we find that it provides accuracy that is on par with modern DFT methods while offering a speedup of up to three orders of magnitude. For the GMTKN55 benchmark set, OrbNet Denali achieves WTMAD-1 and WTMAD-2 scores of 7.19 and 9.84, on par with modern DFT functionals. For several GMTKN55 subsets, which contain chemical problems that are not present in the training set, OrbNet Denali produces a mean absolute error comparable to those of DFT methods. For the Hutchison conformers benchmark set, OrbNet Denali has a median correlation coefficient of R^2=0.90 compared to the reference DLPNO-CCSD(T) calculation, and R^2=0.97 compared to the method used to generate the training data (wB97X-D3/def2-TZVP), exceeding the performance of any other method with a similar cost. Similarly, the model reaches chemical accuracy for non-covalent interactions in the S66x10 dataset. For torsional profiles, OrbNet Denali reproduces the torsion profiles of wB97X-D3/def2-TZVP with an average MAE of 0.12 kcal/mol for the potential energy surfaces of the diverse fragments in the TorsionNet500 dataset.

physics.chem-ph↗

Stern and Diffuse Layer Interactions During Ionic Strength Cycling

Second harmonic generation amplitude and phase measurements are acquired in real time from fused silica:water interfaces that are subjected to ionic strength transitions conducted at pH 5.8. In conjunction with atomistic modeling, we identify correlations between structure in the Stern layer, encoded in the total second-order nonlinear susceptibility, chi(2)tot, and in the diffuse layer, encoded in the product of chi(2)tot and the total interfacial potential, phi(0)tot. chi(2)tot:phi(0)tot correlation plots indicate that the dynamics in the Stern and diffuse layers are decoupled from one another under some conditions (large change in ionic strength), while they change in lockstep under others (smaller change in ionic strength) as the ionic strength in the aqueous bulk solution varies. The quantitative structural and electrostatic information obtained also informs on the molecular origin of hysteresis in ionic strength cycling over fused silica. Atomistic simulations suggest a prominent role of contact ion pairs (as opposed to solvent-separated ion pairs) in the Stern layer. Those simulations also indicate that net water alignment is limited to the first 2 nm from the interface, even at 0 M ionic strength, highlighting water's polarization as an important contributor to nonlinear optical signal generation.

cond-mat.mtrl-sci↗

A New Imaginary Term in the 2nd Order Nonlinear Susceptibility from Charged Interfaces

Non-resonant second harmonic generation phase and amplitude measurements obtained from the silica:water interface at varying pH and 0.5 M ionic strength point to the existence of a nonlinear susceptibility term, which we call chi(3)X, that is associated with a 90 deg phase shift. Including this contribution in a model for the total effective second-order nonlinear susceptibility produces reasonable point estimates for interfacial potentials and second-order nonlinear susceptibilities when chi(3)Xis about 1.5 times chi(3)water. A model without this term and containing only traditional chi(2) and chi(3) terms cannot recapitulate the experimental data. The new model also provides a demonstrated utility for distinguishing apparent differences in the second-order nonlinear susceptibility when the electrolyte is NaCl vs MgSO4, pointing to the possibility of using HD-SHG to investigate ion-specificity in interfacial processes.

physics.chem-ph↗

Analytical Gradients for Molecular-Orbital-Based Machine Learning

Molecular-orbital-based machine learning (MOB-ML) enables the prediction of accurate correlation energies at the cost of obtaining molecular orbitals. Here, we present the derivation, implementation, and numerical demonstration of MOB-ML analytical nuclear gradients which are formulated in a general Lagrangian framework to enforce orthogonality, localization, and Brillouin constraints on the molecular orbitals. The MOB-ML gradient framework is general with respect to the regression technique (e.g., Gaussian process regression or neural networks) and the MOB feature design. We show that MOB-ML gradients are highly accurate compared to other ML methods on the ISO17 data set while only being trained on energies for hundreds of molecules compared to energies and gradients for hundreds of thousands of molecules for the other ML methods. The MOB-ML gradients are also shown to yield accurate optimized structures, at a computational cost for the gradient evaluation that is comparable to Hartree-Fock theory or hybrid DFT.

physics.chem-ph↗

Multi-task learning for electronic structure to predict and explore molecular potential energy surfaces

We refine the OrbNet model to accurately predict energy, forces, and other response properties for molecules using a graph neural-network architecture based on features from low-cost approximated quantum operators in the symmetry-adapted atomic orbital basis. The model is end-to-end differentiable due to the derivation of analytic gradients for all electronic structure terms, and is shown to be transferable across chemical space due to the use of domain-specific features. The learning efficiency is improved by incorporating physically motivated constraints on the electronic structure through multi-task learning. The model outperforms existing methods on energy prediction tasks for the QM9 dataset and for molecular geometry optimizations on conformer datasets, at a computational cost that is thousand-fold or more reduced compared to conventional quantum-chemistry calculations (such as density functional theory) that offer similar accuracy.

physics.chem-ph↗

A generalized class of strongly stable and dimension-free T-RPMD integrators

Recent work shows that strong stability and dimensionality freedom are essential for robust numerical integration of thermostatted ring-polymer molecular dynamics (T-RPMD) and path-integral molecular dynamics (PIMD), without which standard integrators exhibit non-ergodicity and other pathologies [J. Chem. Phys. 151, 124103 (2019); J. Chem. Phys. 152, 104102 (2020)]. In particular, the BCOCB scheme, obtained via Cayley modification of the standard BAOAB scheme, features a simple reparametrization of the free ring-polymer sub-step that confers strong stability and dimensionality freedom and has been shown to yield excellent numerical accuracy in condensed-phase systems with large time-steps. Here, we introduce a broader class of T-RPMD numerical integrators that exhibit strong stability and dimensionality freedom, irrespective of the Ornstein-Uhlenbeck friction schedule. In addition to considering equilibrium accuracy and time-step stability as in previous work, we evaluate the integrators on the basis of their rates of convergence to equilibrium and their efficiency at evaluating equilibrium expectation values. Within the generalized class, we find BCOCB to be superior with respect to accuracy and efficiency for various configuration-dependent observables, although other integrators within the generalized class perform better for velocity-dependent quantities. Extensive numerical evidence indicates that the stated performance guarantees hold for the strongly anharmonic case of liquid water. Both analytical and numerical results indicate that BCOCB excels over other known integrators in terms of accuracy, efficiency, and stability with respect to time-step for practical applications.

physics.chem-ph↗

Improved accuracy and transferability of molecular-orbital-based machine learning: Organics, transition-metal complexes, non-covalent interactions, and transition states

Molecular-orbital-based machine learning (MOB-ML) provides a general framework for the prediction of accurate correlation energies at the cost of obtaining molecular orbitals. We demonstrate the importance of preserving physical constraints, including invariance conditions and size consistency, when generating the input for the machine learning model. Numerical improvements are demonstrated for different data sets covering total and relative energies for thermally accessible organic and transition-metal containing molecules, non-covalent interactions, and transition-state energies. MOB-ML requires training data from only 1% of the QM7b-T data set (i.e., only 70 organic molecules with seven and fewer heavy atoms) to predict the total energy of the remaining 99% of this data set with sub-kcal/mol accuracy. This MOB-ML model is significantly more accurate than other methods when transferred to a data set comprised of thirteen heavy atom molecules, exhibiting no loss of accuracy on a size intensive (i.e., per-electron) basis. It is shown that MOB-ML also works well for extrapolating to transition-state structures, predicting the barrier region for malonaldehyde intramolecular proton-transfer to within 0.35 kcal/mol when only trained on reactant/product-like structures. Finally, the use of the Gaussian process variance enables an active learning strategy for extending MOB-ML model to new regions of chemical space with minimal effort. We demonstrate this active learning strategy by extending a QM7b-T model to describe non-covalent interactions in the protein backbone-backbone interaction data set to an accuracy of 0.28 kcal/mol.

physics.chem-ph↗