SearcharxivSearch

arXiv subjects

Ryan Pederson

Publications and source records attributed to Ryan Pederson.

13 recordsLinked to original sources

Conditional probability density functional theory for solids

A recently developed approach, conditional probability density functional theory (CP-DFT), yields direct access to the exchange-correlation hole of a system, an important correlation function that is not available from any standard DFT calculation. We present the first results for extended materials with periodic boundary conditions. We demonstrate that CP-DFT works on weakly correlated materials (Na, Si). When applied to the prototypical Kagome material $CsV_3Sb_5$, we find $d$-orbital correlations that are not captured by standard DFT. Such distribution leads to a positive finding probability between two separated electrons and an enhanced charge density wave signal, suggesting a useful approach for strongly correlated systems.

cond-mat.mtrl-sci

TerraBind: Fast and Accurate Binding Affinity Prediction through Coarse Structural Representations

We present TerraBind, a foundation model for protein-ligand structure and binding affinity prediction that achieves 26-fold faster inference than state-of-the-art methods while improving affinity prediction accuracy by $\sim$20\%. Current deep learning approaches to structure-based drug design rely on expensive all-atom diffusion to generate 3D coordinates, creating inference bottlenecks that render large-scale compound screening computationally intractable. We challenge this paradigm with a critical hypothesis: full all-atom resolution is unnecessary for accurate small molecule pose and binding affinity prediction. TerraBind tests this hypothesis through a coarse pocket-level representation (protein C$_\beta$ atoms and ligand heavy atoms only) within a multimodal architecture combining COATI-3 molecular encodings and ESM-2 protein embeddings that learns rich structural representations, which are used in a diffusion-free optimization module for pose generation and a binding affinity likelihood prediction module. On structure prediction benchmarks (FoldBench, PoseBusters, Runs N' Poses), TerraBind matches diffusion-based baselines in ligand pose accuracy. Crucially, TerraBind outperforms Boltz-2 by $\sim$20\% in Pearson correlation for binding affinity prediction on both a public benchmark (CASP16) and a diverse proprietary dataset (18 biochemical/cell assays). We show that the affinity prediction module also provides well-calibrated affinity uncertainty estimates, addressing a critical gap in reliable compound prioritization for drug discovery. Furthermore, this module enables a continual learning framework and a hedged batch selection strategy that, in simulated drug discovery cycles, achieves 6$\times$ greater affinity improvement of selected molecules over greedy-based approaches.

cs.LG

Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs

Batched synthesis and testing of molecular designs is the key bottleneck of drug development. There has been great interest in leveraging biomolecular foundation models as surrogates to accelerate this process. In this work, we show how to obtain scalable probabilistic surrogates of binding affinity for use in Batch Bayesian Optimization (Batch BO). This demands parallel acquisition functions that hedge between designs and the ability to rapidly sample from a joint predictive density to approximate them. Through the framework of Epistemic Neural Networks (ENNs), we obtain scalable joint predictive distributions of binding affinity on top of representations taken from large structure-informed models. Key to this work is an investigation into the importance of prior networks in ENNs and how to pretrain them on synthetic data to improve downstream performance in Batch BO. Their utility is demonstrated by rediscovering known potent EGFR inhibitors on a semi-synthetic benchmark in up to 5x fewer iterations, as well as potent inhibitors from a real-world small-molecule library in up to 10x fewer iterations, offering a promising solution for large-scale drug discovery applications.

cs.LG

The difference between molecules and materials: Reassessing the role of exact conditions in density functional theory

Exact conditions have long been used to guide the construction of density functional approximations. But hundreds of empirical-based approximations tailored for chemistry are in use, many of which neglect these conditions in their design. We analyze well-known conditions and revive several obscure ones. Two crucial distinctions are drawn: that between necessary and sufficient conditions, and between all electronic densities and the subset of realistic Coulombic ground states. Simple search algorithms find that many empirical approximations satisfy many exact conditions for realistic densities and non-empirical approximations satisfy even more conditions than those enforced in their construction. The role of exact conditions in developing approximations is revisited.

physics.chem-ph

Large scale quantum chemistry with Tensor Processing Units

We demonstrate the use of Google's cloud-based Tensor Processing Units (TPUs) to accelerate and scale up conventional (cubic-scaling) density functional theory (DFT) calculations. Utilizing 512 TPU cores, we accomplish the largest such DFT computation to date, with 247848 orbitals, corresponding to a cluster of 10327 water molecules with 103270 electrons, all treated explicitly. Our work thus paves the way towards accessible and systematic use of conventional DFT, free of any system-specific constraints, at unprecedented scales.

physics.comp-ph

Attractor Stability in Finite Asynchronous Biological System Models

We present mathematical techniques for exhaustive studies of long-term dynamics of asynchronous biological system models. Specifically, we extend the notion of $κ$-equivalence developed for graph dynamical systems to support systematic analysis of all possible attractor configurations that can be generated when varying the asynchronous update order (Macauley and Mortveit (2009)). We extend earlier work by Veliz-Cuba and Stigler (2011), Goles et al. (2014), and others by comparing long-term dynamics up to topological conjugation: rather than comparing the exact states and their transitions on attractors, we only compare the attractor structures. In general, obtaining this information is computationally intractable. Here, we adapt and apply combinatorial theory for dynamical systems to develop computational methods that greatly reduce this computational cost. We give a detailed algorithm and apply it to ($i$) the lac operon model for Escherichia coli proposed by Veliz-Cuba and Stigler (2011), and ($ii$) the regulatory network involved in the control of the cell cycle and cell differentiation in the Caenorhabditis elegans vulva precursor cells proposed by Weinstein et al. (2015). In both cases, we uncover all possible limit cycle structures for these networks under sequential updates. Specifically, for the lac operon model, rather than examining all $10! > 3.6 \cdot 10^6$ sequential update orders, we demonstrate that it is sufficient to consider $344$ representative update orders, and, more notably, that these $344$ representatives give rise to $4$ distinct attractor structures. A similar analysis performed for the C. elegans model demonstrates that it has precisely $125$ distinct attractor structures. We conclude with observations on the variety and distribution of the models' attractor structures and use the results to discuss their robustness.

cs.DM

Seven Useful Questions in Density Functional Theory

We explore a variety of unsolved problems in density functional theory, where mathematicians might prove useful. We give the background and context of the different problems, and why progress toward resolving them would help those doing computations using density functional theory. Subjects covered include the magnitude of the kinetic energy in Hartree-Fock calculations, the shape of adiabatic connection curves, using the constrained search with input densities, densities of states, the semiclassical expansion of energies, the tightness of Lieb-Oxford bounds, and how we decide the accuracy of an approximate density.

math-ph

Conditional probability density functional theory

We present conditional probability (CP) density functional theory (DFT) as a formally exact theory. In essence, CP-DFT determines the ground-state energy of a system by finding the CP density from a series of independent Kohn-Sham (KS) DFT calculations. By directly calculating CP densities, we bypass the need for an approximate XC energy functional. In this work we discuss and derive several key properties of the CP density and corresponding CP-KS potential. Illustrative examples are used throughout to help guide the reader through the various concepts and theory presented. We explore a suitable CP-DFT approximation and discuss exact conditions, limitations, and results for selected examples.

physics.chem-ph

Machine learning and density functional theory

Over the past decade machine learning has made significant advances in approximating density functionals, but whether this signals the end of human-designed functionals remains to be seen. Ryan Pederson, Bhupalee Kalita and Kieron Burke discuss the rise of machine learning for functional design.

physics.comp-ph

How Well Does Kohn-Sham Regularizer Work for Weakly Correlated Systems?

Kohn-Sham regularizer (KSR) is a differentiable machine learning approach to finding the exchange-correlation functional in Kohn-Sham density functional theory (DFT) that works for strongly correlated systems. Here we test KSR for weak correlation. We propose spin-adapted KSR (sKSR) with trainable local, semilocal, and nonlocal approximations found by minimizing density and total energy loss. We assess the atoms-to-molecules generalizability by training on one-dimensional (1D) H, He, Li, Be, Be$^{++}$ and testing on 1D hydrogen chains, LiH, BeH$_2$, and helium hydride complexes. The generalization error from our semilocal approximation is comparable to other differentiable approaches, but our nonlocal functional outperforms any existing machine learning functionals, predicting ground-state energies of test systems with a mean absolute error of 2.7 milli-Hartrees.

physics.chem-ph

Kohn-Sham equations as regularizer: building prior knowledge into machine-learned physics

Including prior knowledge is important for effective machine learning models in physics, and is usually achieved by explicitly adding loss terms or constraints on model architectures. Prior knowledge embedded in the physics computation itself rarely draws attention. We show that solving the Kohn-Sham equations when training neural networks for the exchange-correlation functional provides an implicit regularization that greatly improves generalization. Two separations suffice for learning the entire one-dimensional H$_2$ dissociation curve within chemical accuracy, including the strongly correlated region. Our models also generalize to unseen types of molecules and overcome self-interaction error.

physics.comp-ph

Bypassing the energy functional in density functional theory: Direct calculation of electronic energies from conditional probability densities

Density functional calculations can fail for want of an accurate exchange-correlation approximation. The energy can instead be extracted from a sequence of density functional calculations of conditional probabilities (CP-DFT). Simple CP approximations yield usefully accurate results for two-electron ions, the hydrogen dimer, and the uniform gas at all temperatures. CP-DFT has no self-interaction error for one electron, and correctly dissociates H2, both major challenges. For warm dense matter, classical CP-DFT calculations can overcome the convergence problems of Kohn-Sham DFT.

physics.chem-ph

Multireference Ab Initio Studies of Magnetic Properties of Terbium-Based Single-Molecule Magnets

We investigate how different chemical environment influences magnetic properties of terbium(III) (Tb)-based single-molecule magnets (SMMs), using first-principles relativistic multireference methods. Recent experiments showed that Tb-based SMMs can have exceptionally large magnetic anisotropy and that they can be used for experimental realization of quantum information applications, with a judicious choice of chemical environment. Here, we perform complete active space self-consistent field (CASSCF) calculations including relativistic spin-orbit interaction (SOI) for representative Tb-based SMMs such as TbPc$_2$ and TbPcNc in three charge states. We calculate low-energy electronic structure from which we compute the Tb crystal-field parameters and construct an effective pseudospin Hamiltonian. Our calculations show that ligand type and fine points of molecular geometry do not affect the zero-field splitting, while the latter varies weakly with oxidation number. On the other hand, higher-energy levels have a strong dependence on all these characteristics. For neutral TbPc$_2$ and TbPcNc molecules, the Tb magnetic moment and the ligand spin are parallel to each other and the coupling strength between them does not depend much on ligand type and details of atomic structure. However, ligand distortion and molecular symmetry play a crucial role in transverse crystal-field parameters which lead to tunnel splitting. The tunnel splitting induces quantum tunneling of magnetization by itself or by combining with other processes. Our results provide insight into mechanisms of magnetization relaxation in the representative Tb-based SMMs.

physics.chem-ph