SearcharxivSearch

arXiv subjects

Nathan A. Baker

Publications and source records attributed to Nathan A. Baker.

At least 19 recordsLinked to original sources

Clifford-efficient sparse state preparation for molecular wavefunctions

Sparse quantum state preparation concerns an $n$-qubit target state that is a superposition of only $d \ll 2^n$ computational basis states. Existing approaches exploit this sparsity by compressing these $d$ basis states and their amplitudes onto a smaller set of qubits, called the dense register, before expanding the prepared state to the full register. Rather than relying on the permutation-based compression used in prior work, we exploit affine relationships among the binary configurations over the finite field $\operatorname{GF}(2)$ to reduce both the non-Clifford gate count and the ancillary qubit count. Invertible affine transformations over $\operatorname{GF}(2)$, comprising Gaussian elimination and all-ones-row removal, first reduce the dense register from $n$ to the rank $r$ using only Clifford gates and no ancillary qubits. An optional binary encoding stage then trades additional Toffoli gates and ancillary qubits for further compression to the minimum $\lceil\log_2 d\rceil$ dense qubits needed to represent $d$ distinct configurations. For chemically relevant wavefunctions, such as those obtained from selected configuration interaction calculations, shared electronic excitation patterns produce many of these affine relationships, enabling substantial Clifford-only compression before binary encoding. Across the molecular benchmarks, our method requires the fewest ancillary qubits among the evaluated sparse state preparation methods while maintaining comparable non-Clifford gate counts when using binary encoding.

quant-ph

QDK/Chemistry: A Modular Toolkit for Quantum Chemistry Applications

We present QDK/Chemistry, a software toolkit for quantum chemistry workflows targeting quantum computers. The toolkit addresses a key challenge in the field: while quantum algorithms for chemistry have matured considerably, the infrastructure connecting classical electronic structure calculations to quantum circuit execution remains fragmented. QDK/Chemistry provides this infrastructure through a modular architecture that separates data representations from computational methods, enabling researchers to compose workflows from interchangeable components. In addition to providing native implementations of targeted algorithms in the quantum-classical pipeline, the toolkit builds upon and integrates with widely used open-source quantum chemistry packages and quantum computing frameworks through a plugin system, allowing users to combine methods from different sources without modifying workflow logic. This paper describes the design philosophy, current capabilities, and role of QDK/Chemistry as a foundation for reproducible quantum chemistry experiments.

quant-ph

Acceleration without Disruption: DFT Software as a Service

Density functional theory (DFT) has been a cornerstone in computational chemistry, physics, and materials science for decades, benefiting from advancements in computational power and theoretical methods. This paper introduces a novel, cloud-native application, Accelerated DFT, which offers an order of magnitude acceleration in DFT simulations. By integrating state-of-the-art cloud infrastructure and redesigning algorithms for graphic processing units (GPUs), Accelerated DFT achieves high-speed calculations without sacrificing accuracy. It provides an accessible and scalable solution for the increasing demands of DFT calculations in scientific communities. The implementation details, examples, and benchmark results illustrate how Accelerated DFT can significantly expedite scientific discovery across various domains.

physics.chem-ph

Accelerating computational materials discovery with artificial intelligence and cloud high-performance computing: from large-scale screening to experimental validation

High-throughput computational materials discovery has promised significant acceleration of the design and discovery of new materials for many years. Despite a surge in interest and activity, the constraints imposed by large-scale computational resources present a significant bottleneck. Furthermore, examples of large-scale computational discovery carried through experimental validation remain scarce, especially for materials with product applicability. Here we demonstrate how this vision became reality by first combining state-of-the-art artificial intelligence (AI) models and traditional physics-based models on cloud high-performance computing (HPC) resources to quickly navigate through more than 32 million candidates and predict around half a million potentially stable materials. By focusing on solid-state electrolytes for battery applications, our discovery pipeline further identified 18 promising candidates with new compositions and rediscovered a decade's worth of collective knowledge in the field as a byproduct. By employing around one thousand virtual machines (VMs) in the cloud, this process took less than 80 hours. We then synthesized and experimentally characterized the structures and conductivities of our top candidates, the Na$_x$Li$_{3-x}$YCl$_6$ ($0 < x < 3$) series, demonstrating the potential of these compounds to serve as solid electrolytes. Additional candidate materials that are currently under experimental investigation could offer more examples of the computational discovery of new phases of Li- and Na-conducting solid electrolytes. We believe that this unprecedented approach of synergistically integrating AI models and cloud HPC not only accelerates materials discovery but also showcases the potency of AI-guided experimentation in unlocking transformative scientific breakthroughs with real-world applications.

cond-mat.mtrl-sci

A clustering-based biased Monte Carlo approach to protein titration curve prediction

In this work, we developed an efficient approach to compute ensemble averages in systems with pairwise-additive energetic interactions between the entities. Methods involving full enumeration of the configuration space result in exponential complexity. Sampling methods such as Markov Chain Monte Carlo (MCMC) algorithms have been proposed to tackle the exponential complexity of these problems; however, in certain scenarios where significant energetic coupling exists between the entities, the efficiency of the such algorithms can be diminished. We used a strategy to improve the efficiency of MCMC by taking advantage of the cluster structure in the interaction energy matrix to bias the sampling. We pursued two different schemes for the biased MCMC runs and show that they are valid MCMC schemes. We used both synthesized and real-world systems to show the improved performance of our biased MCMC methods when compared to the regular MCMC method. In particular, we applied these algorithms to the problem of estimating protonation ensemble averages and titration curves of residues in a protein.

q-bio.BM

Towards quantum computing for high-energy excited states in molecular systems: quantum phase estimations of core-level states

This paper explores the utility of the quantum phase estimation (QPE) in calculating high-energy excited states characterized by promotions of electrons occupying inner energy shells. These states have been intensively studied over the last few decades especially in supporting the experimental effort at light sources. Results obtained with the QPE are compared with various high-accuracy many-body techniques developed to describe core-level states. The feasibility of the quantum phase estimator in identifying classes of challenging shake-up states characterized by the presence of higher-order excitation effects is also discussed.

quant-ph

Data-driven molecular modeling with the generalized Langevin equation

The complexity of molecular dynamics simulations necessitates dimension reduction and coarse-graining techniques to enable tractable computation. The generalized Langevin equation (GLE) describes coarse-grained dynamics in reduced dimensions. In spite of playing a crucial role in non-equilibrium dynamics, the memory kernel of the GLE is often ignored because it is difficult to characterize and expensive to solve. To address these issues, we construct a data-driven rational approximation to the GLE. Building upon previous work leveraging the GLE to simulate simple systems, we extend these results to more complex molecules, whose many degrees of freedom and complicated dynamics require approximation methods. We demonstrate the effectiveness of our approximation by testing it against exact methods and comparing observables such as autocorrelation and transition rates.

physics.comp-ph

Visualizing biomolecular electrostatics in virtual reality with UnityMol-APBS

Virtual reality is a powerful tool with the ability to immerse a user within a completely external environment. This immersion is particularly useful when visualizing and analyzing interactions between small organic molecules, molecular inorganic complexes, and biomolecular systems such as redox proteins and enzymes. A common tool used in the biomedical community to analyze such interactions is the APBS software, which was developed to solve the equations of continuum electrostatics for large biomolecular assemblages. Numerous applications exist for using APBS in the biomedical community including analysis of protein ligand interactions and APBS has enjoyed widespread adoption throughout the biomedical community. Currently, typical use of the full APBS toolset is completed via the command line followed by visualization using a variety of two-dimensional external molecular visualization software. This process has inherent limitations: visualization of three-dimensional objects using a two-dimensional interface masks important information within the depth component. Herein, we have developed a single application, UnityMol-APBS, that provides a dual experience where users can utilize the full range of the APBS toolset, without the use of a command line interface, by use of a simple \ac{GUI} for either a standard desktop or immersive virtual reality experience.

q-bio.BM

Q# and NWChem: Tools for Scalable Quantum Chemistry on Quantum Computers

Fault-tolerant quantum computation promises to solve outstanding problems in quantum chemistry within the next decade. Realizing this promise requires scalable tools that allow users to translate descriptions of electronic structure problems to optimized quantum gate sequences executed on physical hardware, without requiring specialized quantum computing knowledge. To this end, we present a quantum chemistry library, under the open-source MIT license, that implements and enables straightforward use of state-of-art quantum simulation algorithms. The library is implemented in Q#, a language designed to express quantum algorithms at scale, and interfaces with NWChem, a leading electronic structure package. We define a standardized schema for this interface, Broombridge, that describes second-quantized Hamiltonians, along with metadata required for effective quantum simulation, such as trial wavefunction ansatzes. This schema is generated for arbitrary molecules by NWChem, conveniently accessible, for instance, through Docker containers and a recently developed web interface EMSL Arrows. We illustrate use of the library with various examples, including ground- and excited-state calculations for LiH, H$_{10}$, and C$_{20}$ with an active-space simplification, and automatically obtain resource estimates for classically intractable examples.

quant-ph

Atomic radius and charge parameter uncertainty in biomolecular solvation energy calculations

Atomic radii and charges are two major parameters used in implicit solvent electrostatics and energy calculations. The optimization problem for charges and radii is under-determined, leading to uncertainty in the values of these parameters and in the results of solvation energy calculations using these parameters. This paper presents a new method for quantifying this uncertainty in implicit solvation calculations of small molecules using surrogate models based on generalized polynomial chaos (gPC) expansions. There are relatively few atom types used to specify radii parameters in implicit solvation calculations; therefore, surrogate models for these low-dimensional spaces could be constructed using least-squares fitting. However, there are many more types of atomic charges; therefore, construction of surrogate models for the charge parameter space requires compressed sensing combined with an iterative rotation method to enhance problem sparsity. We demonstrate the application of the method by presenting results for the uncertainties in small molecule solvation energies based on these approaches. The method presented in this paper is a promising approach for efficiently quantifying uncertainty in a wide range of force field parameterization problems, including those beyond continuum solvation calculations.The intent of this study is to provide a way for developers of implicit solvent model parameter sets to understand the sensitivity of their target properties (solvation energy) on underlying choices for solute radius and charge parameters.

q-bio.BM

Improvements to the APBS biomolecular solvation software suite

The Adaptive Poisson-Boltzmann Solver (APBS) software was developed to solve the equations of continuum electrostatics for large biomolecular assemblages that has provided impact in the study of a broad range of chemical, biological, and biomedical applications. APBS addresses three key technology challenges for understanding solvation and electrostatics in biomedical applications: accurate and efficient models for biomolecular solvation and electrostatics, robust and scalable software for applying those theories to biomolecular systems, and mechanisms for sharing and analyzing biomolecular electrostatics data in the scientific community. To address new research applications and advancing computational capabilities, we have continually updated APBS and its suite of accompanying software since its release in 2001. In this manuscript, we discuss the models and capabilities that have recently been implemented within the APBS software package including: a Poisson-Boltzmann analytical and a semi-analytical solver, an optimized boundary element solver, a geometry-based geometric flow solvation model, a graph theory based algorithm for determining p$K_a$ values, and an improved web-based visualization tool for viewing electrostatics.

q-bio.BM

GIBS: A grand-canonical Monte Carlo simulation program for simulating ion-biomolecule interactions

The ionic environment of biomolecules strongly influences their structure, conformational stability, and inter-molecular interactions.This paper introduces GIBS, a grand-canonical Monte Carlo (GCMC) simulation program for computing the thermodynamic properties of ion solutions and their distributions around biomolecules. This software implements algorithms that automate the excess chemical potential calculations for a given target salt concentration. GIBS uses a cavity-bias algorithm to achieve high sampling acceptance rates for inserting ions and solvent hard spheres in simulating dense ionic systems. In the current version, ion-ion interactions are described using Coulomb, hard-sphere, or Lennard-Jones (L-J) potentials; solvent-ion interactions are described using hard-sphere, L-J and attractive square-well potentials; and, solvent-solvent interactions are described using hard-sphere repulsions. This paper and the software package includes examples of using GIBS to compute the ion excess chemical potentials and mean activity coefficients of sodium chloride as well as to compute the cylindrical radial distribution functions of monovalent (Na$^+$, Rb$^+$), divalent (Sr$^{2+}$), and trivalent (CoHex$^{3+}$) around fixed all-atom models of 25 base-pair nucleic acid duplexes. GIBS is written in C++ and is freely available community use; it can be downloaded at https://github.com/Electrostatics/GIBS.

q-bio.BM

Bayesian Model Averaging for Ensemble-Based Estimates of Solvation Free Energies

This paper applies the Bayesian Model Averaging (BMA) statistical ensemble technique to estimate small molecule solvation free energies. There is a wide range of methods available for predicting solvation free energies, ranging from empirical statistical models to ab initio quantum mechanical approaches. Each of these methods is based on a set of conceptual assumptions that can affect predictive accuracy and transferability. Using an iterative statistical process, we have selected and combined solvation energy estimates using an ensemble of 17 diverse methods from the fourth Statistical Assessment of Modeling of Proteins and Ligands (SAMPL) blind prediction study to form a single, aggregated solvation energy estimate. The ensemble design process evaluates the statistical information in each individual method as well as the performance of the aggregate estimate obtained from the ensemble as a whole. Methods that possess minimal or redundant information are pruned from the ensemble and the evaluation process repeats until aggregate predictive performance can no longer be improved. We show that this process results in a final aggregate estimate that outperforms all individual methods by reducing estimate errors by as much as 91% to 1.2 kcal/mol accuracy. We also compare our iterative refinement approach to other statistical ensemble approaches and demonstrate that this iterative process reduces estimate errors by as much as 61%. This work provides a new approach for accurate solvation free energy prediction and lays the foundation for future work on aggregate models that can balance computational cost with prediction accuracy.

q-bio.BM

Energy Minimization of Discrete Protein Titration State Models Using Graph Theory

There are several applications in computational biophysics which require the optimization of discrete interacting states; e.g., amino acid titration states, ligand oxidation states, or discrete rotamer angles. Such optimization can be very time-consuming as it scales exponentially in the number of sites to be optimized. In this paper, we describe a new polynomial-time algorithm for optimization of discrete states in macromolecular systems. This algorithm was adapted from image processing and uses techniques from discrete mathematics and graph theory to restate the optimization problem in terms of "maximum flow-minimum cut" graph analysis. The interaction energy graph, a graph in which vertices (amino acids) and edges (interactions) are weighted with their respective energies, is transformed into a flow network in which the value of the minimum cut in the network equals the minimum free energy of the protein, and the cut itself encodes the state that achieves the minimum free energy. Because of its deterministic nature and polynomial-time performance, this algorithm has the potential to allow for the ionization state of larger proteins to be discovered.

q-bio.BM

Multi-shell model of ion-induced nucleic acid condensation

We present a semi-quantitative model of condensation of short nucleic acid (NA) duplexes induced by tri-valent cobalt(III) hexammine (CoHex) ions. The model is based on partitioning of bound counterion distribution around singleNA duplex into "external" and "internal" ion binding shells distinguished by the proximity to duplex helical axis. In the aggregated phase the shells overlap, which leads to significantly increased attraction of CoHex ions in these overlaps with the neighboring duplexes. The duplex aggregation free energy is decomposed into attractive and repulsive components in such a way that they can be represented by simple analytical expressions with parameters derived from molecular dynamic (MD) simulations and numerical solutions of Poisson equation. The short-range interactions described by the attractive term depend on the fractions of bound ions in the overlapping shells and affinity of CoHex to the "external" shell of nearly neutralized duplex. The repulsive components of the free energy are duplex configurational entropy loss upon the aggregation and the electrostatic repulsion of the duplexes that remains after neutralization by bound CoHex ions. The estimates of the aggregation free energy are consistent with the experimental range of NA duplex condensation propensities, including the unusually poor condensation of RNA structures and subtle sequence effects upon DNA condensation. The model predicts that, in contrast to DNA, RNA duplexes may condense into tighter packed aggregates with a higher degree of duplex neutralization. The model also predicts that longer NA fragments will condense more readily than shorter ones. The ability of this model to explain experimentally observed trends in NA condensation, lends support to proposed NA condensation picture based on the multivalent "ion binding shells".

q-bio.BM

Continuum Electrostatics Approaches to Calculating p$K_a$s and $E_m$s in Proteins

Proteins change their charge state through protonation and redox reactions as well as through binding charged ligands. The free energy of these reactions are dominated by solvation and electrostatic energies and modulated by protein conformational relaxation in response to the ionization state changes. Although computational methods for calculating these interactions can provide very powerful tools for predicting protein charge states, they include several critical approximations of which users should be aware. This chapter discusses the strengths, weaknesses, and approximations of popular computational methods for predicting charge states and understanding their underlying electrostatic interactions. The goal of this chapter is to inform users about applications and potential caveats of these methods as well as outline directions for future theoretical and computational research.

q-bio.BM

An ISA-Tab specification for protein titration data exchange

Data curation presents a challenge to all scientific disciplines to ensure public availability and reproducibility of experimental data. Standards for data preservation and exchange are central to addressing this challenge: the Investigation-Study-Assay Tabular (ISA-Tab) project has developed a widely used template for such standards in biological research. This paper describes the application of ISA-Tab to protein titration data. Despite the importance of titration experiments for understanding protein structure, stability, and function and for testing computational approaches to protein electrostatics, no such mechanism currently exists for sharing and preserving biomolecular titration data. We have adapted the ISA-Tab template to provide a structured means of supporting experimental structural chemistry data with a particular emphasis on the calculation and measurement of pKa values. This activity has been performed as part of the broader pKa Cooperative effort, leveraging data that has been collected and curated by the Cooperative members. In this article, we present the details of this specification and its application to a broad range of pKa and electrostatics data obtained for multiple protein systems. The resulting curated data is publicly available at http://pkacoop.org.

q-bio.BM

The role of correlation and solvation in ion interactions with B-DNA

The ionic atmospheres around nucleic acids play important roles in biological function. Large-scale explicit solvent simulations coupled to experimental assays such as anomalous small-angle X-ray scattering (ASAXS) can provide important insights into the structure and energetics of such atmospheres but are time- and resource-intensive. In this paper, we use classical density functional theory (cDFT) to explore the balance between ion-DNA, ion-water, and ion-ion interactions in ionic atmospheres of RbCl, SrCl$_2$, and CoHexCl$_3$ (cobalt hexammine chloride) around a B-form DNA molecule. The accuracy of the cDFT calculations was assessed by comparison between simulated and experimental ASAXS curves, demonstrating that an accurate model should take into account ion-ion correlation and ion hydration forces, DNA topology, and the discrete distribution of charges on DNA strands. As expected, these calculations revealed significant differences between monovalent, divalent, and trivalent cation distributions around DNA. About half of the DNA-bound Rb$^+$ ions penetrate into the minor groove of the DNA and half adsorb on the DNA strands. The fraction of cations in the minor groove decreases for the larger Sr$^{2+}$ ions and becomes zero for CoHex$^{3+}$ ions, which all adsorb on the DNA strands. The distribution of CoHex$^{3+}$ ions is mainly determined by Coulomb and steric interactions, while ion-correlation forces play a central role in the monovalent Rb$^+$ distribution and a combination of ion-correlation and hydration forces affect the Sr$^{2+}$ distribution around DNA.

q-bio.BM