Searcharxiv⌕ Search

arXiv subjects

Logan Ward

Publications and source records attributed to Logan Ward.

35 records · Page 2Linked to original sources

Rapid Production of Accurate Embedded-Atom Method Potentials for Metal Alloys

A critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the limited availability of suitable interatomic potentials. Here, we introduce the Rapid Alloy Method for Producing Accurate General Empirical Potentials or RAMPAGE, a computationally economical procedure to generate binary embedded-atom model potentials from already-existing single-element potentials that can be further combined into multi-component alloy potentials. We present the quality of RAMPAGE calibrated Finnis-Sinclair type EAM potentials using binary Ag-Al and ternary Ag-Au-Cu as case studies. We demonstrate that RAMPAGE potentials can reproduce bulk properties and forces with greater accuracy than that of other alloy potentials. In some simulations, it is observed the quality of the optimized cross interactions can exceed that of the original off-the-shelf elemental potential inputs.

cond-mat.mtrl-sci↗

Mapping Thermoelectric Transport in a Multicomponent Alloy Space

Interest in high entropy alloy thermoelectric materials is predicated on achieving ultralow lattice thermal conductivity $κ\sub{L}$ through large compositional disorder. However, here we show that for a given mechanism, such as mass contrast phonon scattering, $κ\sub{L}$ will be minimized along the binary alloy with the highest mass contrast, such that adding an intermediate-mass atom to increase atomic disorder can increase thermal conductivity. Only when each component adds an independent scattering mechanism (such as adding strain fluctuation to an existing mass fluctuation) is there a benefit. In addition, both charge carriers and heat-carrying phonons are known to experience scattering due to alloying effects, leading to a trade-off in thermoelectric performance. We apply analytic transport models, based on perturbation and effective medium theories, to predict how alloy scattering will affect the thermal and electronic transport across the full compositional range of several pseudo-ternary and pseudo-quaternary alloy systems. To do so, we demonstrate a multicomponent extension to both thermal and electronic binary alloy scattering models based on the virtual crystal approximation. Finally, we show that common functional forms used in computational thermodynamics can be applied to this problem to further generalize the scattering behavior that is modeled.

cond-mat.mtrl-sci↗

Principles of the Battery Data Genome

Electrochemical energy storage is central to modern society -- from consumer electronics to electrified transportation and the power grid. It is no longer just a convenience but a critical enabler of the transition to a resilient, low-carbon economy. The large pluralistic battery research and development community serving these needs has evolved into diverse specialties spanning materials discovery, battery chemistry, design innovation, scale-up, manufacturing and deployment. Despite the maturity and the impact of battery science and technology, the data and software practices among these disparate groups are far behind the state-of-the-art in other fields (e.g. drug discovery), which have enjoyed significant increases in the rate of innovation. Incremental performance gains and lost research productivity, which are the consequences, retard innovation and societal progress. Examples span every field of battery research , from the slow and iterative nature of materials discovery, to the repeated and time-consuming performance testing of cells and the mitigation of degradation and failures. The fundamental issue is that modern data science methods require large amounts of data and the battery community lacks the requisite scalable, standardized data hubs required for immediate use of these approaches. Lack of uniform data practices is a central barrier to the scale problem. In this perspective we identify the data- and software-sharing gaps and propose the unifying principles and tools needed to build a robust community of data hubs, which provide flexible sharing formats to address diverse needs. The Battery Data Genome is offered as a data-centric initiative that will enable the transformative acceleration of battery science and technology, and will ultimately serve as a catalyst to revolutionize our approach to innovation.

physics.soc-ph↗

Colmena: Scalable Machine-Learning-Based Steering of Ensemble Simulations for High Performance Computing

Scientific applications that involve simulation ensembles can be accelerated greatly by using experiment design methods to select the best simulations to perform. Methods that use machine learning (ML) to create proxy models of simulations show particular promise for guiding ensembles but are challenging to deploy because of the need to coordinate dynamic mixes of simulation and learning tasks. We present Colmena, an open-source Python framework that allows users to steer campaigns by providing just the implementations of individual tasks plus the logic used to choose which tasks to execute when. Colmena handles task dispatch, results collation, ML model invocation, and ML model (re)training, using Parsl to execute tasks on HPC systems. We describe the design of Colmena and illustrate its capabilities by applying it to electrolyte design, where it both scales to 65536 CPUs and accelerates the discovery rate for high-performance molecules by a factor of 100 over unguided searches.

cs.DC↗

Evening the Score: Targeting SARS-CoV-2 Protease Inhibition in Graph Generative Models for Therapeutic Candidates

We examine a pair of graph generative models for the therapeutic design of novel drug candidates targeting SARS-CoV-2 viral proteins. Due to a sense of urgency, we chose well-validated models with unique strengths: an autoencoder that generates molecules with similar structures to a dataset of drugs with anti-SARS activity and a reinforcement learning algorithm that generates highly novel molecules. During generation, we explore optimization toward several design targets to balance druglikeness, synthetic accessability, and anti-SARS activity based on \icfifty. This generative framework\footnote{https://github.com/exalearn/covid-drug-design} will accelerate drug discovery in future pandemics through the high-throughput generation of targeted therapeutic candidates.

q-bio.BM↗

Benchmarking Deep Graph Generative Models for Optimizing New Drug Molecules for COVID-19

Design of new drug compounds with target properties is a key area of research in generative modeling. We present a small drug molecule design pipeline based on graph-generative models and a comparison study of two state-of-the-art graph generative models for designing COVID-19 targeted drug candidates: 1) a variational autoencoder-based approach (VAE) that uses prior knowledge of molecules that have been shown to be effective for earlier coronavirus treatments and 2) a deep Q-learning method (DQN) that generates optimized molecules without any proximity constraints. We evaluate the novelty of the automated molecule generation approaches by validating the candidate molecules with drug-protein binding affinity models. The VAE method produced two novel molecules with similar structures to the antiretroviral protease inhibitor Indinavir that show potential binding affinity for the SARS-CoV-2 protein target 3-chymotrypsin-like protease (3CL-protease).

cs.LG↗

AI- and HPC-enabled Lead Generation for SARS-CoV-2: Models and Processes to Extract Druglike Molecules Contained in Natural Language Text

Researchers worldwide are seeking to repurpose existing drugs or discover new drugs to counter the disease caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). A promising source of candidates for such studies is molecules that have been reported in the scientific literature to be drug-like in the context of coronavirus research. We report here on a project that leverages both human and artificial intelligence to detect references to drug-like molecules in free text. We engage non-expert humans to create a corpus of labeled text, use this labeled corpus to train a named entity recognition model, and employ the trained model to extract 10912 drug-like molecules from the COVID-19 Open Research Dataset Challenge (CORD-19) corpus of 198875 papers. Performance analyses show that our automated extraction model can achieve performance on par with that of non-expert humans.

cs.CL↗

HydroNet: Benchmark Tasks for Preserving Intermolecular Interactions and Structural Motifs in Predictive and Generative Models for Molecular Data

Intermolecular and long-range interactions are central to phenomena as diverse as gene regulation, topological states of quantum materials, electrolyte transport in batteries, and the universal solvation properties of water. We present a set of challenge problems for preserving intermolecular interactions and structural motifs in machine-learning approaches to chemical problems, through the use of a recently published dataset of 4.95 million water clusters held together by hydrogen bonding interactions and resulting in longer range structural patterns. The dataset provides spatial coordinates as well as two types of graph representations, to accommodate a variety of machine-learning practices.

cs.LG↗

A high-throughput structural and electrochemical study of metallic glass formation in Ni-Ti-Al

Based on a set of machine learning predictions of glass formation in the Ni-Ti-Al system, we have undertaken a high-throughput experimental study of that system. We utilized rapid synthesis followed by high-throughput structural and electrochemical characterization. Using this dual-modality approach, we are able to better classify the amorphous portion of the library, which we found to be the portion with a full-width-half-maximum (FWHM) of 0.42 A$^{-1}$ for the first sharp x-ray diffraction peak. We demonstrate that the FWHM and corrosion resistance are correlated but that, while chemistry still plays a role, a large FWHM is necessary for the best corrosion resistance.

cond-mat.mtrl-sci↗

A Data Ecosystem to Support Machine Learning in Materials Science

Facilitating the application of machine learning to materials science problems will require enhancing the data ecosystem to enable discovery and collection of data from many sources, automated dissemination of new data across the ecosystem, and the connecting of data with materials-specific machine learning models. Here, we present two projects, the Materials Data Facility (MDF) and the Data and Learning Hub for Science (DLHub), that address these needs. We use examples to show how MDF and DLHub capabilities can be leveraged to link data with machine learning models and how users can access those capabilities through web and programmatic interfaces.

cond-mat.mtrl-sci↗

IRNet: A General Purpose Deep Residual Regression Framework for Materials Discovery

Materials discovery is crucial for making scientific advances in many domains. Collections of data from experiments and first-principle computations have spurred interest in applying machine learning methods to create predictive models capable of mapping from composition and crystal structures to materials properties. Generally, these are regression problems with the input being a 1D vector composed of numerical attributes representing the material composition and/or crystal structure. While neural networks consisting of fully connected layers have been applied to such problems, their performance often suffers from the vanishing gradient problem when network depth is increased. In this paper, we study and propose design principles for building deep regression networks composed of fully connected layers with numerical vectors as input. We introduce a novel deep regression network with individual residual learning, IRNet, that places shortcut connections after each layer so that each layer learns the residual mapping between its output and input. We use the problem of learning properties of inorganic materials from numerical attributes derived from material composition and/or crystal structure to compare IRNet's performance against that of other machine learning techniques. Using multiple datasets from the Open Quantum Materials Database (OQMD) and Materials Project for training and evaluation, we show that IRNet provides significantly better prediction performance than the state-of-the-art machine learning approaches currently used by domain scientists. We also show that IRNet's use of individual residual learning leads to better convergence during the training phase than when shortcut connections are between multi-layer stacks while maintaining the same number of parameters.

physics.comp-ph↗

Machine Learning Prediction of Accurate Atomization Energies of Organic Molecules from Low-Fidelity Quantum Chemical Calculations

Recent studies illustrate how machine learning (ML) can be used to bypass a core challenge of molecular modeling: the tradeoff between accuracy and computational cost. Here, we assess multiple ML approaches for predicting the atomization energy of organic molecules. Our resulting models learn the difference between low-fidelity, B3LYP, and high-accuracy, G4MP2, atomization energies, and predict the G4MP2 atomization energy to 0.005 eV (mean absolute error) for molecules with less than 9 heavy atoms and 0.012 eV for a small set of molecules with between 10 and 14 heavy atoms. Our two best models, which have different accuracy/speed tradeoffs, enable the efficient prediction of G4MP2-level energies for large molecules and are available through a simple web interface.

physics.comp-ph↗

Ternary mixed-anion semiconductors with tunable band gaps from machine-learning and crystal structure prediction

We report the computational investigation of a series of ternary X$_4$Y$_2$Z and X$_5$Y$_2$Z$_2$ compounds with X={Mg, Ca, Sr, Ba}, Y={P, As, Sb, Bi}, and Z={S, Se, Te}. The compositions for these materials were predicted through a search guided by machine learning, while the structures were resolved using the minima hopping crystal structure prediction method. Based on $\textit{ab initio}$ calculations, we predict that many of these compounds are thermodynamically stable. In particular, 21 of the X$_4$Y$_2$Z compounds crystallize in a tetragonal structure with $\textit{I-42d}$ symmetry, and exhibit band gaps in the range of 0.3 and 1.8 eV, well suited for various energy applications. We show that several candidate compounds (in particular X$_4$Y$_2$Te and X$_4$Sb$_2$Se) exhibit good photo absorption in the visible range, while others (e.g., Ba$_4$Sb$_2$Se) show excellent thermoelectric performance due to a high power factor and extremely low lattice thermal conductivities.

cond-mat.mtrl-sci↗

DLHub: Model and Data Serving for Science

While the Machine Learning (ML) landscape is evolving rapidly, there has been a relative lag in the development of the "learning systems" needed to enable broad adoption. Furthermore, few such systems are designed to support the specialized requirements of scientific ML. Here we present the Data and Learning Hub for science (DLHub), a multi-tenant system that provides both model repository and serving capabilities with a focus on science applications. DLHub addresses two significant shortcomings in current systems. First, its selfservice model repository allows users to share, publish, verify, reproduce, and reuse models, and addresses concerns related to model reproducibility by packaging and distributing models and all constituent components. Second, it implements scalable and low-latency serving capabilities that can leverage parallel and distributed computing resources to democratize access to published models through a simple web interface. Unlike other model serving frameworks, DLHub can store and serve any Python 3-compatible model or processing function, plus multiple-function pipelines. We show that relative to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, DLHub provides greater capabilities, comparable performance without memoization and batching, and significantly better performance when the latter two techniques can be employed. We also describe early uses of DLHub for scientific applications.

cs.LG↗

A General-Purpose Machine Learning Framework for Predicting Properties of Inorganic Materials

A very active area of materials research is to devise methods that use machine learning to automatically extract predictive models from existing materials data. While prior examples have demonstrated successful models for some applications, many more applications exist where machine learning can make a strong impact. To enable faster development of machine-learning-based models for such applications, we have created a framework capable of being applied to a broad range of materials data. Our method works by using a chemically diverse list of attributes, which we demonstrate are suitable for describing a wide variety of properties, and a novel method for partitioning the data set into groups of similar materials in order to boost the predictive accuracy. In this manuscript, we demonstrate how this new method can be used to predict diverse properties of crystalline and amorphous materials, such as band gap energy and glass-forming ability.

cond-mat.mtrl-sci↗

Structural evolution and kinetics in Cu-Zr Metallic Liquids

The atomic structure of the supercooled liquid has often been discussed as a key source of glass formation in metals. The presence of icosahedrally-coordinated clusters and their tendency to form networks have been identified as one possible structural trait leading to glass forming ability in the Cu-Zr binary system. In this work, we show that this theory is insufficient to explain glass formation at all compositions in that binary system. Instead, we propose that the formation of ideally-packed clusters at the expense of atomic arrangements with excess or deficient free volume can explain glass-forming by a similar mechanism. We show that this behavior is reflected in the structural relaxation of a metallic glass during constant pressure cooling and the time evolution of structure at a constant volume. We then demonstrate that this theory is sufficient to explain slowed diffusivity in compositions across the range of Cu-Zr metallic glasses.

cond-mat.mtrl-sci↗

Rapid Production of Accurate Embedded-Atom Method Potentials for Metal Alloys

The most critical limitation to the wide-scale use of classical molecular dynamics for alloy design is the availability of suitable interatomic potentials. In this work, we demonstrate a simple procedure to generate a library of accurate binary potentials using already-existing single-element potentials that can be easily combined to form multi-component alloy potentials. For the Al-Ni, Cu-Au, and Cu-Al-Zr systems, we show that this method produces results comparable in accuracy to alloy potentials where all parts have been fitted simultaneously, without the additional computational expense. Furthermore, we demonstrate applicability to both crystalline and amorphous phases.

cond-mat.mtrl-sci↗