SearcharxivSearch

arXiv subjects

Evan R. Antoniuk

Publications and source records attributed to Evan R. Antoniuk.

6 recordsLinked to original sources

Active Learning Enables Generation of Molecules that Advance the Known Pareto Front

Although generative models hold promise for discovering molecules with optimized desired properties, they often fail to suggest synthesizable molecules that improve upon the properties of the structures represented in the training distribution. We find that this limitation arises not only from the molecule generation process itself, but also from the poor generalization capabilities of molecular property predictors. We address this challenge by creating a closed-loop molecule generation pipeline with iterative retraining on new quantum chemical simulation data. Compared against static, single-pass generative modeling approaches, only our closed-loop iterative workflow generates molecules with properties extending beyond the training distribution (up to 0.44 standard deviations beyond the original range) and achieves a 79% improvement in out-of-distribution molecule classification accuracy. Furthermore, by conditioning molecular generation on thermodynamic stability data obtained during the iterative loop, the proportion of stable and hence potentially synthesizable molecules generated is 3.5x higher than the next-best model.

cs.LG

BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models

Data-driven molecular discovery leverages artificial intelligence/machine learning (AI/ML) and generative modeling to filter and design novel molecules. Discovering novel molecules requires accurate out-of-distribution (OOD) predictions, but ML models struggle to generalize OOD. Currently, no systematic benchmarks exist for molecular OOD prediction tasks. We present $\mathbf{BOOM}$, $\mathbf{b}$enchmarks for $\mathbf{o}$ut-$\mathbf{o}$f-distribution $\mathbf{m}$olecular property predictions: a chemically-informed benchmark for OOD performance on common molecular property prediction tasks. We evaluate over 150 model-task combinations to benchmark deep learning models on OOD performance. Overall, we find that no existing model achieves strong generalization across all tasks: even the top-performing model exhibited an average OOD error 3x higher than in-distribution. Current chemical foundation models do not show strong OOD extrapolation, while models with high inductive bias can perform well on OOD tasks with simple, specific properties. We perform extensive ablation experiments, highlighting how data generation, pre-training, hyperparameter optimization, model architecture, and molecular representation impact OOD performance. Developing models with strong OOD generalization is a new frontier challenge in chemical ML. This open-source benchmark is available at https://github.com/FLASK-LLNL/BOOM

cs.LG

The Open Polymers 2026 (OPoly26) Dataset and Evaluations

Polymers-macromolecular systems composed of repeating chemical units-constitute the molecular foundation of living organisms, while their synthetic counterparts drive transformative advances across medicine, consumer products, and energy technologies. While machine learning (ML) models have been trained on millions of quantum chemical atomistic simulations for materials and/or small molecular structures to enable efficient, accurate, and transferable predictions of chemical properties, polymers have largely not been included in prior datasets due to the computational expense of high quality electronic structure calculations on representative polymeric structures. Here, we address this shortcoming with the creation of the Open Polymers 2026 (OPoly26) dataset, which contains more than 6.57 million density functional theory (DFT) calculations on up to 360 atom clusters derived from polymeric systems, comprising over 1.2 billion total atoms. OPoly26 captures the chemical diversity that makes polymers intrinsically tunable and versatile materials, encompassing variations in monomer composition, degree of polymerization, chain architectures, and solvation environments. We show that augmenting ML model training with the OPoly26 dataset improves model performance for polymer prediction tasks. We also publicly release the OPoly26 dataset to help further the development of ML models for polymers, and more broadly, strive towards universal atomistic models.

physics.chem-ph

Discovery of stable surfaces with extreme work functions by high-throughput density functional theory and machine learning

The work function is the key surface property that determines how much energy is required for an electron to escape the surface of a material. This property is crucial for thermionic energy conversion, band alignment in heterostructures, and electron emission devices. Here, we present a high-throughput workflow using density functional theory (DFT) to calculate the work function and cleavage energy of 33,631 slabs (58,332 work functions) that we created from 3,716 bulk materials, including up to ternary compounds. The number of materials for which we calculated surface properties surpasses the previously largest database, the Materials Project, by a factor of $\sim$27. On the tail ends of the work function distribution we identify 34 and 56 surfaces with an ultra-low (<2 eV) and ultra-high (>7 eV) work function, respectively. Further, we discover that the $(100)$-Ba-O surface of BaMoO$_3$ and the $(001)$-F surface of Ag$_2$F have record-low (1.25 eV) and record-high (9.06 eV) steady-state work functions without requiring coatings, respectively. Based on this database we develop a physics-based approach to featurize surfaces and use supervised machine learning to predict the work function. We find that physical choice of features improves prediction performance far more than choice of model. Our random forest model achieves a mean absolute test error of 0.09 eV, which is more than 6 times better than the baseline and comparable to the accuracy of DFT. This surrogate model enables rapid predictions of the work function ($\sim 10^5$ faster than DFT) across a vast chemical space and facilitates the discovery of material surfaces with extreme work functions for energy conversion, electronic applications, and contacts in 2-dimensional devices.

cond-mat.mtrl-sci

Representing Polymers as Periodic Graphs with Learned Descriptors for Accurate Polymer Property Predictions

One of the grand challenges of utilizing machine learning for the discovery of innovative new polymers lies in the difficulty of accurately representing the complex structures of polymeric materials. Although a wide array of hand-designed polymer representations have been explored, there has yet to be an ideal solution for how to capture the periodicity of polymer structures, and how to develop polymer descriptors without the need for human feature design. In this work, we tackle these problems through the development of our periodic polymer graph representation. Our pipeline for polymer property predictions is comprised of our polymer graph representation that naturally accounts for the periodicity of polymers, followed by a message-passing neural network (MPNN) that leverages the power of graph deep learning to automatically learn chemically-relevant polymer descriptors. Across a diverse dataset of 10 polymer properties, we find that this polymer graph representation consistently outperforms hand-designed representations with a 20% average reduction in prediction error. Our results illustrate how the incorporation of chemical intuition through directly encoding periodicity into our polymer graph representation leads to a considerable improvement in the accuracy and reliability of polymer property predictions. We also demonstrate how combining polymer graph representations with message-passing neural network architectures can automatically extract meaningful polymer features that are consistent with human intuition, while outperforming human-derived features. This work highlights the advancement in predictive capability that is possible if using chemical descriptors that are specifically optimized for capturing the unique chemical structure of polymers.

cond-mat.mtrl-sci

Machine learning-assisted discovery of many new solid Li-ion conducting materials

We discover many new crystalline solid materials with fast single crystal Li ion conductivity at room temperature, discovered through density functional theory simulations guided by machine learning-based methods. The discovery of new solid Li superionic conductors is of critical importance to the development of safe all-solid-state Li-ion batteries. With a predictive universal structure-property relationship for fast ion conduction not well understood, the search for new solid Li ion conductors has relied largely on trial-and-error computational and experimental searches over the last several decades. In this work, we perform a guided search of materials space with a machine learning (ML)-based prediction model for material selection and density functional theory molecular dynamics (DFT-MD) simulations for calculating ionic conductivity. These materials are screened from over 12,000 experimentally synthesized and characterized candidates with very diverse structures and compositions. When compared to a random search of materials space, we find that the ML-guided search is 2.7 times more likely to identify fast Li ion conductors, with at least a 45x improvement in the log-average of room temperature Li ion conductivity. The F1 score of the ML-based model is 0.50, 3.5 times better than the F1 score expected from completely random guesswork. In a head-to-head competition against six Ph.D. students working in the field, we find that the ML-based model doubles the F1 score of human experts in its ability to identify fast Li-ion conductors from atomistic structure with a thousand- fold increase in speed, clearly demonstrating the utility of this model for the research community. All conducting materials reported here lack transition metals and are predicted to exhibit low electronic conduction, high stability against oxidation, and high thermodynamic stability.

cond-mat.mtrl-sci