SearcharxivSearch

arXiv subjects

Harikrishna Sahu

Publications and source records attributed to Harikrishna Sahu.

5 recordsLinked to original sources

A Synthetically-accessible Universe of Chemically Recyclable Polymers

Polymers synthesized via ring-opening polymerization (ROP) of cyclic monomers represent an important class of materials due to their chemical recyclability and possible insertion in several critical applications. We present a dataset of 1 million synthetically realizable ROP polymer structures generated through a combination of Virtual Forward Synthesis (VFS) and polymer expert language models and qualified by stringent chemical heuristics. VFS is used to generate ROP polymers by applying known reactions to existing monomers. The polymer foundation models polyBART and POLYT5 further enable the generation of ROP candidates, with polyBART exploring its learned latent space and POLYT5 producing candidates via sequence-to-sequence generation. The resulting ROP polymers are subjected to robust filtering criteria to ensure novelty, validity and overall data quality through a combination of automated validation pipelines and a comprehensive set of chemist-informed heuristic rules introduced in this work for the first time. We hope that this dataset will serve as a valuable resource for downstream sustainable applications.

cond-mat.soft

Load-dependent Hardness Prediction for Materials using Machine Learning

Superhard materials are critical for wear-resistant and high-stress applications. Conventional approaches correlating hardness with elastic moduli derived from DFT calculations enable rapid screening but overlook the strong load dependence of hardness. In this work, machine learning (ML) models were developed using a large, curated dataset of load-dependent experimental Vickers hardness (Hv) measurements. Moderate correlation was observed between experimental and DFT-based Hv values, whereas a single-task ML model trained solely on experimental data outperformed multi-task models that combined experimental and computed data. The superior performance of the single-task model highlights that explicit inclusion of indentation load, along with compositional, electronic, and structural descriptors, is essential and sufficient for accurate hardness prediction, beyond what can be achieved using DFT-accessible bulk and shear moduli alone (or in tandem with experimental data). These results emphasize the importance of high-quality experimental data and explicit inclusion of measurement conditions, particularly load, in the development of reliable hardness prediction models.

cond-mat.mtrl-sci

On-the-Fly Machine-Learned Force Fields for High-Fidelity Polymer Glass Transition Simulations

Predicting polymer glass transition temperatures (Tg) with first-principles fidelity has long remained out of reach, as cooling multi-thousand-atom systems over a broad temperature range at acceptable rates exceeds the computational limits of ab initio molecular dynamics (AIMD). Here we employ a hybrid scheme that merges AIMD with accelerated on-the-fly (OTF) machine-learned force-field (MLFF) construction, enabling Tg prediction at quantum-mechanical accuracy with near-classical computational cost. The OTF protocol to construct MLFFs adaptively triggers first-principles calculations only when newly encountered configurations lie outside the current model's domain of confidence, allowing robust, parameter-free MLFFs to be built from merely 1000 AIMD-sampled configurations per polymer. These MLFFs are then utilized to perform long-time cooling simulations on amorphous supercells containing several thousand atoms. Applied across twelve polymers spanning aromatic, aliphatic, heteroatomic, and branched chemistries, the method yields predictions in excellent accord with experiment while reducing computational cost by approximately six orders of magnitude relative to AIMD. This work establishes a new paradigm for predictive polymer modeling, demonstrating that OTF-MLFFs provide a generalizable, accurate, and scalable route to simulating the thermophysical behavior of complex disordered materials at near quantum-mechanical fidelity.

cond-mat.mtrl-sci

An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design

Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.

cond-mat.mtrl-sci

polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Design

Designing polymers for targeted applications and accurately predicting their properties is a key challenge in materials science owing to the vast and complex polymer chemical space. While molecular language models have proven effective in solving analogous problems for molecular discovery, similar advancements for polymers are limited. To address this gap, we propose polyBART, a language model-driven polymer discovery capability that enables rapid and accurate exploration of the polymer design space. Central to our approach is Pseudo-polymer SELFIES (PSELFIES), a novel representation that allows for the transfer of molecular language models to the polymer space. polyBART is, to the best of our knowledge, the first language model capable of bidirectional translation between polymer structures and properties, achieving state-of-the-art results in property prediction and design of novel polymers for electrostatic energy storage. Further, polyBART is validated through a combination of both computational and laboratory experiments. We report what we believe is the first successful synthesis and validation of a polymer designed by a language model, predicted to exhibit high thermal degradation temperature and confirmed by our laboratory measurements. Our work presents a generalizable strategy for adapting molecular language models to the polymer space and introduces a polymer foundation model, advancing generative polymer design that may be adapted for a variety of applications.

cond-mat.soft