SearcharxivSearch

arXiv subjects

Rampi Ramprasad

Publications and source records attributed to Rampi Ramprasad.

At least 19 recordsLinked to original sources

A Synthetically-accessible Universe of Chemically Recyclable Polymers

Polymers synthesized via ring-opening polymerization (ROP) of cyclic monomers represent an important class of materials due to their chemical recyclability and possible insertion in several critical applications. We present a dataset of 1 million synthetically realizable ROP polymer structures generated through a combination of Virtual Forward Synthesis (VFS) and polymer expert language models and qualified by stringent chemical heuristics. VFS is used to generate ROP polymers by applying known reactions to existing monomers. The polymer foundation models polyBART and POLYT5 further enable the generation of ROP candidates, with polyBART exploring its learned latent space and POLYT5 producing candidates via sequence-to-sequence generation. The resulting ROP polymers are subjected to robust filtering criteria to ensure novelty, validity and overall data quality through a combination of automated validation pipelines and a comprehensive set of chemist-informed heuristic rules introduced in this work for the first time. We hope that this dataset will serve as a valuable resource for downstream sustainable applications.

cond-mat.soft

On-the-Fly Machine-Learned Force Fields for High-Fidelity Polymer Glass Transition Simulations

Predicting polymer glass transition temperatures (Tg) with first-principles fidelity has long remained out of reach, as cooling multi-thousand-atom systems over a broad temperature range at acceptable rates exceeds the computational limits of ab initio molecular dynamics (AIMD). Here we employ a hybrid scheme that merges AIMD with accelerated on-the-fly (OTF) machine-learned force-field (MLFF) construction, enabling Tg prediction at quantum-mechanical accuracy with near-classical computational cost. The OTF protocol to construct MLFFs adaptively triggers first-principles calculations only when newly encountered configurations lie outside the current model's domain of confidence, allowing robust, parameter-free MLFFs to be built from merely 1000 AIMD-sampled configurations per polymer. These MLFFs are then utilized to perform long-time cooling simulations on amorphous supercells containing several thousand atoms. Applied across twelve polymers spanning aromatic, aliphatic, heteroatomic, and branched chemistries, the method yields predictions in excellent accord with experiment while reducing computational cost by approximately six orders of magnitude relative to AIMD. This work establishes a new paradigm for predictive polymer modeling, demonstrating that OTF-MLFFs provide a generalizable, accurate, and scalable route to simulating the thermophysical behavior of complex disordered materials at near quantum-mechanical fidelity.

cond-mat.mtrl-sci

LLM-Augmented Chemical Synthesis and Design Decision Programs

Retrosynthesis, the process of breaking down a target molecule into simpler precursors through a series of valid reactions, stands at the core of organic chemistry and drug development. Although recent machine learning (ML) research has advanced single-step retrosynthetic modeling and subsequent route searches, these solutions remain restricted by the extensive combinatorial space of possible pathways. Concurrently, large language models (LLMs) have exhibited remarkable chemical knowledge, hinting at their potential to tackle complex decision-making tasks in chemistry. In this work, we explore whether LLMs can successfully navigate the highly constrained, multi-step retrosynthesis planning problem. We introduce an efficient scheme for encoding reaction pathways and present a new route-level search strategy, moving beyond the conventional step-by-step reactant prediction. Through comprehensive evaluations, we show that our LLM-augmented approach excels at retrosynthesis planning and extends naturally to the broader challenge of synthesizable molecular design.

cs.AI

Load-dependent Hardness Prediction for Materials using Machine Learning

Superhard materials are critical for wear-resistant and high-stress applications. Conventional approaches correlating hardness with elastic moduli derived from DFT calculations enable rapid screening but overlook the strong load dependence of hardness. In this work, machine learning (ML) models were developed using a large, curated dataset of load-dependent experimental Vickers hardness (Hv) measurements. Moderate correlation was observed between experimental and DFT-based Hv values, whereas a single-task ML model trained solely on experimental data outperformed multi-task models that combined experimental and computed data. The superior performance of the single-task model highlights that explicit inclusion of indentation load, along with compositional, electronic, and structural descriptors, is essential and sufficient for accurate hardness prediction, beyond what can be achieved using DFT-accessible bulk and shear moduli alone (or in tandem with experimental data). These results emphasize the importance of high-quality experimental data and explicit inclusion of measurement conditions, particularly load, in the development of reliable hardness prediction models.

cond-mat.mtrl-sci

Retrieval Augmented Generation of Literature-derived Polymer Knowledge: The Example of a Biodegradable Polymer Expert System

Polymer literature contains a large and growing body of experimental knowledge, yet much of it is buried in unstructured text and inconsistent terminology, making systematic retrieval and reasoning difficult. Existing tools typically extract narrow, study-specific facts in isolation, failing to preserve the cross-study context required to answer broader scientific questions. Retrieval-augmented generation (RAG) offers a promising way to overcome this limitation by combining large language models (LLMs) with external retrieval, but its effectiveness depends strongly on how domain knowledge is represented. In this work, we develop two retrieval pipelines: a dense semantic vector-based approach (VectorRAG) and a graph-based approach (GraphRAG). Using over 1,000 polyhydroxyalkanoate (PHA) papers, we construct context-preserving paragraph embeddings and a canonicalized structured knowledge graph supporting entity disambiguation and multi-hop reasoning. We evaluate these pipelines through standard retrieval metrics, comparisons with general state-of-the-art systems such as GPT and Gemini, and qualitative validation by a domain chemist. The results show that GraphRAG achieves higher precision and interpretability, while VectorRAG provides broader recall, highlighting complementary trade-offs. Expert validation further confirms that the tailored pipelines, particularly GraphRAG, produce well-grounded, citation-reliable responses with strong domain relevance. By grounding every statement in evidence, these systems enable researchers to navigate the literature, compare findings across studies, and uncover patterns that are difficult to extract manually. More broadly, this work establishes a practical framework for building materials science assistants using curated corpora and retrieval design, reducing reliance on proprietary models while enabling trustworthy literature analysis at scale.

cs.CE

Accelerated design of proton exchange membranes for green hydrogen production with artificial intelligence

Water electrolysis is an eco-friendly method for hydrogen production that has reached significant levels of technological maturity. Among commercialized water-electrolysis technologies, proton-exchange membrane electrolyzers offer high current density, fast dynamic response, and compact system design, among other advantages. On the other hand, managing their high capital cost and the ``forever-chemistry'' nature of Nafion, a perfluorinated proton-exchange membrane widely used in such devices, remains a major challenge. Searches for fluorine-free replacements for Nafion, pursued largely through physical experimentation, have been active for decades with limited success. In this work, we develop and demonstrate an AI-based strategy for designing proton-exchange membranes for electrolyzers. Two key components of this strategy are an implementation of the virtual forward-synthesis approach and a set of machine-learning predictive models for essential application-inspired membrane properties; the former generates a vast space of millions of synthesizable polymers, which are then evaluated and screened by the latter. The strategy is validated against experimental data for known membranes and then applied to design over 1700 synthesizable candidates. This article concludes with a forward-looking vision in which the strategy could be elevated into an interactive and iterative scheme that is based on large language models to facilitate materials design in multiple ways.

cond-mat.soft

polyRETRO: a Language Model Approach to predict Polymerization Class and Monomer(s) for a Target Polymer

While machine learning has transformed polymer design by enabling rapid property prediction and candidate generation, translating these designs into experimentally realizable materials remains a critical challenge. Traditionally, the synthesis of target polymers has relied heavily on expert intuition and prior experience. The lack of automated retrosynthetic tools to assist chemists, limit the rapid practical impact of data-driven polymer discovery. To expedite lab-scale validation and beyond, we present a retrosynthetic framework that leverages large language models (LLMs) to guide polymer synthesis. Our approach, which we call polyRETRO, involves two key steps: 1) predicting the most likely polymerization reaction class of a target polymer and 2) identifying the underlying chemical transformation templates and the corresponding monomers, using primarily natural-language based constructs. This LLM-driven framework enables direct retrosynthetic analysis given just the target polymer SMILES string. polyRETRO constitutes a initial step towards a scalable, interpretable, and generalizable approach to bridge the gap between computational design and experimental synthesis.

cond-mat.soft

AI-assisted design of chemically recyclable polymers for food packaging

Polymer packaging plays a crucial role in food preservation but poses major challenges in recycling and environmental persistence. To address the need for sustainable, high-performance alternatives, we employed a polymer informatics workflow to identify single- and multi-layer drop-in replacements for polymer-based packaging materials. Machine learning (ML) models, trained on carefully curated polymer datasets, predicted eight key properties across a library of approximately 7.4 million ring-opening polymerization (ROP) polymers generated by virtual forward synthesis (VFS). Candidates were prioritized by the enthalpy of polymerization, a critical metric for chemical recyclability. This screening yielded thousands of promising candidates, demonstrating the feasibility of replacing diverse packaging architectures. We then experimentally validated poly(p-dioxanone) (poly-PDO), an existing ROP polymer whose barrier performance had not been previously reported. Validation showed that poly-PDO exhibits strong water barrier performance, mechanical and thermal properties consistent with predictions, and excellent chemical recyclability (95% monomer recovery), thereby meeting the design targets and underscoring its potential for sustainable packaging. These findings highlight the power of informatics-driven approaches to accelerate the discovery of sustainable polymers by uncovering opportunities in both existing and novel chemistries.

cond-mat.soft

AI-Driven Design of poly(ethylene terephthalate)-replacement copolymers

Poly(ethylene terephthalate) (PET), a widely used thermoplastic in packaging, textiles, and engineering applications, is valued for its strength, clarity, and chemical resistance. Increasing environmental impact concerns and regulatory pressures drive the search for alternatives with comparable or superior performance. We present an AI-driven polymer design pipeline employing virtual forward synthesis (VFS) to generate PET-replacement copolymers. Inspired by the esterification route of PET synthesis, we systematically combined a down-selected set of Toxic Substances Control Act (TSCA)-listed monomers to create 12,100 PET-like polymers. Machine learning models predicted glass transition temperature (Tg), band gap, and tendency to crystallize, for all designs. Multi-objective screening identified 1,108 candidates predicted to match or exceed PET in $T_{\rm g}$ and band gap, including the ``rediscovery'' of other known commercial PET-alternate polymers (e.g., PETG, Tritan, Ecozen) that provide retrospective validation of our design pipeline, demonstrating a capability to rapidly design experimentally feasible polymers at a scale. Furthermore, selected, entirely new (previously unknown) candidates designed here have been synthesized and characterized, providing a definitive validation of the design framework.

cond-mat.soft

An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design

Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.

cond-mat.mtrl-sci

polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Design

Designing polymers for targeted applications and accurately predicting their properties is a key challenge in materials science owing to the vast and complex polymer chemical space. While molecular language models have proven effective in solving analogous problems for molecular discovery, similar advancements for polymers are limited. To address this gap, we propose polyBART, a language model-driven polymer discovery capability that enables rapid and accurate exploration of the polymer design space. Central to our approach is Pseudo-polymer SELFIES (PSELFIES), a novel representation that allows for the transfer of molecular language models to the polymer space. polyBART is, to the best of our knowledge, the first language model capable of bidirectional translation between polymer structures and properties, achieving state-of-the-art results in property prediction and design of novel polymers for electrostatic energy storage. Further, polyBART is validated through a combination of both computational and laboratory experiments. We report what we believe is the first successful synthesis and validation of a polymer designed by a language model, predicted to exhibit high thermal degradation temperature and confirmed by our laboratory measurements. Our work presents a generalizable strategy for adapting molecular language models to the polymer space and introduces a polymer foundation model, advancing generative polymer design that may be adapted for a variety of applications.

cond-mat.soft

AI-Assisted Physics-Informed Predictions of Degradation Behavior of Polymeric Anion Exchange Membranes

The global transition to hydrogen-based energy infrastructures faces significant hurdles. Chief among these are the high costs and sustainability issues associated with acid-based proton exchange membrane fuel cells. Anion exchange membrane (AEM) fuel cells offer promising cost-effective alternatives, yet their widespread adoption is limited by rapid degradation in alkaline environments. Here, we develop a framework that integrates mechanistic insights with machine learning, enabling the identification of generalized degradation behavior across diverse polymeric AEM chemistries and operating conditions. Our model successfully predicts long-term hydroxide conductivity degradation (up to 10,000 hours) from minimal early-time experimental data. This capability significantly reduces experimental burdens and may expedite the design of high-performance, durable AEM materials.

cond-mat.soft

polyGen: A Learning Framework for Atomic-level Polymer Structure Generation

Synthetic polymeric materials underpin fundamental technologies in the energy, electronics, consumer goods, and medical sectors, yet their development still suffers from prolonged design timelines. Although polymer informatics tools have supported speedup, polymer simulation protocols continue to face significant challenges in the on-demand generation of realistic 3D atomic structures that respect conformational diversity. Generative algorithms for 3D structures of inorganic crystals, bio-polymers, and small molecules exist, but have not addressed synthetic polymers because of challenges in representation and dataset constraints. In this work, we introduce polyGen, the first generative model designed specifically for polymer structures from minimal inputs such as the repeat unit chemistry alone. polyGen combines graph-based encodings with a latent diffusion transformer using positional biased attention for realistic conformation generation. Given the limited dataset of 3,855 DFT-optimized polymer structures, we incorporate joint training with small molecule data to enhance generation quality. We also establish structure matching criteria to benchmark our approach on this novel problem. polyGen overcomes the limitations of traditional crystal structure prediction methods for polymers, successfully generating realistic and diverse linear and branched conformations, with promising performance even on challenging large repeat units. As the first atomic-level proof-of-concept capturing intrinsic polymer flexibility, it marks a new capability in material structure generation.

cs.CE

Benchmarking Large Language Models for Polymer Property Predictions

Machine learning has revolutionized polymer science by enabling rapid property prediction and generative design. Large language models (LLMs) offer further opportunities in polymer informatics by simplifying workflows that traditionally rely on large labeled datasets, handcrafted representations, and complex feature engineering. LLMs leverage natural language inputs through transfer learning, eliminating the need for explicit fingerprinting and streamlining training. In this study, we finetune general purpose LLMs -- open-source LLaMA-3-8B and commercial GPT-3.5 -- on a curated dataset of 11,740 entries to predict key thermal properties: glass transition, melting, and decomposition temperatures. Using parameter-efficient fine-tuning and hyperparameter optimization, we benchmark these models against traditional fingerprinting-based approaches -- Polymer Genome, polyGNN, and polyBERT -- under single-task (ST) and multi-task (MT) learning. We find that while LLM-based methods approach traditional models in performance, they generally underperform in predictive accuracy and efficiency. LLaMA-3 consistently outperforms GPT-3.5, likely due to its tunable open-source architecture. Additionally, ST learning proves more effective than MT, as LLMs struggle to capture cross-property correlations, a key strength of traditional methods. Analysis of molecular embeddings reveals limitations of general purpose LLMs in representing nuanced chemo-structural information compared to handcrafted features and domain-specific embeddings. These findings provide insight into the interplay between molecular embeddings and natural language processing, guiding LLM selection for polymer informatics.

cs.CE

MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge

Despite recent advances in large language models (LLMs) for materials science, there is a lack of benchmarks for evaluating their domain-specific knowledge and complex reasoning abilities. To bridge this gap, we introduce MSQA, a comprehensive evaluation benchmark of 1,757 graduate-level materials science questions in two formats: detailed explanatory responses and binary True/False assessments. MSQA distinctively challenges LLMs by requiring both precise factual knowledge and multi-step reasoning across seven materials science sub-fields, such as structure-property relationships, synthesis processes, and computational modeling. Through experiments with 10 state-of-the-art LLMs, we identify significant gaps in current LLM performance. While API-based proprietary LLMs achieve up to 84.5% accuracy, open-source (OSS) LLMs peak around 60.5%, and domain-specific LLMs often underperform significantly due to overfitting and distributional shifts. MSQA represents the first benchmark to jointly evaluate the factual and reasoning capabilities of LLMs crucial for LLMs in advanced materials science.

cs.AI

Polymer Composites Informatics for Flammability, Thermal, Mechanical and Electrical Property Predictions

Polymer composite performance depends significantly on the polymer matrix, additives, processing conditions, and measurement setups. Traditional physics-based optimization methods for these parameters can be slow, labor-intensive, and costly, as they require physical manufacturing and testing. Here, we introduce a first step in extending Polymer Informatics, an AI-based approach proven effective for neat polymer design, into the realm of polymer composites. We curate a comprehensive database of commercially available polymer composites, develop a scheme for machine-readable data representation, and train machine-learning models for 15 flame-resistant, mechanical, thermal, and electrical properties, validating them on entirely unseen data. Future advancements are planned to drive the AI-assisted design of functional and sustainable polymer composites.

cond-mat.soft

An Informatics Framework for the Design of Sustainable, Chemically Recyclable, Synthetically-Accessible and Durable Polymers

We present a novel approach to design durable and chemically recyclable ring-opening polymerization (ROP) class polymers. This approach employs digital reactions using virtual forward synthesis (VFS) to generate over 7 million ROP polymers and machine learning techniques to rapidly predict thermal, thermodynamic and mechanical properties crucial for application-specific performance and recyclability. This combined methodology enables the generation and evaluation of millions of hypothetical ROP polymers from known and commercially available molecules, guiding the selection of approximately 35,000 candidates with optimal features for sustainability and practical utility. Three of these recommended candidates have passed validation tests in the physical lab - two of the three by others, as published previously elsewhere, and one of them is a new thiocane polymer synthesized, tested and reported here. This paper presents the framework, methodology, and initial findings of our study, highlighting the potential of VFS and machine learning to enable a large-scale search of the polymer universe and advance the development of recyclable and environmentally benign polymers.

physics.chem-ph

A Physics-Enforced Neural Network to Predict Polymer Melt Viscosity

Achieving superior polymeric components through additive manufacturing (AM) relies on precise control of rheology. One key rheological property particularly relevant to AM is melt viscosity ($η$). Melt viscosity is influenced by polymer chemistry, molecular weight ($M_w$), polydispersity, induced shear rate ($\dotγ$), and processing temperature ($T$). The relationship of $η$ with $M_w$, $\dotγ$, and $T$ may be captured by parameterized equations. Several physical experiments are required to fit the parameters, so predicting $η$ of a new polymer material in unexplored physical domains is a laborious process. Here, we develop a Physics-Enforced Neural Network (PENN) model that predicts the empirical parameters and encodes the parametrized equations to calculate $η$ as a function of polymer chemistry, $M_w$, polydispersity, $\dotγ$, and $T$. We benchmark our PENN against physics-unaware Artificial Neural Network (ANN) and Gaussian Process Regression (GPR) models. Finally, we demonstrate that the PENN offers superior values of $η$ when extrapolating to unseen values of $M_w$, $\dotγ$, and $T$ for sparsely seen polymers.

cs.CE