SearcharxivSearch

arXiv subjects

Rohit Batra

Publications and source records attributed to Rohit Batra.

At least 19 recordsLinked to original sources

Efficient Semi-Automated Material Microstructure Analysis Using Deep Learning: A Case Study in Additive Manufacturing

Image segmentation is fundamental to microstructural analysis for defect identification and structure-property correlation, yet remains challenging due to pronounced heterogeneity in materials images arising from varied processing and testing conditions. Conventional image processing techniques often fail to capture such complex features rendering them ineffective for large-scale analysis. Even deep learning approaches struggle to generalize across heterogeneous datasets due to scarcity of high-quality labeled data. Consequently, segmentation workflows often rely on manual expert-driven annotations which are labor intensive and difficult to scale. Using an additive manufacturing (AM) dataset as a case study, we present a semi-automated active learning based segmentation pipeline that integrates a U-Net based convolutional neural network with an interactive user annotation and correction interface and a representative core-set image selection strategy. The active learning workflow iteratively updates the model by incorporating user corrected segmentations into the training pool while the core-set strategy identifies representative images for annotation. Three subset selection strategies, manual selection, uncertainty driven sampling and proposed maximin Latin hypercube sampling from embeddings (SMILE) method were evaluated over six refinement rounds. The SMILE strategy consistently outperformed other approaches, improving the macro F1 score from 0.74 to 0.93 while reducing manual annotation time by about 65 percent. The segmented defect regions were further analyzed using a coupled classification model to categorize defects based on microstructural characteristics and map them to corresponding AM process parameters. The proposed framework reduces labeling effort while maintaining scalability and robustness and is broadly applicable to image based analysis across diverse materials systems.

cs.CV

Automated Extraction of Multicomponent Alloy Data Using Large Language Models for Sustainable Design

The design of sustainable materials requires access to materials performance and sustainability data from literature corpus in an organized, structured and automated manner. Natural language processing approaches, particularly large language models (LLMs), have been explored for materials data extraction from the literature, yet often suffer from limited accuracy or narrow scope. In this work, an LLM-based pipeline is developed to accurately extract alloy-related information from both textual descriptions and tabular data across the literature on high-entropy (or multicomponent) alloys (HEA). Specifically two databases with 37,711 and 148,069 entries respectively are retrieved; one from the literature text, consisting of alloy composition, processing conditions, characterization methods, and reported properties, and other from the literature tables, consisting of property names, values, and units. The pipeline enhances materials-domain sensitivity through prompt engineering and retrieval-augmented generation and achieves F1-scores of 0.83 for textual extraction and 0.88 for tabular extraction, surpassing or matching existing approaches. Application of the pipeline to over 10,000 articles yields the largest publicly available multicomponent alloy database and reveals compositional and processing-property trends. The database is further employed for sustainability-aware materials selection in three application domains, i.e., lightweighting, soft magnetic, and corrosion-resistant, identifying multicomponent alloy candidates with more sustainable production while maintaining or exceeding benchmark performance. The pipeline developed can be easily generalized to other class of materials, and assist in development of comprehensive, accurate and usable databases for sustainable materials design.

cond-mat.mtrl-sci

Sustainable Materials Discovery in the Era of Artificial Intelligence

Artificial intelligence (AI) has transformed materials discovery, enabling rapid exploration of chemical space through generative models and surrogate screening. Yet current generative AI models for materials discovery, which now drive exploration of vast chemical and structural spaces, optimize candidates exclusively for structural stability and functional properties, with no integration of environmental assessment at any stage of the design loop. Prospective and ex-ante life cycle assessment methods exist and have been applied to emerging technologies, but they operate as standalone downstream analyses, not as active constraints within generative or active-learning pipelines. The result is that environmental feedback, even when produced, arrives after design decisions have been made rather than informing them. The disconnect between atomic-scale design and lifecycle assessment (LCA) reflects fundamental challenges: (i) data scarcity across heterogeneous sources, (ii) scale gaps from atoms to industrial systems, (iii) uncertainty in synthesis pathways, and (iv) the absence of frameworks that co-optimize performance with environmental impact. In this Perspective, we propose integrating upstream ML-assisted materials discovery with downstream LCA into the ML-LCA framework, comprising five components: information extraction for building materials-environment knowledge bases, harmonized databases linking properties to sustainability metrics, multi-scale models bridging atomic properties to lifecycle impacts, ensemble prediction of manufacturing pathways with uncertainty quantification, and uncertainty-aware optimization enabling simultaneous performance-sustainability navigation. Case studies spanning polymers, glass, photoresists, and cement demonstrate both necessity and feasibility while identifying material-specific integration challenges.

cond-mat.mtrl-sci

Active Learning Guided Computational Discovery of 2D Materials with Large Spin Hall Conductivity

Two-dimensional (2D) materials are promising candidates for next-generation spintronic devices due to their tunable properties and potential for efficient spin-charge interconversion. However, discovering materials with intrinsically high spin Hall conductivity (SHC) is hindered by the vast chemical space and expensive nature of conventional experimental and first-principles methods. In this work, we employ an active learning framework to accelerate the discovery of high-SHC 2D materials. Machine learning (ML) models were trained on SHC values computed from density functional theory calculations, incorporating the Kubo formalism via tight-binding Hamiltonians constructed from maximally localized Wannier functions, with explicit treatment of spin-orbit coupling. Starting from random but chemically diverse 24 2D systems, the dataset was expanded to 41 cases (from an overall pool of around 2000 materials) over three active learning loops using an expected improvement acquisition strategy. The ML technique successfully identified several high SHC candidates with the best candidate exhibiting a SHC of 271.52 (hbar/e) Ohm^-1, nearly 23 times higher than the top performer in the initial round. Beyond candidate discovery, several features such as orbital symmetry near the Fermi energy, types of atomic species, material composition, covalent radii, and electronegativity of constituent atoms were found to play critical role in shaping the spin Hall response in 2D systems. The data generated is made publicly available to facilitate further advances in 2D spintronics.

cond-mat.mtrl-sci

Physically Interpretable Interatomic Potentials via Symbolic Regression and Reinforcement Learning

The development of next-generation molecular simulation models requires moving beyond pre-defined functional forms toward machine learning (ML) techniques that directly capture multiscale physics. Here, we demonstrate such an approach using symbolic regression (SR) with equation learner networks and a reinforcement learning search engine to derive interpretable equations for interatomic interactions. Training data were generated through nested ensemble sampling with density functional theory (DFT) energetics, spanning crystalline to highly disordered states. The optimization of the learner network employed continuous-action Monte Carlo Tree Search (MCTS) combined with gradient descent, enabling efficient exploration of function space. For copper as a representative transition metal, an unconstrained search produced models that outperformed fixed-form Sutton-Chen EAM potentials. The SR-derived models (SR1 and SR2) reproduced key material properties - lattice constants, cohesive energies, equations of state, elastic constants, phonon dispersion, defect formation energies, surface/bulk energetics, and phase transformation with significantly improved accuracy. Furthermore, stringent melting simulations using two-phase solid-amorphous interfaces confirmed that SR models accurately capture the interplay of vibrational entropy, cohesive energy, and structural dynamics, surpassing SC-EAM in both qualitative and quantitative predictions. This highlights the potential of SR to deliver fast, accurate, flexible, and physically meaningful potentials, advancing predictive modeling across scales.

cond-mat.mtrl-sci

Prediction of Mechanical Properties and Thermodynamic Stability of Ti-N system using MTP Interatomic Potential

Ti-N material system have range of compounds with different stoichiometry like Ti2N, Ti3N2, Ti6N5, Ti4N3 alongwith Ti , TiN and solid solutions of N in Ti with a maximum of 23% solubility. In this work, we develop an interatomic potential based on moment tensor potential (MTP) that could reliably predict mechanical properties and thermodynamic stability of all Ti-N system. Taking into account the structural similarity and dissimilarity of various Ti-N system to choose training dataset was crucial for development of the potential. Root mean square error (RMSE) in prediction of formation energy using MTP potential compared to one calculated using density functional theory (DFT) for training dataset is 2.1 meV/atom and for testing dataset is 6.8 meV/atom. The frequency of absolute error in formation energy peaks at a maximum value of 3.8 meV/atom for system that was part of training dataset, while it peaks at 7.6 meV/atom for systems that are not part of the training dataset. Furthermore, the distribution and variability of elastic constants across compositions are systematically evaluated, revealing trends consistent with DFT benchmarks. The developed potential was used to predict energy of new phases in Ti-N system. We show that structures with N/Ti ratios ranging from 0 to 1 can be thermodynamically stable. A maximum deviation of 10 meV/atom from the convex hull plot of formation energy 0K was observed for a few system.

cond-mat.mtrl-sci

Enhancing Experimental Efficiency in Materials Design: A Comparative Study of Taguchi and Machine Learning Methods

Materials design problems often require optimizing multiple variables, rendering full factorial exploration impractical. Design of experiment (DOE) methods, such as Taguchi technique, are commonly used to efficiently sample the design space but they inherently lack the ability to capture non-linear dependency of process variables. In this work, we demonstrate how machine learning (ML) methods can be used to overcome these limitations. We compare the performance of Taguchi method against an active learning based Gaussian process regression (GPR) model in a wire arc additive manufacturing (WAAM) process to accurately predict aspects of bead geometry, including penetration depth, bead width, and height. While Taguchi method utilized a three-factor, five-level L25 orthogonal array to suggest weld parameters, the GPR model used an uncertainty-based exploration acquisition function coupled with latin hypercube sampling for initial training data. Accuracy and efficiency of both models was evaluated on 15 test cases, with GPR outperforming Taguchi in both metrics. This work applies to broader materials processing domain requiring efficient exploration of complex parameters.

cs.LG

Accelerated Design of Block Copolymers: An Unbiased Exploration Strategy via Fusion of Molecular Dynamics Simulations and Machine Learning

Star block copolymers (s-BCPs) have potential applications as novel surfactants or amphiphiles for emulsification, compatbilization, chemical transformations and separations. s-BCPs are star-shaped macromolecules comprised of linear chains of different chemical blocks (e.g., solvophilic and solvophobic blocks) that are covalently joined at one junction point. Various parameters of these macromolecules can be tuned to obtain desired surface properties, including the number of arms, composition of the arms, and the degree-of-polymerization of the blocks (or the length of the arm). This makes identification of the optimal s-BCP design highly non-trivial as the total number of plausible s-BCPs architectures is experimentally or computationally intractable. In this work, we use molecular dynamics (MD) simulations coupled with reinforcement learning based Monte Carlo tree search (MCTS) to identify s-BCPs designs that minimize the interfacial tension between polar and non-polar solvents. We first validate the MCTS approach for design of small- and medium-sized s-BCPs, and then use it to efficiently identify sequences of copolymer blocks for large-sized s-BCPs. The structural origins of interfacial tension in these systems are also identified using the configurations obtained from MD simulations. Chemical insights on the arrangement of copolymer blocks that promote lower interfacial tension were mined using machine learning (ML) techniques. Overall, this work provides an efficient approach to solve design problems via fusion of simulations and ML and provide important groundwork for future experimental investigation of s-BCPs sequences for various applications.

cond-mat.soft

Efficient Probabilistic Computing with Stochastic Perovskite Nickelates

Probabilistic computing has emerged as a viable approach to solve hard optimization problems. Devices with inherent stochasticity can greatly simplify their implementation in electronic hardware. Here, we demonstrate intrinsic stochastic resistance switching controlled via electric fields in perovskite nickelates doped with hydrogen. The ability of hydrogen ions to reside in various metastable configurations in the lattice leads to a distribution of transport gaps. With experimentally characterized p-bits, a shared-synapse p-bit architecture demonstrates highly-parallelized and energy-efficient solutions to optimization problems such as integer factorization and Boolean-satisfiability. The results introduce perovskite nickelates as scalable potential candidates for probabilistic computing and showcase the potential of light-element dopants in next-generation correlated semiconductors.

cond-mat.mtrl-sci

Learning with Delayed Rewards -- A case study on inverse defect design in 2D materials

Defect dynamics in materials are of central importance to a broad range of technologies from catalysis to energy storage systems to microelectronics. Material functionality depends strongly on the nature and organization of defects, their arrangements often involve intermediate or transient states that present a high barrier for transformation. The lack of knowledge of these intermediate states and the presence of this energy barrier presents a serious challenge for inverse defect design, especially for gradient-based approaches. Here, we present a reinforcement learning (Monte Carlo Tree Search) based on delayed rewards that allow for efficient search of the defect configurational space and allows us to identify optimal defect arrangements in low dimensional materials. Using a representative case of 2D MoS2, we demonstrate that the use of delayed rewards allows us to efficiently sample the defect configurational space and overcome the energy barrier for a wide range of defect concentrations (from 1.5% to 8% S vacancies), the system evolves from an initial randomly distributed S vacancies to one with extended S line defects consistent with previous experimental studies. Detailed analysis in the feature space allows us to identify the optimal pathways for this defect transformation and arrangement. Comparison with other global optimization schemes like genetic algorithms suggests that the MCTS with delayed rewards takes fewer evaluations and arrives at a better quality of the solution. The implications of the various sampled defect configurations on the 2H to 1T phase transitions in MoS2 are discussed. Overall, we introduce a Reinforcement Learning (RL) strategy employing delayed rewards that can accelerate the inverse design of defects in materials for achieving targeted functionality.

cond-mat.mtrl-sci

Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders

The design/discovery of new materials is highly non-trivial owing to the near-infinite possibilities of material candidates, and multiple required property/performance objectives. Thus, machine learning tools are now commonly employed to virtually screen material candidates with desired properties by learning a theoretical mapping from material-to-property space, referred to as the \emph{forward} problem. However, this approach is inefficient, and severely constrained by the candidates that human imagination can conceive. Thus, in this work on polymers, we tackle the materials discovery challenge by solving the \emph{inverse} problem: directly generating candidates that satisfy desired property/performance objectives. We utilize syntax-directed variational autoencoders (VAE) in tandem with Gaussian process regression (GPR) models to discover polymers expected to be robust under three extreme conditions: (1) high temperatures, (2) high electric field, and (3) high temperature \emph{and} high electric field, useful for critical structural, electrical and energy storage applications. This approach to learn from (and augment) human ingenuity is general, and can be extended to discover polymers with other targeted properties and performance measures.

cond-mat.soft

Polymer Informatics: Current Status and Critical Next Steps

Artificial intelligence (AI) based approaches are beginning to impact several domains of human life, science and technology. Polymer informatics is one such domain where AI and machine learning (ML) tools are being used in the efficient development, design and discovery of polymers. Surrogate models are trained on available polymer data for instant property prediction, allowing screening of promising polymer candidates with specific target property requirements. Questions regarding synthesizability, and potential (retro)synthesis steps to create a target polymer, are being explored using statistical means. Data-driven strategies to tackle unique challenges resulting from the extraordinary chemical and physical diversity of polymers at small and large scales are being explored. Other major hurdles for polymer informatics are the lack of widespread availability of curated and organized data, and approaches to create machine-readable representations that capture not just the structure of complex polymeric situations but also synthesis and processing conditions. Methods to solve inverse problems, wherein polymer recommendations are made using advanced AI algorithms that meet application targets, are being investigated. As various parts of the burgeoning polymer informatics ecosystem mature and become integrated, efficiency improvements, accelerated discoveries and increased productivity can result. Here, we review emergent components of this polymer informatics ecosystem and discuss imminent challenges and opportunities.

cond-mat.soft

Machine Learning the Metastable Phase Diagram of Materials

Phase diagrams are an invaluable tool for material synthesis and provide information on the phases of the material at any given thermodynamic condition. Conventional phase diagram generation involves experimentation to provide an initial estimate of thermodynamically accessible phases, followed by use of phenomenological models to interpolate between the available experimental data points and extrapolate to inaccessible regions. Such an approach, combined with first-principles calculations and data-mining techniques, has led to exhaustive thermodynamic databases albeit at distinct thermodynamic equilibria. In contrast, materials during their synthesis, operation, or processing, may not reach their thermodynamic equilibrium state but, instead, remain trapped in a local free energy minimum, that may exhibit desirable properties. Mapping these metastable phases and their thermodynamic behavior is highly desirable but currently lacking. Here, we introduce an automated workflow that integrates first principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the metastable phases of a given elemental composition. Using a representative material, carbon, with a vast number of metastable phases without parent in equilibrium, we demonstrate automatic mapping of hundreds of metastable states ranging from near equilibrium to those far-from-equilibrium. Moreover, we incorporate the free energy calculations into a neural-network-based learning of the equations of state that allows for construction of metastable phase diagrams. High temperature high pressure experiments using a diamond anvil cell on graphite sample coupled with high-resolution transmission electron microscopy are used to validate our metastable phase predictions. Our introduced approach is general and broadly applicable to single and multi-component systems.

cond-mat.mtrl-sci

Screening of Therapeutic Agents for COVID-19 using Machine Learning and Ensemble Docking Simulations

The world has witnessed unprecedented human and economic loss from the COVID-19 disease, caused by the novel coronavirus SARS-CoV-2. Extensive research is being conducted across the globe to identify therapeutic agents against the SARS-CoV-2. Here, we use a powerful and efficient computational strategy by combining machine learning (ML) based models and high-fidelity ensemble docking simulations to enable rapid screening of possible therapeutic molecules (or ligands). Our screening is based on the binding affinity to either the isolated SARS-CoV-2 S-protein at its host receptor region or to the Sprotein-human ACE2 interface complex, thereby potentially limiting and/or disrupting the host-virus interactions. We first apply our screening strategy to two drug datasets (CureFFI and DrugCentral) to identify hundreds of ligands that bind strongly to the aforementioned two systems. Candidate ligands were then validated by all atom docking simulations. The validated ML models were subsequently used to screen a large bio-molecule dataset (with nearly a million entries) to provide a rank-ordered list of ~19,000 potentially useful compounds for further validation. Overall, this work not only expands our knowledge of small-molecule treatment against COVID-19, but also provides an efficient pathway to perform high-throughput computational drug screening by combining quick ML surrogate models with expensive high-fidelity simulations, for accelerating the therapeutic cure of diseases.

q-bio.BM

Machine Learning for Multi-fidelity Scale Bridging and Dynamical Simulations of Materials

Molecular dynamics (MD) is a powerful and popular tool for understanding the dynamical evolution of materials at the nano and mesoscopic scales. There are various flavors of MD ranging from the high fidelity albeit computationally expensive ab-initio MD to relatively lower fidelity but much more efficient classical MD such as atomistic and coarse-grained models. Each of these different flavors of MD have been independently used by materials scientists to bring about breakthroughs in materials discovery and design. A significant gulf exists between the various MD flavors, each having varying levels of fidelity. The accuracy of DFT or ab-initio MD is generally much higher than that of classical atomistic simulations which is higher than that of coarse-grained models. Multi-fidelity scale bridging to combine the accuracy and flexibility of ab-initio MD with efficiency classical MD has been a longstanding goal. The advent of big-data analytics has brought to the forefront powerful machine learning methods that can be deployed to achieve this goal. Here, we provide our perspective on the challenges in multi-fidelity scale bridging and trace the developments leading up to the use of machine learning algorithms and data-science towards addressing this grand challenge.

physics.comp-ph

Machine Learning Models for the Lattice Thermal Conductivity Prediction of Inorganic Materials

The lattice thermal conductivity ($κ_{\rm L} $) is a critical property of thermoelectrics, thermal barrier coating materials and semiconductors. While accurate empirical measurements of $κ_{\rm L} $ are extremely challenging, it is usually approximated through computational approaches, such as semi-empirical models, Green-Kubo formalism coupled with molecular dynamics simulations, and first-principles based methods. However, these theoretical methods are not only limited in terms of their accuracy, but sometimes become computationally intractable owing to their cost. Thus, in this work, we build a machine learning (ML)-based model to accurately and instantly predict $κ_{\rm L}$ of inorganic materials, using a benchmark data set of experimentally measured $κ_{\rm L} $ of about 100 inorganic solids. We use advanced and universal feature engineering techniques along with the Gaussian process regression algorithm, and compare the performance of our ML model with past theoretical works. The trained ML model is not only helpful for rational design and screening of novel materials, but we also identify key features governing the thermal transport behavior in non-metals.

cond-mat.mtrl-sci

Machine Learning and Materials Informatics: Recent Applications and Prospects

Propelled partly by the Materials Genome Initiative, and partly by the algorithmic developments and the resounding successes of data-driven efforts in other domains, informatics strategies are beginning to take shape within materials science. These approaches lead to surrogate machine learning models that enable rapid predictions based purely on past data rather than by direct experimentation or by computations/simulations in which fundamental equations are explicitly solved. Data-centric informatics methods are becoming useful to determine material properties that are hard to measure or compute using traditional methods--due to the cost, time or effort involved--but for which reliable data either already exists or can be generated for at least a subset of the critical cases. Predictions are typically interpolative, involving fingerprinting a material numerically first, and then following a mapping (established via a learning algorithm) between the fingerprint and the property of interest. Fingerprints may be of many types and scales, as dictated by the application domain and needs. Predictions may also be extrapolative--extending into new materials spaces--provided prediction uncertainties are properly taken into account. This article attempts to provide an overview of some of the recent successful data-driven "materials informatics" strategies undertaken in the last decade, and identifies some challenges the community is facing and those that should be overcome in the near future.

cond-mat.mtrl-sci

Dopants Promoting Ferroelectricity in Hafnia: Insights From A Comprehensive Chemical Space Exploration

Although dopants have been extensively employed to promote ferroelectricity in hafnia films, their role in stabilizing the responsible ferroelectric non-equilibrium Pca21 phase is not well understood. In this work, using first principles computations, we investigate the influence of nearly 40 dopants on the phase stability in bulk hafnia to identify dopants that can favor formation of the polar Pca21 phase. Although no dopant was found to stabilize this polar phase as the ground state, suggesting that dopants alone cannot induce ferroelectricity in hafnia, Ca, Sr, Ba, La, Y and Gd were found to significantly lower the energy of the polar phase with respect to the equilibrium monoclinic phase. These results are consistent with the empirical measurements of large remnant polarization in hafnia films doped with these elements. Additionally, clear chemical trends of dopants with larger ionic radii and lower electronegativity favoring the polar Pca21 phase in hafnia were identified. For this polar phase, an additional bond between the dopant cation and the 2nd nearest oxygen neighbor was identified as the root-cause of these trends. Further, trivalent dopants (Y, La, and Gd) were revealed to stabilize the polar Pca21 phase at lower strains when compared to divalent dopants (Sr and Ba). Based on these insights, we predict that the lanthanide series metals, the lower half of alkaline earth metals (Ca, Sr and Ba) and Y as the most suitable dopants to promote ferroelectricity in hafnia.

cond-mat.mtrl-sci