Searcharxiv⌕ Search

arXiv subjects

Sudheesh Kumar Ethirajan

Publications and source records attributed to Sudheesh Kumar Ethirajan.

3 recordsLinked to original sources

LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature

Wide access to advanced experimental methods in materials science has given rise to an abundance of procedural knowledge, which is scattered across decades of scientific literature and recorded in unstructured formats that are challenging to analyze systematically. In this work, we present LeMat-Synth Parser, a modular, open-source, and multi-modal extraction toolbox that utilizes large language models (LLMs) and vision language models (VLMs) to automatically structure synthesis protocols and performance metrics extracted from both text and figures of publications. Applying LeMat-Synth Parser to 81K open-access publications, we curate LeMat-Synth, an extensive dataset of 58K synthesis procedures and to our knowledge the largest and most diverse structured inorganic materials synthesis dataset to date, covering 35 synthesis methods and 16 material classes based on a domain-specific ontology. We validate extraction quality against annotations by domain experts and a scalable LLM-as-a-judge framework, and benchmark a suite of models to identify optimal configurations and characterize cross-model biases. To demonstrate the extensibility of LeMat-Synth Parser, we apply it to two distinct domains. First, we link synthesis protocols and catalyst identity to thermocatalytic performance across a corpus of ammonia-decomposition publications. Second, we cross-validate text- and figure-reported critical transition temperatures across 1,384 superconductivity papers, then use the validated pipeline to recover the critical transition temperature for every composition in a sample series. We release LeMat-Synth Parser and the LeMat-Synth dataset openly on GitHub and Hugging Face

cs.DL↗

Electrostatic Phenomenology Benchmarks for Machine-Learned Interatomic Potentials in Electrochemistry: Beyond the Energy-Force Metric

Accurate treatment of long-range interactions in machine learning interatomic potentials (MLIPs) is essential for electrochemical simulations. However, aggregate energy and force errors alone are insufficient to establish an MLIP's physical accuracy since they do not detect qualitative inconsistencies in the model such as the prediction of image-charge attraction, dielectric screening, or charge transfer. We introduce a benchmark suite EPhEct (Electrostatic Phenomena for Electrochemistry) of focused test cases designed to evaluate MLIPs on electrochemically relevant physical phenomena. The tests probe for image-charge attraction at a metal electrode, the splitting between longitudinal and transverse optical phonons as a probe of ionic and electronic screening, the dipole moment of interfacial water, and Fermi-level pinning during ion discharge. These tests establish a qualitative diagnostic routine complementary to aggregate energy-force metrics.

cond-mat.mtrl-sci↗

Generalization of Graph-Based Active Learning Relaxation Strategies Across Materials

Although density functional theory (DFT) has aided in accelerating the discovery of new materials, such calculations are computationally expensive, especially for high-throughput efforts. This has prompted an explosion in exploration of machine learning assisted techniques to improve the computational efficiency of DFT. In this study, we present a comprehensive investigation of the broader application of Finetuna, an active learning framework to accelerate structural relaxation in DFT with prior information from Open Catalyst Project pretrained graph neural networks. We explore the challenges associated with out-of-domain systems: alcohol ($C_{>2}$) on metal surfaces as larger adsorbates, metal-oxides with spin polarization, and three-dimensional (3D) structures like zeolites and metal-organic-frameworks. By pre-training machine learning models on large datasets and fine-tuning the model along the simulation, we demonstrate the framework's ability to conduct relaxations with fewer DFT calculations. Depending on the similarity of the test systems to the training systems, a more conservative querying strategy is applied. Our best-performing Finetuna strategy reduces the number of DFT single-point calculations by 80% for alcohols and 3D structures, and 42% for oxide systems.

cond-mat.mtrl-sci↗