SearcharxivSearch

arXiv subjects

Haewon Kim

Publications and source records attributed to Haewon Kim.

2 recordsLinked to original sources

Data-driven Prediction of Ionic Conductivity in Solid-State Electrolytes with Machine Learning and Large Language Models

Solid-state electrolytes (SSEs) are attractive for next-generation lithium-ion batteries due to improved safety and stability but their low room-temperature ionic conductivity hinders practical application. Experimental synthesis and testing of new SSEs remain time-consuming and resource intensive. Machine learning (ML) offers an accelerated route for SSE discovery; however, composition-only models neglect structural factors important for ion transport while graph neural networks (GNNs) are challenged by the scarcity of structure-labeled conductivity data and the prevalence of crystallographic disorder in CIFs. Here, we train two complementary predictors on the same room-temperature, structure-labeled dataset (n = 499). A gradient-boosted tree regressor (GBR) combining stoichiometric and geometric descriptors achieves best performance (MAE = 0.543 in log(S cm-1)), and Shapley Additive exPlanations (SHAP) identifies probe-occupiable volume (POAV) and lattice parameters as key correlations for conductivity. In parallel, we fine-tune large language models (LLMs) using compact text prompts derived from CIF metadata (formula with optional symmetry and disorder tags), avoiding direct use of raw atomic coordinates. Notably, Llama-3.1-8B-Instruct achieves high accuracy (MAE = 0.657 in log(S cm-1)) using formula and symmetry information, eliminating the need for numerical feature extraction from CIF files. Together, these results show that global geometric descriptors improve tree-based predictions and enable interpretable structure-property analysis, while LLMs provide a competitive low-preprocessing alternative for rapid SSE screening.

cond-mat.mtrl-sci

Investigation of finite-sample properties of robust location and scale estimators

When the experimental data set is contaminated, we usually employ robust alternatives to common location and scale estimators such as the sample median and Hodges-Lehmann estimators for location and the sample median absolute deviation and Shamos estimators for scale. It is well known that these estimators have high positive asymptotic breakdown points and are Fisher-consistent as the sample size tends to infinity. To the best of our knowledge, the finite-sample properties of these estimators, depending on the sample size, have not well been studied in the literature. In this paper, we fill this gap by providing their closed-form finite-sample breakdown points and calculating the unbiasing factors and relative efficiencies of the robust estimators through the extensive Monte Carlo simulations up to the sample size 100. The numerical study shows that the unbiasing factor improves the finite-sample performance significantly. In addition, we provide the predicted values for the unbiasing factors obtained by using the least squares method which can be used for the case of sample size more than 100.

stat.ME