SearcharxivSearch

arXiv subjects

Youngwoo Cho

Publications and source records attributed to Youngwoo Cho.

8 recordsLinked to original sources

Robust and Interpretable Adaptation of Equivariant Materials Foundation Models via Sparsity-promoting Fine-tuning

Pre-trained materials foundation models, or machine learning interatomic potentials, leverage general physicochemical knowledge to effectively approximate potential energy surfaces. However, they often require domain-specific calibration due to physicochemical diversity as well as mismatches between practical computational settings and those used in constructing the pre-training data. To address this, we propose a sparsity-promoting fine-tuning method that selectively updates model parameters by exploiting the structural properties of E(3)-equivariant materials foundation models. On energy and force prediction tasks across molecular and crystalline benchmarks, our method matches or surpasses full fine-tuning and equivariant low-rank adaptation while updating only $\sim$3~\% of parameters, and in some cases as little as $\sim$0.5~\%. Beyond energy and force calibration, we further demonstrate task generalizability by applying our method to magnetic moment prediction and magnetism-aware total energy modeling. Finally, analysis of sparsity patterns reveals physically interpretable signatures, such as enhanced $d$-orbital contributions in transition metal systems. Overall, our results establish sparsity-promoting fine-tuning as a flexible and interpretable method for domain specialization of equivariant materials foundation models.

cs.LG

Plasma GraphRAG: Physics-Grounded Parameter Selection for Gyrokinetic Simulations

Accurate parameter selection is fundamental to gyrokinetic plasma simulations, yet current practices rely heavily on manual literature reviews, leading to inefficiencies and inconsistencies. We introduce Plasma GraphRAG, a novel framework that integrates Graph Retrieval-Augmented Generation (GraphRAG) with large language models (LLMs) for automated, physics-grounded parameter range identification. By constructing a domain-specific knowledge graph from curated plasma literature and enabling structured retrieval over graph-anchored entities and relations, Plasma GraphRAG enables LLMs to generate accurate, context-aware recommendations. Extensive evaluations across five metrics, comprehensiveness, diversity, grounding, hallucination, and empowerment, demonstrate that Plasma GraphRAG outperforms vanilla RAG by over $10\%$ in overall quality and reduces hallucination rates by up to $25\%$. {Beyond enhancing simulation reliability, Plasma GraphRAG offers a methodology for accelerating scientific discovery across complex, data-rich domains.

physics.plasm-ph

Machine learning prediction of plasma behavior from discharge configurations on WEST

Accurately predicting plasma behavior based on discharge configurations is essential for the safe and efficient operation of tokamak experiments. While physics-based integrated modeling codes provide valuable insights, their high computational cost limits their applicability for fast scenario design and control optimization. In this study, we propose a transformer-based machine learning model to predict key global plasma parameters on the WEST tokamak, including the normalized beta ($\beta_{n}$), toroidal beta ($\beta_{t}$), poloidal beta ($\beta_{p}$), plasma stored energy ($W_{\mathrm{mhd}}$), safety factor at the magnetic axis ($q_{0}$), and safety factor at the 95% flux surface ($q_{95}$). The model uses only signals that can be defined before the discharge, such as magnetic coil currents, auxiliary heating power, plasma current reference, and line-averaged plasma density. Trained on 550 discharges from the WEST campaigns, the model demonstrates an average mean square error (MSE) loss of 0.026, an average coefficient of determination $R^{2}$ of 0.94, and achieves inference times on the order of 0.1 seconds. These results highlight the potential of data-driven surrogate models for assisting in discharge planning, scenario evaluation, and real-time control of tokamak plasmas.

physics.plasm-ph

Reconstructing High-fidelity Plasma Turbulence with Data-driven Tuning of Diffusion in Low Resolution Grids

Developing physically consistent closure models is a longstanding challenge in simulating plasma turbulence, even in minimal systems such as the two-field Hasegawa-Wakatani (HW) model, which captures essential features of drift-wave turbulence with a reduced set of variables. In this work, we leverage theoretical insights from Direct Interaction Approximation (DIA) to construct a six-term closure structure that captures the dominant turbulent transport processes, including both diffusion and hyper-diffusion. While the mathematical form of the closure is fully prescribed by DIA, the corresponding transport coefficients are learned from data using physics-informed neural networks (PINNs). The resulting Extended HW model with Closure (EHW-C) model reveals several nontrivial features of plasma turbulence: notably, some inferred coefficients become negative in certain regimes, indicating inverse transport, a phenomenon absent in conventional closure models. Moreover, the EHW-C model accurately reproduces the spectral and flux characteristics of high-resolution Direct Numerical Simulations (DNS), while requiring only one-eighth the spatial resolution per direction, yielding a tenfold speed-up. This work demonstrates how theory-guided machine learning can both enhance computational efficiency and uncover emergent transport mechanisms in strongly nonlinear plasma systems.

physics.plasm-ph

A high-fidelity surrogate model for the ion temperature gradient (ITG) instability using a small expensive simulation dataset

One of the main challenges in building high-fidelity surrogate models of tokamak turbulence is the substantial demand for high-quality data. Typically, producing high-quality data involves simulating complex physical processes, which requires extensive computing resources. In this work, we propose a fine tuning-based approach to develop the surrogate model that reduces the amount of high-quality data required by 80\%. We demonstrate the effectiveness of this approach by constructing a proof-of-principle ITG surrogate model using datasets generated from two gyrokinetic codes, GKW and GX. GX needs in terms of computing resources are much lighter than GKW. Remarkably, the surrogate models' performance remain nearly the same whether trained on 798 GKW results alone or 159 GKW results plus an additional 11979 GX results. These encouraging outcomes indicate that fine tuning methods can significantly decrease the high-quality data needed to develop the simulation-driven surrogate model. Moreover, the approach presented here has the potential to facilitate surrogate model development for heavy codes and may ultimately pave the way for digital twin systems of tokamaks.

physics.plasm-ph

HistRED: A Historical Document-Level Relation Extraction Dataset

Despite the extensive applications of relation extraction (RE) tasks in various domains, little has been explored in the historical context, which contains promising data across hundreds and thousands of years. To promote the historical RE research, we present HistRED constructed from Yeonhaengnok. Yeonhaengnok is a collection of records originally written in Hanja, the classical Chinese writing, which has later been translated into Korean. HistRED provides bilingual annotations such that RE can be performed on Korean and Hanja texts. In addition, HistRED supports various self-contained subtexts with different lengths, from a sentence level to a document level, supporting diverse context settings for researchers to evaluate the robustness of their RE models. To demonstrate the usefulness of our dataset, we propose a bilingual RE model that leverages both Korean and Hanja contexts to predict relations between entities. Our model outperforms monolingual baselines on HistRED, showing that employing multiple language contexts supplements the RE predictions. The dataset is publicly available at: https://huggingface.co/datasets/Soyoung/HistRED under CC BY-NC-ND 4.0 license.

cs.CL

Enemy Spotted: in-game gun sound dataset for gunshot classification and localization

Recently, deep learning-based methods have drawn huge attention due to their simple yet high performance without domain knowledge in sound classification and localization tasks. However, a lack of gun sounds in existing datasets has been a major obstacle to implementing a support system to spot criminals from their gunshots by leveraging deep learning models. Since the occurrence of gunshot is rare and unpredictable, it is impractical to collect gun sounds in the real world. As an alternative, gun sounds can be obtained from an FPS game that is designed to mimic real-world warfare. The recent FPS game offers a realistic environment where we can safely collect gunshot data while simulating even dangerous situations. By exploiting the advantage of the game environment, we construct a gunshot dataset, namely BGG, for the firearm classification and gunshot localization tasks. The BGG dataset consists of 37 different types of firearms, distances, and directions between the sound source and a receiver. We carefully verify that the in-game gunshot data has sufficient information to identify the location and type of gunshots by training several sound classification and localization baselines on the BGG dataset. Afterward, we demonstrate that the accuracy of real-world firearm classification and localization tasks can be enhanced by utilizing the BGG dataset.

cs.SD

Knowledge Graph-based Question Answering with Electronic Health Records

Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine. In this light, QA on Electronic Health Records (EHR), namely EHR QA, can work as a crucial milestone towards developing an intelligent agent in healthcare. EHR data are typically stored in a relational database, which can also be converted to a directed acyclic graph, allowing two approaches for EHR QA: Table-based QA and Knowledge Graph-based QA. We hypothesize that the graph-based approach is more suitable for EHR QA as graphs can represent relations between entities and values more naturally compared to tables, which essentially require JOIN operations. In this paper, we propose a graph-based EHR QA where natural language queries are converted to SPARQL instead of SQL. To validate our hypothesis, we create four EHR QA datasets (graph-based VS table-based, and simplified database schema VS original database schema), based on a table-based dataset MIMICSQL. We test both a simple Seq2Seq model and a state-of-the-art EHR QA model on all datasets where the graph-based datasets facilitated up to 34% higher accuracy than the table-based dataset without any modification to the model architectures. Finally, all datasets are open-sourced to encourage further EHR QA research in both directions.

cs.DB