SearcharxivSearch

arXiv subjects

Honghui Kim

Publications and source records attributed to Honghui Kim.

7 recordsLinked to original sources

MADField: Multi-fidelity Amortized Density Field for Adsorption in Nanoporous Materials

High-throughput computational screening of nanoporous materials for gas storage and separation requires fast and accurate characterization of adsorption equilibrium. Particle-based grand canonical Monte Carlo (GCMC) and density-based classical density functional theory (cDFT) provide simulation-based estimates of gas uptake and adsorbate density fields, but their speed-accuracy tradeoff remains insufficient for large-scale screening. In this work, we address this gap with Multi-fidelity Amortized Density Field for Adsorption in Nanoporous Materials (MADField), which reframes adsorption prediction as equilibrium density-field estimation. MADField learns from two complementary fidelities, combining broad and scalable cDFT density supervision with higher-fidelity GCMC density labels, and recovers gas uptake by integrating the predicted density field. MADField improves uptake accuracy over the strongest baselines by 6.0x for cDFT and 15.4x for GCMC, and its predicted fields accelerate cDFT solvers with 2.0x fewer iteration steps while recovering 42 percent of cases that fail under standard settings. Finally, we evaluate MADField for conventional CH4 working capacity screening on the 270k-structure ARC-MOF database. Within this space of extremely rare high-capacity targets, 167 in total, the model achieves 56x higher average precision than the strongest baseline and accelerates inference by five orders of magnitude compared to GCMC. By prioritizing the MADField rankings, selecting the top 1.7 percent of candidates recovers 95 percent of all targets, while the top 6 percent ensures 100 percent recall.

physics.comp-ph

Machine Learning Hamiltonians are Accurate Energy-Force Predictors

Recently, machine learning Hamiltonian (MLH) models have gained traction as fast approximations of electronic structures such as orbitals and electron densities, while also enabling direct evaluation of energies and forces from their predictions. However, despite their physical grounding, existing Hamiltonian models are evaluated mainly by reconstruction metrics, leaving it unclear how well they perform as energy-force predictors. We address this gap with a benchmark that computes energies and forces directly from predicted Hamiltonians. Within this framework, we propose QHFlow2, a state-of-the-art Hamiltonian model with an SO(2)-equivariant backbone and a two-stage edge update. QHFlow2 achieves $40\%$ lower Hamiltonian error than the previous best model with fewer parameters. Under direct evaluation on MD17/rMD17, it is the first Hamiltonian model to reach NequIP-level force accuracy while achieving up to $20\times$ lower energy MAE. On QH9, QHFlow2 reduces energy error by up to $20\times$ compared to MACE. Finally, we demonstrate that QHFlow2 exhibits consistent scaling behavior with respect to model capacity and data, and that improvements in Hamiltonian accuracy effectively translate into more accurate energy and force computations.

physics.comp-ph

AtomMOF: All-Atom Flow Matching for MOF-Adsorbate Structure Prediction

Deep generative models have shown promise for modeling metal-organic frameworks (MOFs), but existing approaches (1) rely on coarse-grained representations that assume fixed bond lengths and angles, and (2) neglect the MOF-adsorbate interactions, which are critical for downstream applications. We introduce AtomMOF, a scalable flow-based model built on an all-atom Diffusion Transformer that maps 2D molecular graphs of building blocks and adsorbates directly to equilibrium 3D structures without imposing structural constraints. We further present scaling laws for porous crystal generation, indicating predictable performance gains with increased model capacity, and introduce Feynman-Kac steering guided by machine-learned interatomic potentials to improve geometric validity and sampling stability. On the (MOF-only) BW dataset, AtomMOF increases the match rate by 35.00% and reduces RMSD by 32.64%. On the ODAC25 dataset (MOF-adsorbate), AtomMOF is substantially more sample-efficient than grand canonical Monte Carlo in recovering adsorption configurations and can identify candidates with lower adsorption energies than the reference dataset. Code is available at https://github.com/nayoung10/AtomMOF.

cond-mat.mtrl-sci

CatFlow: Co-generation of Slab-Adsorbate Systems via Flow Matching

Discovering heterogeneous catalysts tailored for specific reaction intermediates remains a fundamental bottleneck in materials science. While traditional trial-and-error methods and recent generative models have shown promise, they struggle to capture the intrinsic coupling between surface geometry and adsorbate interactions. To address this limitation, we propose CatFlow, a flow matching-based framework for de novo design and structure prediction of heterogeneous catalysts. Our model operates on a primitive cell-based factorized representation of the slab-adsorbate complex, reducing the number of learnable variables by an average of 9.2x while explicitly encoding the surface orientation of the slab-adsorbate interface. Experiments on the Open Catalyst 2020 dataset demonstrate that CatFlow significantly improves the structural fidelity of generated catalysts compared to autoregressive and sequential baselines. Further experiments show that the generated structures accurately capture the adsorption energy distributions of physically plausible interfaces and lie closer to thermodynamic local minima.

cond-mat.mtrl-sci

A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning

Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit from cross-modal or relational supervision. Protein-ligand co-folding provides a molecular analogue of such supervision by exposing models to atom-level ligand-protein interactions, raising the question of whether co-folding models can yield strong small-molecule representations. We study this question using Boltz2, a modern co-folding model, by transferring its atom-level ligand representations to standalone small-molecule tasks. Through systematic probing and distillation, we show that Boltz2 representations match or outperform existing models on the ADMET benchmark, accelerate molecular generative modeling, and improve sample efficiency in structure-guided ligand optimization. We further find that Boltz2 representations are complementary to those learned from conventional standalone molecular supervision, including 3D conformers, bioassay labels, and quantum-chemical properties. Finally, we extend representation alignment to reinforcement learning, showing that dense representation-level supervision can complement scalar rewards in molecular discovery. These results identify protein-ligand co-folding as a promising pretraining paradigm for small-molecule representation learning and position Boltz2 as a strong, off-the-shelf molecular foundation model.

q-bio.BM

LitMOF: An LLM Multi-Agent for Literature-Validated Metal-Organic Frameworks Database Correction and Expansion

Metal-organic framework (MOF) databases have grown rapidly through experimental deposition and large-scale literature extraction, but recent analyses show that nearly half of their entries contain substantial structural errors. These inaccuracies propagate through high-throughput screening and machine-learning workflows, limiting the reliability of data-driven MOF discovery. Correcting such errors is exceptionally difficult because true repairs require integrating crystallographic files, synthesis descriptions, and contextual evidence scattered across the literature. Here we introduce LitMOF, a large language model-driven multi-agent framework that validates crystallographic information directly from the original literature and cross-validates it with database entries to repair structural errors. Applying LitMOF to the experimental MOF database (the CSD MOF Subset), we constructed LitMOF-DB, a curated set of 189,567 computation-ready structures, including the successful repair of 9,227 invalid entries, which accounts for 69.1% of the CSD-derived not-computation-ready MOFs in the latest CoRE MOF DB. Additionally, the system uncovered 8,771 experimentally reported MOFs absent from existing resources, substantially expanding the known experimental design space. Using direct air capture screening as a case study, we demonstrate that structural errors severely distort predicted adsorption energies and CO2/H2O selectivity, leading to systematic misranking of materials, false positives, and the omission of high-performance candidates. This work establishes a scalable pathway toward self-correcting scientific databases and a generalizable approach for LLM-driven curation in materials science.

cs.DB

Assessing exchange-correlation functionals for heterogeneous catalysis of nitrogen species

Increasing interest in sustainable synthesis of ammonia, nitrates, and urea has led to an increase in studies of catalytic conversion between nitrogen-containing compounds using heterogeneous catalysts. Density functional theory (DFT) is commonly employed to obtain molecular-scale insight into these reactions, but there have been relatively few assessments of the exchange-correlation functionals that are best suited for heterogeneous catalysis of nitrogen compounds. Here, we assess a range of functionals ranging from the generalized gradient approximation (GGA) to the random phase approximation (RPA) for the formation energies of gas-phase nitrogen species, the lattice constants of representative solids from several common classes of catalysts (metals, oxides, and metal-organic frameworks (MOFs)), and the adsorption energies of a range of nitrogen-containing intermediates on these materials. The results reveal that the choice of exchange-correlation functional and van der Waals correction can have a surprisingly large effect and that increasing the level of theory does not always improve the accuracy for nitrogen-containing compounds. This suggests that the selection of functionals should be carefully evaluated on the basis of the specific reaction and material being studied.

cond-mat.mtrl-sci