SearcharxivSearch

arXiv subjects

Jong-Hoon Park

Publications and source records attributed to Jong-Hoon Park.

3 recordsLinked to original sources

Multi-View Molecular Representation Learning with Hierarchical Graphs and Contextualized Fingerprints

Molecular property prediction requires representations that generalize from limited labeled data to structurally novel compounds. Existing molecular pretraining methods often rely on a single view: graph-based approaches model atom-bond topology but provide limited fragment-level supervision, whereas fingerprint descriptors encode chemical patterns but are typically used as fixed auxiliary features. We propose HiFi-Mol, a multi-view framework that separately pretrains a hierarchical graph encoder and a contextualized fingerprint encoder before downstream integration. The graph branch uses fragment-aware masking with multi-resolution supervision to capture substructure-aware representations, while the fingerprint branch tokenizes active entries from seven fingerprint families and applies masked language modeling to learn contextualized embeddings. During fine-tuning, HiFi-Mol combines projected multi-resolution graph features with fingerprint embeddings for downstream prediction. Evaluated on MoleculeNet benchmarks under the scaffold split, HiFi-Mol achieves a 2.77% improvement in average ROC-AUC over the best baseline across eight classification tasks while maintaining competitive performance on three regression tasks. Further analyses reveal that fragment-aware masking improves graph representation quality, and classification results demonstrate dataset-dependent strengths of the individual graph and fingerprint variants, confirming that the two views provide complementary predictive signals.

cs.LG

Finely Tunable Thermal Expansion of NiTi by Stress-Induced Martensitic Transformation and Thermomechanical Training

Tailoring the thermal expansion of martensitic materials by crystallographic texture and anisotropic variation of lattice parameters is a promising route to a flexible design of thermally stable systems. NiTi alloys are prototype materials in this respect, with shape-memory and superelastic properties owing to their thermoelastic martensitic transformations. Here, we propose a method to realize finely tunable coefficients of thermal expansion (CTE) for the NiTi alloy based upon a special combination of mechanical and thermal training. We achieve a near-zero in-plane CTE that is smaller in value than that of the FeNi-based Invar alloy. Atomistic simulations and theoretical calculations guide the method design and clarify the underlying mechanisms of the relationship between the processing conditions, the microstructural evolution, and the thermal expansion behavior. The directions for further, finer adjustments of the CTE without constraints on the shape of the materials are indicated.

cond-mat.mtrl-sci

Domain Knowledge Infused Conditional Generative Models for Accelerating Drug Discovery

The role of Artificial Intelligence (AI) is growing in every stage of drug development. Nevertheless, a major challenge in drug discovery AI remains: Drug pharmacokinetic (PK) and Drug-Target Interaction (DTI) datasets collected in different studies often exhibit limited overlap, creating data overlap sparsity. Thus, data curation becomes difficult, negatively impacting downstream research investigations in high-throughput screening, polypharmacy, and drug combination. We propose xImagand-DKI, a novel SMILES/Protein-to-Pharmacokinetic/DTI (SP2PKDTI) diffusion model capable of generating an array of PK and DTI target properties conditioned on SMILES and protein inputs that exhibit data overlap sparsity. We infuse additional molecular and genomic domain knowledge from the Gene Ontology (GO) and molecular fingerprints to further improve our model performance. We show that xImagand-DKI-generated synthetic PK data closely resemble real data univariate and bivariate distributions, and can adequately fill in gaps among PK and DTI datasets. As such, xImagand-DKI is a promising solution for data overlap sparsity and may improve performance for downstream drug discovery research tasks. Code available at: https://github.com/GenerativeDrugDiscovery/xImagand-DKI

q-bio.QM