SearcharxivSearch

arXiv subjects

Chiho Im

Publications and source records attributed to Chiho Im.

2 recordsLinked to original sources

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries without explicitly reasoning about which residues or interactions are functionally essential. As a result, design decisions are entangled with continuous sampling dynamics, limiting interpretability, controllability, and systematic reuse of biochemical knowledge. We introduce Proteo-R1, a reasoning-guided protein design framework that explicitly decouples molecular understanding from geometric generation. Proteo-R1 adopts a dual-expert architecture in which a multimodal large language model (MLLM) serves as an understanding expert, analyzing protein sequences, structures, and textual context to identify key functional residues that govern binding and specificity. These residue-level decisions are then passed as hard constraints to a separate diffusion-based generation expert, which performs conditional co-design while respecting the fixed interaction anchors. This factorization mirrors how human experts approach molecular engineering: first, reasoning about critical interactions, then optimizing geometry subject to those constraints. By operationalizing reasoning as explicit residue-level commitments rather than latent textual guidance, Proteo-R1 achieves stable, interpretable, and modular integration of LLM reasoning with state-of-the-art geometric generative models. Code, data, and demos are available at https://smiles724.github.io/r1/.

cs.LG

A Graph Completion Method that Jointly Predicts Geometry and Topology Enables Effective Molecule Assembly

A common starting point for drug design is to find small chemical groups or "fragments" that form interactions with distinct subregions in a protein binding pocket. The subsequent challenge is to assemble these fragments into a molecule that has high affinity to the protein, by adding chemical bonds between atoms in different fragments. This "molecule assembly" task is particularly challenging because, initially, fragment positions are known only approximately. Prior methods for spatial graph completion-adding missing edges to a graph whose nodes have associated spatial coordinates-either treat node positions as fixed or adjust node positions before predicting edges. The fact that these methods treat geometry and topology prediction separately limits their ability to reconcile noisy geometries and plausible connectivities. To address this limitation, we introduce EdGr, a spatial graph diffusion model that reasons jointly over geometry and topology of molecules to simultaneously predict fragment positions and inter-fragment bonds. Importantly, predicted edge likelihoods directly influence node position updates during the diffusion denoising process, allowing connectivity cues to guide spatial movements, and vice versa. EdGr substantially outperforms previous methods on the molecule assembly task and maintains robust performance as noise levels increase. Beyond drug discovery, our approach of explicitly coupling geometry and topology prediction is broadly applicable to spatial graph completion problems, such as neural circuit reconstruction, 3D scene understanding, and sensor network design.

q-bio.QM