SearcharxivSearch

arXiv subjects

Samir Darouich

Publications and source records attributed to Samir Darouich.

7 recordsLinked to original sources

Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity

Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships. Despite the known importance of 3D structure to determine chemical and physical properties, the most widely used molecular fingerprints encode only two-dimensional connectivity. Such representations fail to distinguish similar but distinct stereoisomers and conformers. Alternative 3D methods are typically defined pairwise, making their application to large chemical spaces prohibitive, while deep learning embeddings are expressive but uninterpretable and limited by their training data diversity. Here, we introduce novel physics-inspired molecular fingerprints based on principles from spectral graph theory. We represent molecules as a complete graph in 3D space, with edge weights encoding heuristic physical interactions. Eigenvalue decomposition of the resulting graph Laplacian matrix results in a computationally efficient fixed-length chemical fingerprint that encodes 3D structure while obeying necessary physical symmetries of permutation and E(3) invariance. Spectral fingerprints differentiate between unique molecular structures with identical 2D connectivity, overcoming a limitation of 2D descriptors, while maintaining the low computational cost needed for efficient screening of vast chemical spaces. We evaluate our fingerprints with community detection algorithms and observe strong performance against representative baselines across datasets from organic, inorganic, biological, reticular, and reaction chemistry. Nearest-neighbor property estimation and applicability domain analyses reveal the utility of our molecular representation in machine learning and cheminformatics. We anticipate that spectral fingerprints will serve as generalizable, interpretable, and efficient measures of chemical similarity that incorporate 3D information at minimal cost.

physics.chem-ph

ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification Across the Periodic Table

Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species. This work introduces ElemeNet, a unified, general-purpose software package for molecular machine learning. The ElemeNet software package enables the training of advanced ML models for diverse properties and datasets with an enlarged range of elemental compositions. We define molecular representations compatible with elements 1-100, supporting diverse organometallic and biological systems in addition to organic chemistry already well-served by the Chemprop ML toolkit. As well as more common atom-, bond-, and molecule-level predictions, we introduce moiety predictions. We also natively define optional conditioning on charge and spin states. Advanced E(3)-equivariant and transformer architectures are supported, as well as classical 2D models, with all classes including built-in uncertainty quantification through deterministic and statistical measures. We benchmark our protocols for ML model training against representative datasets from organic, inorganic, coordination, and biological chemistry, achieving competitive and SOTA performance relative to literature baselines and favorable scaling to millions of molecules. The entire workflow is exposed through a concise command-line interface, lowering the barrier to entry for non-expert users. We anticipate ElemeNet will empower non-computational researchers to leverage modern deep learning methods across the chemical and physical sciences.

physics.chem-ph

ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation

Distilling conditional diffusion models aims to transfer the behavior of a large teacher to a smaller student while preserving alignment across conditioning inputs. Unlike recognition tasks, knowledge distillation in conditional diffusion often struggles to transfer knowledge beyond the training distribution, since the predicted noise strongly depends on the conditioning signal. As a result, effective distillation requires exploring a large conditioning space. In practical settings, this creates a major bottleneck. Paired image-condition data may be limited, and generating synthetic images for every available condition is often computationally infeasible, while the pool of conditions, such as text prompts, can be extremely large. Recent work addresses this issue by switching conditions during training, exposing the student to a broader conditioning space without changing the distillation objective. Yet this raises a complementary question: once a large conditioning corpus is available, how should the training effort be allocated? In this work, we introduce ARIA, a framework that adaptively allocates training effort across coarse regions of the conditioning space. By maintaining online estimates of teacher-student discrepancy at the region level, ARIA focuses updates where misalignment persists while preserving the original distillation objective. Empirically, ARIA improves over RC across most architectures and settings, with the clearest gains observed in unseen and underrepresented regimes. We also provide a theoretical analysis showing that the proposed tracking mechanism follows the evolving discrepancy during training under bounded variance and drift assumptions.

cs.LG

SymDrift: One-Shot Generative Modeling under Symmetries

Generative modeling of physical systems, such as molecules, requires learning distributions that are invariant under global symmetries, such as rotations in three-dimensional space. Equivariant diffusion and flow matching models can incorporate such invariances effectively, even when trained on a non-invariant empirical distribution, but they typically rely on costly multi-step sampling. Recently, drifting models have emerged as an efficient alternative, enabling single-step generation and achieving state-of-the-art performance in generative modeling tasks. However, we show that drifting models face a symmetry-specific challenge, since an equivariant generator does not generally produce the same drifting field as the one obtained from the symmetrized target distribution. Addressing this issue would require expensive symmetrization of the empirical distribution. To avoid this cost, we propose SymDrift, a framework that makes the drifting field itself symmetry-aware. We introduce two complementary strategies: (i) a symmetrized drift in coordinate space based on optimal alignment, and (ii) a $G$-invariant embedding that removes symmetry ambiguity by construction. Empirically, SymDrift outperforms existing one-shot methods on standard benchmarks for conformer and transition state generation, while remaining competitive with significantly more expensive multi-step approaches. By enabling one-shot inference, SymDrift reduces computational overhead by up to 40$\times$ compared to existing baselines, making it promising for high-throughput applications such as virtual drug screening and large-scale reaction network exploration.

cs.LG

Beyond the Training Domain: Robust Generative Transition State Models for Unseen Chemistry

Transition states (TSs) govern the rates and outcomes of chemical reactions, making their accurate prediction a central challenge in computational chemistry. Although recent machine-learning models achieve near chemical accuracy in the prediction of TS structures and the associated reaction barriers for small organic reactions, their ability to generalize beyond the training domain remains largely unexplored. Here, we introduce targeted benchmarks to probe chemical and structural novelty in generative TS prediction. Building on Transition1x, a large-scale dataset of reactions involving small organic molecules, we construct curated extensions incorporating controlled elemental substitutions and diverse transition-metal complexes (TMC). These benchmarks reveal fundamental limitations of generative models in the generalization to previously unseen elements. As a result, they produce unphysical geometries and large energetic errors, even for reactions structurally similar to well-predicted organic systems. To address this challenge, we introduce a self-supervised pretraining strategy based on equilibrium conformers that exposes generative TS models to novel chemical environments prior to targeted fine-tuning. Across the newly proposed benchmarks, self-supervised pretraining substantially improves TS prediction for previously unseen systems, lowering the median RMSD of TS geometries on T1x-TMC reactions from 0.39 to 0.19 $\mathring{A}$ and reducing fine-tuning data requirements by up to 75%, enabling reliable performance even in low-data regimes. Overall, the integration of generative TS models with self-supervised pseudo-reaction pretraining provides an efficient, scalable, and chemically robust framework for elucidating TSs well beyond the small organic molecule domain, establishing a foundation for investigating complex and catalytically relevant reaction landscapes.

physics.chem-ph

INTERFACE Force Field for Alumina with Validated Bulk Phases and a pH-Resolved Surface Model Database for Electrolyte and Organic Interfaces

Alumina and aluminum oxyhydroxides underpin chemical-engineering technologies from heterogeneous catalysis, corrosion protection, functional coatings, energy-storage devices, to biomedical components. Yet molecular models that predictively connect phase structure, pH-dependent surface chemistry, electrolyte organization, and adsorption across operating conditions remain limited. Here we introduce a unified INTERFACE Force Field (IFF) parameterization together with a curated, ready-to-use pH-resolved surface model database that provides the most accurate and transferable atomistic description of major alumina phases to date. The framework covers a-Al2O3, g-Al2O3, boehmite, diaspore, and gibbsite using a single, physically interpretable parameter set that is directly compatible with CHARMM, AMBER, OPLS-AA, CVFF, and PCFF. Across structural, thermodynamic, mechanical, and interfacial benchmarks, simulations reproduce experimental reference data with more than 95 percent accuracy, exceeding existing force fields and the reliability of current density-functional approaches. A key advance is the first transferable treatment of surface ionization and charge regulation across alumina phases over a broad range of pH values, enabling simulations of realistic solid electrolyte interfaces without phase-specific reparameterization. Quantitative reliability is demonstrated by reproducing trends in zeta potentials and pH-dependent adsorption of a corrosion inhibitor at alumina-water interfaces. Predicted adsorption free energies and surface contact times correlate with experiments across more than an order of magnitude. Relative to ML-DFT workflows, the speed 100 to 1000 times faster, reaching system sizes and time scales inaccessible to quantum methods. The results establish a predictive computational platform to design alumina-containing functional materials under realistic process conditions.

cond-mat.mtrl-sci

Adaptive Transition State Refinement with Learned Equilibrium Flows

Identifying transition states (TSs), the high-energy configurations that molecules pass through during chemical reactions, is essential for understanding and designing chemical processes. However, accurately and efficiently identifying these states remains one of the most challenging problems in computational chemistry. In this work, we introduce a new generative AI approach that improves the quality of initial guesses for TS structures. Our method can be combined with a variety of existing techniques, including both machine learning models and fast, approximate quantum methods, to refine their predictions and bring them closer to chemically accurate results. Applied to TS guesses from a state-of-the-art machine learning model, our approach reduces the median structural error to just 0.088 $\unicode{x212B}$ and lowers the median absolute error in reaction barrier heights to 0.79 kcal mol$^{-1}$. When starting from a widely used tight-binding approximation, it increases the success rate of locating valid TSs by 41\% and speeds up high-level quantum optimization by a factor of three. By making TS searches more accurate, robust, and efficient, this method could accelerate reaction mechanism discovery and support the development of new materials, catalysts, and pharmaceuticals.

physics.chem-ph