SearcharxivSearch

arXiv subjects

Dominic Mashak

Publications and source records attributed to Dominic Mashak.

4 recordsLinked to original sources

CVT Archives and Chemical Embedding Measures for Multi-Objective Quality Diversity in Molecular Design

Nonlinear optical (NLO) materials are essential for photonic technologies, yet discovering optimal NLO molecules requires balancing multiple competing objectives across vast chemical spaces. Previous work showed that Multi-Objective MAP-Elites (MOME) with grid-based archives discovers diverse, high-quality molecules for electro-optic applications. However, uniform grid partitioning wastes archive capacity on chemically infeasible regions while undersampling high-density areas. We apply MOME with Centroidal Voronoi Tessellation (CVT) archives whose cells are defined by learned embeddings from ChemBERTa-2 Multi-Task Regression reduced via UMAP, capturing chemical similarity beyond simple structural features. We investigate a four-objective NLO molecular design problem: maximizing the $\beta / \gamma$ hyperpolarizability ratio, constraining HOMO-LUMO gap and linear polarizability to target ranges, and minimizing energy per atom. Our results demonstrate that embedding-based measures in CVT archives yield significantly higher median global hypervolume and multi-objective quality diversity scores, while filling nearly all native archive niches.

physics.comp-ph

Multi-Objective Evolutionary Design of Molecules with Enhanced Nonlinear Optical Properties

Nonlinear optical (NLO) materials are essential for many photonic, telecommunication, and laser technologies, yet discovering better NLO molecules is computationally challenging due to the vast chemical space and competing objectives. We compare evolutionary algorithms for molecular design, targeting four objectives: maximizing the ratio of first-to-second hyperpolarizability $(\beta/\gamma)$, optimizing HOMO-LUMO gap and linear polarizability to target ranges, and minimizing energy per atom. We encode molecules as SMILES strings and evaluate their properties using quantum-chemical calculations. We compare NSGA-II, MAP-Elites, MOME, a single-objective $(\mu+\lambda)$ evolutionary algorithm, and simulated annealing. Quality diversity methods maintain archives across a measure space defined by atom and bond count, enabling the discovery of structurally diverse molecules. Our results demonstrate that NSGA-II consistently earns high scores in every objective, leading to high-quality molecules, but MOME does a better job exploring a wide range of possibilities, resulting in higher global hypervolume and MOQD scores. However, each method has strengths and weaknesses, and produced many promising molecules.

physics.comp-ph

Finding Molecules with Specific Properties: Simulated Annealing vs. Evolution

We compare the ability of a simulated annealing program and an evolutionary algorithm to find molecules with large molecular average hyperpolarizabilities. This property is an important component of nonlinear optical materials. Both optimization programs represent molecules as SMILES strings, a method that is widely used by chemists to describe molecular structure using short ASCII strings. Our results suggest that both approaches are comparable and can be used to solve a variety of more realistic problems of interest to chemists and material scientists.

physics.comp-ph

Benchmarking Hartree-Fock and DFT for Molecular Hyperpolarizability: Implications for Evolutionary Design

Evolutionary algorithms for molecular design require computationally efficient yet accurate fitness functions. We systematically benchmark Hartree-Fock and density functional theory for predicting molecular first hyperpolarizability ($\beta$), evaluating five functionals (HF, PBE0, B3LYP, CAM-B3LYP, M06-2X) across six basis sets against experimental data from five organic push-pull chromophores. For this dataset, HF/3-21G achieves 45.5% mean absolute percentage error with perfect pairwise ranking in 7.4 minutes per molecule. All 30 tested combinations of functional and basis sets maintain perfect pairwise agreement, validating their use as evolutionary fitness functions despite moderate absolute errors. Larger basis sets yield a lower percentage error compared to the experimental values than the difference with the functional. The preservation of pairwise rankings across all combinations of functionals and basis sets provides crucial guidance for evolutionary optimization of nonlinear optical materials.

physics.chem-ph