SearcharxivSearch

arXiv subjects

Rose K. Cersonsky

Publications and source records attributed to Rose K. Cersonsky.

10 recordsLinked to original sources

A Generalized Approach for Incorporating Geometry and Directionality into Coarse-Grained Machine-Learned Potentials

Machine-learned interatomic potentials have enabled highly accurate atomistic simulations, but extending these capabilities to coarse-grained systems remains challenging due to the loss of geometric and orientational information during coarse-graining. In this work, we present a generalized framework for incorporating molecular geometry and directionality into coarse-grained machine-learned potentials through two complementary approaches: anisotropic density-based descriptors (AniSOAP) and symmetry-adapted equivariant message-passing neural networks (MACE-CG). Using Gay-Berne particles and coarse-grained representations of benzene, formamide, and water, we demonstrate that explicitly retaining molecular anisotropy substantially improves the prediction of energies, forces, and torques relative to isotropic representations. AniSOAP provides an effective linear baseline when molecular shape is well approximated by ellipsoidal symmetry, while symmetry-adapted MACE-CG enables the incorporation of arbitrary molecular point-group symmetries. For water, whose orientational degrees of freedom are poorly represented by ellipsoidal descriptors alone, symmetry-adapted rigid-body features improve energy, force, and torque prediction by resolving orientational degeneracies inherent to isotropic and moment-of-inertia-based representations. These results show that information loss in coarse-grained modeling is governed not only by mapping resolution but also by the symmetry and geometric information retained in the representation, providing a systematic route toward more expressive and transferable coarse-grained machine-learned potentials.

physics.chem-ph

Interpretable Visualizations of Data Spaces for Classification Problems

How do classification models "see" our data? Based on their success in delineating behaviors, there must be some lens through which it is easy to see the boundary between classes; however, our current set of visualization techniques makes this prospect difficult. In this work, we propose a hybrid supervised-unsupervised technique distinctly suited to visualizing the decision boundaries determined by classification problems. This method provides a human-interpretable map that can be analyzed qualitatively and quantitatively, which we demonstrate through visualizing and interpreting a decision boundary for chemical neurotoxicity. While we discuss this method in the context of chemistry-driven problems, its application can be generalized across subfields for "unboxing" the operations of machine-learning classification models.

cs.LG

Extrapolation of Machine-Learning Interatomic Potentials for Organic and Polymeric Systems

Machine-Learning Interatomic Potentials (MLIPs) have surged in popularity due to their promise of expanding the spatiotemporal scales possible for simulating molecules with high fidelity. The accuracy of any MLIP is dependent on the data used for its training; thus, for large molecules, like polymers, where accurate training data is prohibitively difficult to obtain, it becomes necessary to pursue non-traditional methods to construct MLIPs, many of which are based on constructing MLIPs using smaller, analogous chemical systems. However, we have yet to understand the limits to which smaller molecules can be used as a proxy for extrapolating macromolecular energetics. Here, we provide a ``control study'' for such experiments, exploring the ability of MLIP approaches to extrapolate between n=1-8 n-polyalkanes at identical conditions. Through Principal Covariates Classification, we quantitatively demonstrate how convergence in chemical environments between training and testing datasets coincides with an MLIP's transferability. Additionally, we show how careful attention to the construction of an MLIP's neighbor list can promote greater transferability when considering various levels of the energetic hierarchy. Our results establish a roadmap for how one can create transferable MLIPs for macromolecular systems without the prohibitive cost of constructing system-specific training data.

cond-mat.soft

Expanding Density-Correlation Machine Learning Representations for Anisotropic Coarse-Grained Particles

Physics-based, atom-centered machine learning (ML) representations have been instrumental to the effective integration of ML within the atomistic simulation community. Many of these representations build off the idea of atoms as having spherical, or isotropic, interactions. In many communities, there is often a need to represent groups of atoms, either to increase the computational efficiency of simulation via coarse-graining or to understand molecular influences on system behavior. In such cases, atom-centered representations will have limited utility, as groups of atoms may not be well-approximated as spheres. In this work, we extend the popular Smooth Overlap of Atomic Positions (SOAP) ML representation for systems consisting of non-spherical anisotropic particles or clusters of atoms. We show the power of this anisotropic extension of SOAP, which we deem \AniSOAP, in accurately characterizing liquid crystal systems and predicting the energetics of Gay-Berne ellipsoids and coarse-grained benzene crystals. With our study of these prototypical anisotropic systems, we derive fundamental insights into how molecular shape influences mesoscale behavior and explain how to reincorporate important atom-atom interactions typically not captured by coarse-grained models. Moving forward, we propose \AniSOAP as a flexible, unified framework for coarse-graining in complex, multiscale simulation.

physics.comp-ph

The rule of four: anomalous stoichiometries of inorganic compounds

Why are materials with specific characteristics more abundant than others? This is a fundamental question in materials science and one that is traditionally difficult to tackle, given the vastness of compositional and configurational space. We highlight here the anomalous abundance of inorganic compounds whose primitive unit cell contains a number of atoms that is a multiple of four. This occurrence - named here the 'rule of four' - has to our knowledge not previously been reported or studied. Here, we first highlight the rule's existence, especially notable when restricting oneself to experimentally known compounds, and explore its possible relationship with established descriptors of crystal structures, from symmetries to energies. We then investigate this relative abundance by looking at structural descriptors, both of global (packing configurations) and local (the smooth overlap of atomic positions) nature. Contrary to intuition, the overabundance does not correlate with low-energy or high-symmetry structures; in fact, structures which obey the 'rule of four' are characterized by low symmetries and loosely packed arrangements maximizing the free volume. We are able to correlate this abundance with local structural symmetries, and visualize the results using a hybrid supervised-unsupervised machine learning method.

cond-mat.mtrl-sci

A data-driven interpretation of the stability of molecular crystals

Due to the subtle balance of intermolecular interactions that govern structure-property relations, predicting the stability of crystal structures formed from molecular building blocks is a highly non-trivial scientific problem. A particularly active and fruitful approach involves classifying the different combinations of interacting chemical moieties, as understanding the relative energetics of different interactions enables the design of molecular crystals and fine-tuning their stabilities. While this is usually performed based on the empirical observation of the most commonly encountered motifs in known crystal structures, we propose to apply a combination of supervised and unsupervised machine-learning techniques to automate the construction of an extensive library of molecular building blocks. We introduce a structural descriptor tailored to the prediction of the binding (lattice) energy and apply it to a curated dataset of organic crystals and exploit its atom-centered nature to obtain a data-driven assessment of the contribution of different chemical groups to the lattice energy of the crystal. We then interpret this library using a low-dimensional representation of the structure-energy landscape and discuss selected examples of the insights into crystal engineering that can be extracted from this analysis, providing a complete database to guide the design of molecular materials.

physics.chem-ph

A Route to Hierarchical Assembly of Colloidal Diamond

Photonic crystals, appealing for their ability to control light, are constructed by periodic regions of different dielectric constants. Yet, the structural holy grail in photonic materials, diamond, remains challenging to synthesize at the colloidal length scale. Here we explore new ways to assemble diamond using modified gyrobifastigial (mGBF) nanoparticles, a shape that resembles two antialigned triangular prisms. We investigate the parameter space that leads to the self-assembly of diamond, and we compare the likelihood of defects in diamond self-assembled via mGBF vs. the nanoparticle shape that is the current focus for assembling diamond, the truncated tetrahedra. We introduce a potential route for realizing mGBF particles by dimerizing triangular prisms using attractive patches, and we report the impact of this superstructure on the photonic properties.

cond-mat.mtrl-sci

Improving Sample and Feature Selection with Principal Covariates Regression

Selecting the most relevant features and samples out of a large set of candidates is a task that occurs very often in the context of automated data analysis, where it can be used to improve the computational performance, and also often the transferability, of a model. Here we focus on two popular sub-selection schemes which have been applied to this end: CUR decomposition, that is based on a low-rank approximation of the feature matrix and Farthest Point Sampling, that relies on the iterative identification of the most diverse samples and discriminating features. We modify these unsupervised approaches, incorporating a supervised component following the same spirit as the Principal Covariates Regression (PCovR) method. We show that incorporating target information provides selections that perform better in supervised tasks, which we demonstrate with ridge regression, kernel ridge regression, and sparse kernel regression. We also show that incorporating aspects of simple supervised learning models can improve the accuracy of more complex models, such as feed-forward neural networks. We present adjustments to minimize the impact that any subselection may incur when performing unsupervised tasks. We demonstrate the significant improvements associated with the use of PCov-CUR and PCov-FPS selections for applications to chemistry and materials science, typically reducing by a factor of two the number of features and samples which are required to achieve a given level of regression accuracy.

physics.chem-ph

Structure-Property Maps with Kernel Principal Covariates Regression

Data analyses based on linear methods constitute the simplest, most robust, and transparent approaches to the automatic processing of large amounts of data for building supervised or unsupervised machine learning models. Principal covariates regression (PCovR) is an underappreciated method that interpolates between principal component analysis and linear regression, and can be used to conveniently reveal structure-property relations in terms of simple-to-interpret, low-dimensional maps. Here we provide a pedagogic overview of these data analysis schemes, including the use of the kernel trick to introduce an element of non-linearity, while maintaining most of the convenience and the simplicity of linear approaches. We then introduce a kernelized version of PCovR and a sparsified extension, and demonstrate the performance of this approach in revealing and predicting structure-property relations in chemistry and materials science, showing a variety of examples including elemental carbon, porous silicate frameworks, organic molecules, amino acid conformers, and molecular materials.

stat.ML

Pressure-Tunable Photonic Band Gaps in an Entropic Colloidal Crystal

Materials adopting the diamond structure possess useful properties in atomic and colloidal systems, and are a popular target for synthesis in colloids where a photonic band gap is possible. The desirable photonic properties of the diamond structure pose an interesting opportunity for reconfigurable matter: can we create a colloidal crystal able to switch reversibly to and from the diamond structure? Drawing inspiration from high-pressure transitions of diamond-forming atomic systems, we design a system of polyhedrally-shaped particles that transitions from diamond to a tetragonal diamond derivative upon a small pressure change. The transition can alternatively be triggered by changing the shape of the particle in-situ. We propose that the transition provides a reversible reconfiguration process for a potential new colloidal material, and draw parallels between this transition and phase behavior of the atomic transitions from which we take inspiration.

cond-mat.mtrl-sci