SearcharxivSearch

arXiv subjects

Arthur Y. Lin

Publications and source records attributed to Arthur Y. Lin.

4 recordsLinked to original sources

A Generalized Approach for Incorporating Geometry and Directionality into Coarse-Grained Machine-Learned Potentials

Machine-learned interatomic potentials have enabled highly accurate atomistic simulations, but extending these capabilities to coarse-grained systems remains challenging due to the loss of geometric and orientational information during coarse-graining. In this work, we present a generalized framework for incorporating molecular geometry and directionality into coarse-grained machine-learned potentials through two complementary approaches: anisotropic density-based descriptors (AniSOAP) and symmetry-adapted equivariant message-passing neural networks (MACE-CG). Using Gay-Berne particles and coarse-grained representations of benzene, formamide, and water, we demonstrate that explicitly retaining molecular anisotropy substantially improves the prediction of energies, forces, and torques relative to isotropic representations. AniSOAP provides an effective linear baseline when molecular shape is well approximated by ellipsoidal symmetry, while symmetry-adapted MACE-CG enables the incorporation of arbitrary molecular point-group symmetries. For water, whose orientational degrees of freedom are poorly represented by ellipsoidal descriptors alone, symmetry-adapted rigid-body features improve energy, force, and torque prediction by resolving orientational degeneracies inherent to isotropic and moment-of-inertia-based representations. These results show that information loss in coarse-grained modeling is governed not only by mapping resolution but also by the symmetry and geometric information retained in the representation, providing a systematic route toward more expressive and transferable coarse-grained machine-learned potentials.

physics.chem-ph

Extrapolation of Machine-Learning Interatomic Potentials for Organic and Polymeric Systems

Machine-Learning Interatomic Potentials (MLIPs) have surged in popularity due to their promise of expanding the spatiotemporal scales possible for simulating molecules with high fidelity. The accuracy of any MLIP is dependent on the data used for its training; thus, for large molecules, like polymers, where accurate training data is prohibitively difficult to obtain, it becomes necessary to pursue non-traditional methods to construct MLIPs, many of which are based on constructing MLIPs using smaller, analogous chemical systems. However, we have yet to understand the limits to which smaller molecules can be used as a proxy for extrapolating macromolecular energetics. Here, we provide a ``control study'' for such experiments, exploring the ability of MLIP approaches to extrapolate between n=1-8 n-polyalkanes at identical conditions. Through Principal Covariates Classification, we quantitatively demonstrate how convergence in chemical environments between training and testing datasets coincides with an MLIP's transferability. Additionally, we show how careful attention to the construction of an MLIP's neighbor list can promote greater transferability when considering various levels of the energetic hierarchy. Our results establish a roadmap for how one can create transferable MLIPs for macromolecular systems without the prohibitive cost of constructing system-specific training data.

cond-mat.soft

Interpretable Visualizations of Data Spaces for Classification Problems

How do classification models "see" our data? Based on their success in delineating behaviors, there must be some lens through which it is easy to see the boundary between classes; however, our current set of visualization techniques makes this prospect difficult. In this work, we propose a hybrid supervised-unsupervised technique distinctly suited to visualizing the decision boundaries determined by classification problems. This method provides a human-interpretable map that can be analyzed qualitatively and quantitatively, which we demonstrate through visualizing and interpreting a decision boundary for chemical neurotoxicity. While we discuss this method in the context of chemistry-driven problems, its application can be generalized across subfields for "unboxing" the operations of machine-learning classification models.

cs.LG

Expanding Density-Correlation Machine Learning Representations for Anisotropic Coarse-Grained Particles

Physics-based, atom-centered machine learning (ML) representations have been instrumental to the effective integration of ML within the atomistic simulation community. Many of these representations build off the idea of atoms as having spherical, or isotropic, interactions. In many communities, there is often a need to represent groups of atoms, either to increase the computational efficiency of simulation via coarse-graining or to understand molecular influences on system behavior. In such cases, atom-centered representations will have limited utility, as groups of atoms may not be well-approximated as spheres. In this work, we extend the popular Smooth Overlap of Atomic Positions (SOAP) ML representation for systems consisting of non-spherical anisotropic particles or clusters of atoms. We show the power of this anisotropic extension of SOAP, which we deem \AniSOAP, in accurately characterizing liquid crystal systems and predicting the energetics of Gay-Berne ellipsoids and coarse-grained benzene crystals. With our study of these prototypical anisotropic systems, we derive fundamental insights into how molecular shape influences mesoscale behavior and explain how to reincorporate important atom-atom interactions typically not captured by coarse-grained models. Moving forward, we propose \AniSOAP as a flexible, unified framework for coarse-graining in complex, multiscale simulation.

physics.comp-ph