SearcharxivSearch

arXiv subjects

Andrea Anelli

Publications and source records attributed to Andrea Anelli.

8 recordsLinked to original sources

Dynamic language model representations for multi-objective reaction optimisation

Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of high-throughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.

cs.LG

APPA : Agentic Preformulation Pathway Assistant

The design and development of effective drug formulations is a critical process in pharmaceutical research, particularly for small molecule active pharmaceutical ingredients. This paper introduces a novel agentic preformulation pathway assistant (Appa), leveraging large language models coupled to experimental databases and a suite of machine learning models to streamline the preformulation process of drug candidates. Appa successfully integrates domain expertise from scientific publications, databases holding experimental results, and machine learning predictors to reason and propose optimal preformulation strategies based on the current evidence. This results in case-specific user guidance for the developability assessment of a new drug and directs towards the most promising experimental route, significantly reducing the time and resources required for the manual collection and analysis of existing evidence. The approach aims to accelerate the transition of promising compounds from discovery to preclinical and clinical testing.

physics.chem-ph

Automatic solid form classification in pharmaceutical drug development

In materials and pharmaceutical development, rapidly and accurately determining the similarity between X-ray powder diffraction (XRPD) measurements is crucial for efficient solid form screening and analysis. We present SMolNet, a classifier based on a Siamese network architecture, designed to automate the comparison of XRPD patterns. Our results show that training SMolNet on loss functions from the self-supervised learning domain yields a substantial boost in performance with respect to class separability and precision, specifically when classifying phases of previously unseen compounds. The application of SMolNet demonstrates significant improvements in screening efficiency across multiple active pharmaceutical ingredients, providing a powerful tool for scientists to discover and categorize measurements with reliable accuracy.

physics.chem-ph

Exploring the robust extrapolation of high-dimensional machine learning potentials

We show that, contrary to popular assumptions, predictions from machine learning potentials built upon high-dimensional atom-density representations almost exclusively occur in regions of the representation space which lie outside the convex hull defined by the training set points. We then propose a perspective to rationalize the domain of robust extrapolation and accurate prediction of atomistic machine learning potentials in terms of the probability density induced by training points in the representation space

physics.comp-ph

Learning the electronic density of states in condensed matter

The electronic density of states (DOS) quantifies the distribution of the energy levels that can be occupied by electrons in a quasiparticle picture, and is central to modern electronic structure theory. It also underpins the computation and interpretation of experimentally observable material properties such as optical absorption and electrical conductivity. We discuss the challenges inherent in the construction of a machine-learning (ML) framework aimed at predicting the DOS as a combination of local contributions that depend in turn on the geometric configuration of neighbours around each atom, using quasiparticle energy levels from density functional theory as training data. We present a challenging case study that includes configurations of silicon spanning a broad set of thermodynamic conditions, ranging from bulk structures to clusters, and from semiconducting to metallic behavior. We compare different approaches to represent the DOS, and the accuracy of predicting quantities such as the Fermi level, the DOS at the Fermi level, or the band energy, either directly or as a side-product of the evaluation of the DOS. The performance of the model depends crucially on the smoothening of the DOS, and there is a tradeoff to be made between the systematic error associated with the smoothening and the error in the ML model for a specific structure. We demonstrate the usefulness of this approach by computing the density of states of a large amorphous silicon sample, for which it would be prohibitively expensive to compute the DOS by direct electronic structure calculations, and show how the atom-centred decomposition of the DOS that is obtained through our model can be used to extract physical insights into the connections between structural and electronic features.

cond-mat.mtrl-sci

A Bayesian approach to NMR crystal structure determination

Nuclear Magnetic Resonance (NMR) spectroscopy is particularly well-suited to determine the structure of molecules and materials in powdered form. Structure determination usually proceeds by finding the best match between experimentally observed NMR chemical shifts and those of candidate structures. Chemical shifts for the candidate configurations have traditionally been computed by electronic-structure methods, and more recently predicted by machine learning. However, the reliability of the determination depends on the errors in the predicted shifts. Here we propose a Bayesian framework for determining the confidence in the identification of the experimental crystal structure, based on knowledge of the typical error in the electronic structure methods. We also extend the recently-developed ShiftML machine-learning model, including the evaluation of the uncertainty of its predictions. We demonstrate the approach on the determination of the structures of six organic molecular crystals. We critically assess the reliability of the structure determinations, facilitated by the introduction of a visualization of the of similarity between candidate configurations in terms of their chemical shifts and their structures. We also show that the commonly used values for the errors in calculated $^{13}$C shifts are underestimated, and that more accurate, self-consistently determined uncertainties make it possible to use $^{13}$C shifts to improve the accuracy of structure determinations.

physics.chem-ph

Generalized convex hull construction for materials discovery

High-throughput computational materials searches generate large databases of locally-stable structures. Conventionally, the needle-in-a-haystack search for the few experimentally-synthesizable compounds is performed using a convex hull construction, which identifies structures stabilized by manipulation of a particular thermodynamic constraint (for example pressure or composition) chosen based on prior experimental evidence or intuition. To address the biased nature of this procedure we introduce a generalized convex hull framework. Convex hulls are constructed on data-driven principal coordinates, which represent the full structural diversity of the database. Their coupling to experimentally-realizable constraints hints at the conditions that are most likely to stabilize a given configuration. The probabilistic nature of our framework also addresses the uncertainty stemming from the use of approximate models during database construction, and eliminates redundant structures. The remaining small set of candidates that have a high probability of being synthesizable provide a much needed starting point for the determination of viable synthetic pathways.

cond-mat.mtrl-sci

Automatic Selection of Atomic Fingerprints and Reference Configurations for Machine-Learning Potentials

Machine learning of atomic-scale properties is revolutionizing molecular modelling, making it possible to evaluate inter-atomic potentials with first-principles accuracy, at a fraction of the costs. The accuracy, speed and reliability of machine-learning potentials, however, depends strongly on the way atomic configurations are represented, i.e. the choice of descriptors used as input for the machine learning method. The raw Cartesian coordinates are typically transformed in "fingerprints", or "symmetry functions", that are designed to encode, in addition to the structure, important properties of the potential-energy surface like its invariances with respect to rotation, translation and permutation of like atoms. Here we discuss automatic protocols to select a number of fingerprints out of a large pool of candidates, based on the correlations that are intrinsic to the training data. This procedure can greatly simplify the construction of neural network potentials that strike the best balance between accuracy and computational efficiency, and has the potential to accelerate by orders of magnitude the evaluation of Gaussian Approximation Potentials based on the Smooth Overlap of Atomic Positions kernel. We present applications to the construction of neural network potentials for water and for an Al-Mg-Si alloy, and to the prediction of the formation energies of small organic molecules using Gaussian process regression.

physics.comp-ph