SearcharxivSearch

arXiv subjects

Leroy Cronin

Publications and source records attributed to Leroy Cronin.

At least 19 recordsLinked to original sources

Elucidating the Size of Chemical Space with Assembly Theory

Chemical space is unimaginably vast with common heuristic estimates suggesting that there are ca. 10^60 'drug-like' molecules possible below a molecular mass of 500 Da. However, these estimates largely ignore the structural and synthetic complexity of the molecules enumerated. Here we present a first-principles estimate of the size of chemical space using the Assembly Theory, which quantifies the amount of causation required to form a molecule, captured in the assembly Index. This is a measurable molecular complexity measure derived from the minimum number of recursive bond-joining operations required to construct a molecular graph. Assembly Theory partitions chemical space into levels defined by Assembly Index, allowing bounds to be placed on its growth as molecular complexity increases. We show that chemical space (the accumulated Assembly Index level sets) grows at least super-exponentially, and at most, double-exponentially with respect to the Assembly Index. Using the GDB-13 database as a reference for growth-rate estimation, we model how chemical space expands under increasing complexity and contracts under structural constraints, including atom and bond types, number of rings, ring size, and chemical motifs. Under constraints comparable to standard drug-like estimates, including molecular mass below 500 Da, our analysis yields a chemical space of approximately 10117 molecules at Assembly Index 25. Finally, we constrain chemical space by biologically relevant motifs and identify structurally relevant molecules near the accessible boundaries of these assembly-defined spaces.

physics.chem-ph

The Physics of Causation

Assembly theory (AT) introduces causation as a material property and establishes a metrology for objects produced by evolution and selection. The physical scale of causation is quantified by the assembly index, defined as the minimum number of recursive steps necessary to make an object. Observing countable copies of high assembly index objects indicates a mechanism producing them is persistent, such that the object's environment constructs a memory that traps causation within a contingent chain. Copy number and assembly index together underlie a standardized metrology for detecting causation (assembly index) and contingency (copy number). These allow a precise definition of an assembly threshold that demarcates life (and its derivative agential, intelligent, and technological forms and artifacts) as structures with persistent copies in regimes of deep causal possibility. In introducing a fundamental concept of material causation to quantify and measure life, AT represents a departure from prior theories of causation, such as interventional ones, which have so far proven incompatible with fundamental physics. We discuss how AT's concept of causation provides the foundation for a theory of physics that allows precise and testable concept of "life", and in which novelty, contingency and the potential for open-endedness are fundamental, and determinism is emergent from selection along assembled lineages.

physics.hist-ph

Searching for Life-As-We-Don't-Know-It: Mission-relevant Application of Assembly Theory for Exoplanet Life Detection

This white paper introduces a framework for applying Assembly Theory (AT) to planetary atmospheres as a biosignature framework suitable for the Habitable Worlds Observatory (HWO). AT quantifies the minimum combinatorial complexity required to co-construct an observed ensemble of molecular species, providing a measure of how much selection and evolution is encoded in a planetary atmosphere's chemical space, without assuming any specific biochemistry, kinetics nor metabolism. We outline some forthcoming results applying this framework and how it can be extended to population-level exoplanet studies, validated against existing spectroscopic data, and used to directly inform HWO instrumental requirements. Rather than imposing a binary alive/dead classification, AT-based atmospheric analysis would provide a continuous measure of planetary complexity, opening a path toward detecting life-as-we-don't-know-it.

astro-ph.IM

Quantifying the Emergence of Selection Prior to Biological Evolution

Selection is central to biological evolution, yet there has been no general experimental framework for quantifying selection in chemical systems before life. Here we demonstrate that selection in a prebiological chemical system can be directly quantified. Assembly Theory predicts that selection corresponds to a transition from undirected to directed exploration of chemical possibility space, measurable through the amount of Assembly, A, which integrates molecular assembly index with observed copy number. By analysing peptide ensembles produced under diverse polymerisation conditions, we show that undirected reactions explore sequence space almost uniformly, yielding exploration ratios of 0.85-0.95, whereas reactions influenced by evolved proteases generate markedly lower ratios (0.51-0.75) and elevated A, consistent with selective reinforcement of specific assembly pathways. Across multiple environments and amino-acid combinations, the exploration ratio and ensemble assembly A robustly distinguish directed from undirected exploration, establishing a general, experimentally tractable metric for detecting and measuring selection in chemical evolution.

q-bio.MN

Assembly Addition Chains

In this paper we extend the notion of Addition Chains over Z+ to a general set S. We explain how the algebraic structure of Assembly Multi-Magma over the pairs (S,BB proper subset of S) allows to define the concept of Addition Chain over S, called Assembly Addition Chains of S with Building Blocks BB. Analogously to the Z+ case, we introduce the concept of Optimal Assembly Addition Chains over S and prove lower and upper bounds for their lengths, similar to the bounds found by Schonhage for the Z+ case. In the general case the unit 1 is in set Z+ is replaced by the subset BB and the mentioned bounds for the length of an Optimal Assembly Addition Chain of O is in set S are defined in terms of the size of O (i.e. the number of Building Blocks required to construct O). The main examples of S that we consider through this papers are (i) j-Strings (Strings with an alphabeth of j letters), (ii) Colored Connected Graphs and (iii) Colored Polyominoes.

math.CO

Rapid Exploration of Assembly Chemical Space of Molecular Graphs

Quantifying how hard it is to build a molecular graph matters for biosignature detection, chemical complexity, and cheminformatics. We present an exact, scalable algorithm to compute the molecular assembly index (MA) which prioritizes the largest duplicate subgraphs, represents fragmentation with an 'assembly state' array of edge-lists, reuses states via hashing/DAGs, and prunes the search using a dynamic-programming branch-and-bound guided by a conditional-addition-chain lower bound. For organic molecules in the greater than 500 Da range our approach is up to six orders of magnitude faster than prior methods and yields exact MAs where previous algorithms would have timed out. We compute MAs to convergence for ~300k COCONUT natural products with <50 bonds, profiling time and memory scaling. Finally, we exploit the speed of our algorithm to calculate joint assembly spaces and introduce the Joint Assembly Overlap (JAO), a Jaccard-like metric that emphasizes global scaffold reuse and show that the JAO yields substantially different rankings from Tanimoto similarity with ECFP fingerprints and MCS (e.g. in steroids 270-380/Da and short peptides), accounting for substructural similarity beyond local environments. Together, these advances turn the molecular assembly index into a practical tool for large-scale exploration of chemical space.

cs.DS

Exploring molecular assembly as a biosignature using mass spectrometry and machine learning

Molecular assembly offers a promising path to detect life beyond Earth, while minimizing assumptions based on terrestrial life. As mass spectrometers will be central to upcoming Solar System missions, predicting molecular assembly from their data without needing to elucidate unknown structures will be essential for unbiased life detection. An ideal agnostic biosignature must be interpretable and experimentally measurable. Here, we show that molecular assembly, a recently developed approach to measure objects that have been produced by evolution, satisfies both criteria. First, it is interpretable for life detection, as it reflects the assembly of molecules with their bonds as building blocks, in contrast to approaches that discount construction history. Second, it can be determined without structural elucidation, as it can be physically measured by mass spectrometry, a property that distinguishes it from other approaches that use structure-based information measures for molecular complexity. Whilst molecular assembly is directly measurable using mass spectrometry data, there are limits imposed by mission constraints. To address this, we developed a machine learning model that predicts molecular assembly with high accuracy, reducing error by three-fold compared to baseline models. Simulated data shows that even small instrumental inconsistencies can double model error, emphasizing the need for standardization. These results suggest that standardized mass spectrometry databases could enable accurate molecular assembly prediction, without structural elucidation, providing a proof-of-concept for future astrobiology missions.

cs.LG

Chemputer and Chemputation -- A Universal Chemical Compound Synthesis Machine

Chemputation reframes synthesis as the programmable execution of reaction code on a universally re-configurable hardware graph. Here we prove that a chemputer equipped with a finite, but extensible, set of reagents, catalysts and process conditions, together with a chempiler that maps reaction graphs onto hardware, is universal: it can generate any stable, isolable molecule in finite time and in analytically detectable quantity, provided real-time error correction keeps the per-step fidelity above the threshold set by the molecule's assembly index. The proof is constructed by casting the platform as a Chemical Synthesis Turing Machine (CSTM). The CSTM formalism supplies (i) an eight-tuple state definition that unifies reagents, process variables (including catalysts) and tape operations; (ii) the Universal Chemputation Principle; and (iii) a dynamic-error-correction routine ensuring fault tolerant execution. Linking this framework to assembly theory strengthens the definition of a molecule by demanding practical synthesizability and error correction becomes a prerequisite for universality. We validate the abstraction against >100 \c{hi}DL programs executed on a modular chemputer rigs spanning single step to multi-step routes. Mapping each procedure onto CSTM shows that the cumulative number of unit operations grows linearly with synthetic depth. Together, these results elevate chemical synthesis to the status of a general computation: algorithms written in \c{hi}DL are compiled to hardware, executed with closed-loop correction, and produce verifiable molecular outputs. By formalising chemistry in this way, the chemputer offers a path to shareable, executable chemical code, interoperable hardware ecosystems, and ultimately a searchable, provable atlas of chemical space.

cs.ET

Constructing the Molecular Tree of Life using Assembly Theory and Mass Spectrometry

Here we demonstrate the first biochemistry-agnostic approach to map evolutionary relationships at the molecular scale, allowing the construction of phylogenetic models using mass spectrometry (MS) and Assembly Theory (AT) without elucidating molecular identities. AT allows us to estimate the complexity of molecules by deducing the amount of shared information stored within them when . By examining 74 samples from a diverse range of biotic and abiotic sources, we used tandem MS data to detect 24102 analytes (9262 unique) and 59518 molecular fragments (6755 unique). Using this MS dataset, together with AT, we were able to infer the joint assembly spaces (JAS) of samples from molecular analytes. We show how JAS allows agnostic annotation of samples without fingerprinting exact analyte identities, facilitating accurate determination of their biogenicity and taxonomical grouping. Furthermore, we developed an AT-based framework to construct a biochemistry-agnostic phylogenetic tree which is consistent with genome-based models and outperforms other similarity-based algorithms. Finally, we were able to use AT to track colony lineages of a single bacterial species based on phenotypic variation in their molecular composition with high accuracy, which would be challenging to track with genomic data. Our results demonstrate how AT can expand causal molecular inference to non-sequence information without requiring exact molecular identities, thereby opening the possibility to study previously inaccessible biological domains.

q-bio.PE

The Emergence of Chirality from Metabolism

Molecular chirality is critical to biochemical function, but it is unknown when chiral selectivity first became important in the evolutionary transition from geochemistry to biochemistry during the emergence of life. Here, we identify key transitions in the selection of chiral molecules in metabolic evolution, showing how achiral molecules (lacking chiral centers) may have given rise to specific and abundant chiral molecules in the elaboration of metabolic networks from geochemically available precursor molecules. Simulated expansions of biosphere-scale metabolism suggest new hypotheses about the evolution of chiral molecules within biochemistry, including a prominent role for both achiral and chiral compounds as nucleation sites of early metabolic network growth, an increasing enrichment of molecules with more chiral centers as these networks expand, and conservation of broken chiral symmetries along reaction pathways as a general organizing principle. We also find an unexpected enrichment in large, non-polymeric achiral molecules. Leveraging metabolic data of 40,023 genomes and metagenomes, we analyzed the statistics of chiral and achiral molecules in the large-scale organization of metabolism, revealing a chiral-enriched phase of network organization evidenced by system-size dependent chiral scaling laws that differ for individuals and ecosystems. By uncovering how metabolic networks could lead to chiral selection, our findings open new avenues for bridging metabolism and genetics-first approaches to the origin of chirality, allowing tools for better timing of major transitions in molecular organization during the emergence of life, understanding the role of chirality in extant and synthetic metabolisms, and informing targets for chirality-based biosignatures.

q-bio.MN

Quantifying the Complexity of Materials with Assembly Theory

Quantifying the evolution and complexity of materials is of importance in many areas of science and engineering, where a central open challenge is developing experimental complexity measurements to distinguish random structures from evolved or engineered materials. Assembly Theory (AT) was developed to measure complexity produced by selection, evolution and technology. Here, we extend the fundamentals of AT to quantify complexity in inorganic molecules and solid-state periodic objects such as crystals, minerals and microprocessors, showing how the framework of AT can be used to distinguish naturally formed materials from evolved and engineered ones by quantifying the amount of assembly using the assembly equation defined by AT. We show how tracking the Assembly of repeated structures within a material allows us formalizing the complexity of materials in a manner accessible to measurement. We confirm the physical relevance of our formal approach, by applying it to phase transformations in crystals using the HCP to FCC transformation as a model system. To explore this approach, we introduce random stacking faults in closed-packed systems simplified to one-dimensional strings and demonstrate how Assembly can track the phase transformation. We then compare the Assembly of closed-packed structures with random or engineered faults, demonstrating its utility in distinguishing engineered materials from randomly structured ones. Our results have implications for the study of pre-genetic minerals at the origin of life, optimization of material design in the trade-off between complexity and function, and new approaches to explore material technosignatures which can be unambiguously identified as products of engineered design.

cond-mat.mtrl-sci

Achieving Operational Universality through a Turing Complete Chemputer

The most fundamental abstraction underlying all modern computers is the Turing Machine, that is if any modern computer can simulate a Turing Machine, an equivalence which is called Turing completeness, it is theoretically possible to achieve any task that can be algorithmically described by executing a series of discrete unit operations. In chemistry, the ability to program chemical processes is demanding because it is hard to ensure that the process can be understood at a high level of abstraction, and then reduced to practice. Herein we exploit the concept of Turing completeness applied to robotic platforms for chemistry that can be used to synthesise complex molecules through unit operations that execute chemical processes using a chemically-aware programming language, XDL. We leverage the concept of computability by computers to synthesizability of chemical compounds by automated synthesis machines. The results of an interactive demonstration of Turing completeness using the colour gamut and conditional logic are presented and examples of chemical use-cases are discussed. Over 16.7 million combinations of Red, Green, Blue (RGB) colour space were binned into 5 discrete values and measured over 10 regions of interest (ROIs), affording 78 million possible states per step and served as a proxy for conceptual, chemical space exploration. This formal description establishes a formal framework in future chemical programming languages to ensure complex logic operations are expressed and executed correctly, with the possibility of error correction, in the automated and autonomous pursuit of increasingly complex molecules.

cs.CL

Assembly Theory and its Relationship with Computational Complexity

Assembly theory (AT) quantifies selection using the assembly equation and identifies complex objects that occur in abundance based on two measurements, assembly index and copy number, where the assembly index is the minimum number of joining operations necessary to construct an object from basic parts, and the copy number is how many instances of the given object(s) are observed. Together these define a quantity, called Assembly, which captures the amount of causation required to produce objects in abundance in an observed sample. This contrasts with the random generation of objects. Herein we describe how AT's focus on selection as the mechanism for generating complexity offers a distinct approach, and answers different questions, than computational complexity theory with its focus on minimum descriptions via compressibility. To explore formal differences between the two approaches, we show several simple and explicit mathematical examples demonstrating that the assembly index, itself only one piece of the theoretical framework of AT, is formally not equivalent to other commonly used complexity measures from computer science and information theory including Shannon entropy, Huffman encoding, and Lempel-Ziv-Welch compression. We also include proofs that assembly index is not in the same computational complexity class as these compression algorithms and discuss fundamental differences in the ontological basis of AT, and assembly index as a physical observable, which distinguish it from theoretical approaches to formalizing life that are unmoored from measurement.

cs.CC

Validation of the Scientific Literature via Chemputation Augmented by Large Language Models

Chemputation is the process of programming chemical robots to do experiments using a universal symbolic language, but the literature can be error prone and hard to read due to ambiguities. Large Language Models (LLMs) have demonstrated remarkable capabilities in various domains, including natural language processing, robotic control, and more recently, chemistry. Despite significant advancements in standardizing the reporting and collection of synthetic chemistry data, the automatic reproduction of reported syntheses remains a labour-intensive task. In this work, we introduce an LLM-based chemical research agent workflow designed for the automatic validation of synthetic literature procedures. Our workflow can autonomously extract synthetic procedures and analytical data from extensive documents, translate these procedures into universal XDL code, simulate the execution of the procedure in a hardware-specific setup, and ultimately execute the procedure on an XDL-controlled robotic system for synthetic chemistry. This demonstrates the potential of LLM-based workflows for autonomous chemical synthesis with Chemputers. Due to the abstraction of XDL this approach is safe, secure, and scalable since hallucinations will not be chemputable and the XDL can be both verified and encrypted. Unlike previous efforts, which either addressed only a limited portion of the workflow, relied on inflexible hard-coded rules, or lacked validation in physical systems, our approach provides four realistic examples of syntheses directly executed from synthetic literature. We anticipate that our workflow will significantly enhance automation in robotically driven synthetic chemistry research, streamline data extraction, improve the reproducibility, scalability, and safety of synthetic and experimental chemistry.

cs.AI

Mapping Evolution of Molecules Across Biochemistry with Assembly Theory

Evolution is often understood through genetic mutations driving changes in an organism's fitness, but there is potential to extend this understanding beyond the genetic code. We propose that natural products - complex molecules central to Earth's biochemistry can be used to uncover evolutionary mechanisms beyond genes. By applying Assembly Theory (AT), which views selection as a process not limited to biological systems, we can map and measure evolutionary forces in these molecules. AT enables the exploration of the assembly space of natural products, demonstrating how the principles of the selfish gene apply to these complex chemical structures, selecting vastly improbable and complex molecules from a vast space of possibilities. By comparing natural products with a broader molecular database, we can assess the degree of evolutionary contingency, providing insight into how molecular novelty emerges and persists. This approach not only quantifies evolutionary selection at the molecular level but also offers a new avenue for drug discovery by exploring the molecular assembly spaces of natural products. Our method provides a fresh perspective on measuring the evolutionary processes both, shaping and being read out, by the molecular imprint of selection.

q-bio.PE

AI-Driven Robotic Crystal Explorer for Rapid Polymorph Identification

Crystallisation is an important phenomenon which facilitates the purification as well as structural and bulk phase material characterisation using crystallographic methods. However, different conditions can lead to a vast set of different crystal structure polymorphs and these often exhibit different physical properties, allowing materials to be tailored to specific purposes. This means the high dimensionality that can result from variations in the conditions which affect crystallisation, and the interaction between them, means that exhaustive exploration is difficult, time-consuming, and costly to explore. Herein we present a robotic crystal search engine for the automated and efficient high-throughput approach to the exploration of crystallisation conditions. The system comprises a closed-loop computer crystal-vision system that uses machine learning to both identify crystals and classify their identity in a multiplexed robotic platform. By exploring the formation of a well-known polymorph, we were able to show how a robotic system could be used to efficiently search experimental space as a function of relative polymorph amount and efficiently create a high dimensionality phase diagram with minimal experimental budget and without expensive analytical techniques such as crystallography. In this way, we identify the set of polymorphs possible within a set of experimental conditions, as well as the optimal values of these conditions to grow each polymorph.

cs.RO

Experimental Measurement of Assembly Indices are Required to Determine The Threshold for Life

Assembly Theory (AT) was developed to help distinguish living from non-living systems. The theory is simple as it posits that the amount of selection or Assembly is a function of the number of complex objects where their complexity can be objectively determined using assembly indices. The assembly index of a given object relates to the number of recursive joining operations required to build that object and can be not only rigorously defined mathematically but can be experimentally measured. In pervious work we outlined the theoretical basis, but also extensive experimental measurements that demonstrated the predictive power of AT. These measurements showed that is a threshold in assembly indices for organic molecules whereby abiotic chemical systems could not randomly produce molecules with an assembly index greater or equal than 15. In a recent paper by Hazen et al [1] the authors not only confused the concept of AT with the algorithms used to calculate assembly indices, but also attempted to falsify AT by calculating theoretical assembly indices for objects made from inorganic building blocks. A fundamental misunderstanding made by the authors is that the threshold is a requirement of the theory, rather than experimental observation. This means that exploration of inorganic assembly indices similarly requires an experimental observation, correlated with the theoretical calculations. Then and only then can the exploration of complex inorganic molecules be done using AT and the threshold for living systems, as expressed with such building blocks, be determined. Since Hazen et al.[1] present no experimental measurements of assembly theory, their analysis is not falsifiable.

q-bio.OT

On the Salient Limitations of `On the Salient Limitations of the Methods of Assembly Theory and their Classification of Molecular Biosignatures'

Assembly Theory (AT) is a theory that explains how to determine if a complex object is the product of evolution. Here we explain why attempts to compare AT to compression algorithms, ref 1, does not help identify if the object is the product of selection or not. Specifically, we show why aims to perform benchmark comparisons of different compression schemes to compare the performance of Molecular Assembly Indices against standard compression schemes in determining living vs. non-living samples fails to classify the data correctly. In their approach, Uthamacumaran et al., ref 1, compress data from Marshall et al.2 describing the experimental basis for assembly theory and evaluate the difference in the resulting distributions of data. After several computational experiments Uthamacumaran et al., ref 1, conclude that other compression techniques obtain better results in the problem of detecting life and non-life that of assembly pathways. Here we explain why the approach presented by Uthamacumaran et al.1 appears to be lacking, and why it is not to reproduce their analysis with the information given.

cs.IT