SearcharxivSearch

arXiv subjects

Sara I. Walker

Publications and source records attributed to Sara I. Walker.

At least 19 recordsLinked to original sources

Elemental Stoichiometry as an Ecological Biosignature with Applications to Life Detection

The vast chemical space of possible small molecules, estimated at 10^60 compounds for molecules composed of just C, N, O, and S, is only sparsely occupied by biology. We propose that where life selects molecules within this space constitutes a detectable ecological signature: a fingerprint not of specific compounds, but of the statistical structure of elemental composition across molecules sam-pled from ecological systems. Here we introduce a framework combining Van Krevelen diagrams and element scaling laws to characterize the elemental composition of regions of chemical space occupied by biological systems and contrast them with other chemical systems. Applying this framework to 11,834 microbial metagenomic samples, we show that microbial metabolisms occupy a region of chemical space, which is enriched in heteroatoms such as P, S, N, and O relative to C, shifted toward higher O:C and H:C ratios. We observe sublinear element scaling with system size, yielding insights into how elemental constraints dictate how biological systems occupy chemical space. These patterns are distinct from a sample of 18,000 compounds from the comprehensive Reaxys synthetic chemical database. Critically, datasets from molecules detected in planetary science mission data occupy statistically distinct regions from both terrestrial biological and Reaxys distributions, demonstrating that with standardized methods for data collection, the approach could be developed to discriminate biotic from abiotic chemical signatures in small molecule data from planetary science missions. Our work shows how a combination of Van Krevelen fingerprinting and elemental scaling laws can provide a new class of ecological biosignatures for life detection leveraging mass spectrometric data from planetary missions, which could generalize beyond Earth's specific biochemistry.

q-bio.BM

Deep-time consistency in proteome elemental composition across cellular and viral life

Proteins are constructed from a limited alphabet of ~20 amino acids, yet the origins and selection of this specific alphabet are unresolved. One largely overlooked aspect is whether elemental composition constrains the range of viable proteomes. Here, we analyze the elemental composition of thousands of proteomes spanning cellular domains and viral realms. Despite evolutionary divergence and orders-of-magnitude variation in proteome size and gene content, proteomes exhibit strikingly consistent elemental composition. This consistency is substantially more constrained than amino acid frequencies or physicochemical properties and is not explained by evolutionary relatedness, biological function, or amino acid usage alone. Viral proteomes occupy the same elemental composition space observed in cellular organisms despite the absence of a single viral common ancestor, suggesting common biochemical constraints shape proteome organization across life. To investigate the evolutionary origins of this pattern, we compare modern proteomes with multiple independent reconstructions of the Last Universal Common Ancestor (LUCA) and with synthetic reduced-alphabet proteomes generated from primordial amino acid alphabets. LUCA proteomes occupy the same constrained elemental composition space observed in modern Bacteria and Archaea, whereas reduced primordial-like alphabets systematically generated alternative elemental regimes outside the modern range despite retaining high sequence similarity to extant proteins. Reduced alphabets disrupt fold space and reorganize relationships between elemental composition and predicted protein structural organization. Our results suggest that constrained elemental composition represents a fundamental organizational property of proteomes, which emerged early in evolution and may have contributed to the selection and stabilization of the modern amino acid alphabet.

q-bio.BM

The Physics of Causation

Assembly theory (AT) introduces causation as a material property and establishes a metrology for objects produced by evolution and selection. The physical scale of causation is quantified by the assembly index, defined as the minimum number of recursive steps necessary to make an object. Observing countable copies of high assembly index objects indicates a mechanism producing them is persistent, such that the object's environment constructs a memory that traps causation within a contingent chain. Copy number and assembly index together underlie a standardized metrology for detecting causation (assembly index) and contingency (copy number). These allow a precise definition of an assembly threshold that demarcates life (and its derivative agential, intelligent, and technological forms and artifacts) as structures with persistent copies in regimes of deep causal possibility. In introducing a fundamental concept of material causation to quantify and measure life, AT represents a departure from prior theories of causation, such as interventional ones, which have so far proven incompatible with fundamental physics. We discuss how AT's concept of causation provides the foundation for a theory of physics that allows precise and testable concept of "life", and in which novelty, contingency and the potential for open-endedness are fundamental, and determinism is emergent from selection along assembled lineages.

physics.hist-ph

Rapid Exploration of Assembly Chemical Space of Molecular Graphs

Quantifying how hard it is to build a molecular graph matters for biosignature detection, chemical complexity, and cheminformatics. We present an exact, scalable algorithm to compute the molecular assembly index (MA) which prioritizes the largest duplicate subgraphs, represents fragmentation with an 'assembly state' array of edge-lists, reuses states via hashing/DAGs, and prunes the search using a dynamic-programming branch-and-bound guided by a conditional-addition-chain lower bound. For organic molecules in the greater than 500 Da range our approach is up to six orders of magnitude faster than prior methods and yields exact MAs where previous algorithms would have timed out. We compute MAs to convergence for ~300k COCONUT natural products with <50 bonds, profiling time and memory scaling. Finally, we exploit the speed of our algorithm to calculate joint assembly spaces and introduce the Joint Assembly Overlap (JAO), a Jaccard-like metric that emphasizes global scaffold reuse and show that the JAO yields substantially different rankings from Tanimoto similarity with ECFP fingerprints and MCS (e.g. in steroids 270-380/Da and short peptides), accounting for substructural similarity beyond local environments. Together, these advances turn the molecular assembly index into a practical tool for large-scale exploration of chemical space.

cs.DS

Quantifying the Complexity of Materials with Assembly Theory

Quantifying the evolution and complexity of materials is of importance in many areas of science and engineering, where a central open challenge is developing experimental complexity measurements to distinguish random structures from evolved or engineered materials. Assembly Theory (AT) was developed to measure complexity produced by selection, evolution and technology. Here, we extend the fundamentals of AT to quantify complexity in inorganic molecules and solid-state periodic objects such as crystals, minerals and microprocessors, showing how the framework of AT can be used to distinguish naturally formed materials from evolved and engineered ones by quantifying the amount of assembly using the assembly equation defined by AT. We show how tracking the Assembly of repeated structures within a material allows us formalizing the complexity of materials in a manner accessible to measurement. We confirm the physical relevance of our formal approach, by applying it to phase transformations in crystals using the HCP to FCC transformation as a model system. To explore this approach, we introduce random stacking faults in closed-packed systems simplified to one-dimensional strings and demonstrate how Assembly can track the phase transformation. We then compare the Assembly of closed-packed structures with random or engineered faults, demonstrating its utility in distinguishing engineered materials from randomly structured ones. Our results have implications for the study of pre-genetic minerals at the origin of life, optimization of material design in the trade-off between complexity and function, and new approaches to explore material technosignatures which can be unambiguously identified as products of engineered design.

cond-mat.mtrl-sci

Assembly Theory and its Relationship with Computational Complexity

Assembly theory (AT) quantifies selection using the assembly equation and identifies complex objects that occur in abundance based on two measurements, assembly index and copy number, where the assembly index is the minimum number of joining operations necessary to construct an object from basic parts, and the copy number is how many instances of the given object(s) are observed. Together these define a quantity, called Assembly, which captures the amount of causation required to produce objects in abundance in an observed sample. This contrasts with the random generation of objects. Herein we describe how AT's focus on selection as the mechanism for generating complexity offers a distinct approach, and answers different questions, than computational complexity theory with its focus on minimum descriptions via compressibility. To explore formal differences between the two approaches, we show several simple and explicit mathematical examples demonstrating that the assembly index, itself only one piece of the theoretical framework of AT, is formally not equivalent to other commonly used complexity measures from computer science and information theory including Shannon entropy, Huffman encoding, and Lempel-Ziv-Welch compression. We also include proofs that assembly index is not in the same computational complexity class as these compression algorithms and discuss fundamental differences in the ontological basis of AT, and assembly index as a physical observable, which distinguish it from theoretical approaches to formalizing life that are unmoored from measurement.

cs.CC

Experimental Measurement of Assembly Indices are Required to Determine The Threshold for Life

Assembly Theory (AT) was developed to help distinguish living from non-living systems. The theory is simple as it posits that the amount of selection or Assembly is a function of the number of complex objects where their complexity can be objectively determined using assembly indices. The assembly index of a given object relates to the number of recursive joining operations required to build that object and can be not only rigorously defined mathematically but can be experimentally measured. In pervious work we outlined the theoretical basis, but also extensive experimental measurements that demonstrated the predictive power of AT. These measurements showed that is a threshold in assembly indices for organic molecules whereby abiotic chemical systems could not randomly produce molecules with an assembly index greater or equal than 15. In a recent paper by Hazen et al [1] the authors not only confused the concept of AT with the algorithms used to calculate assembly indices, but also attempted to falsify AT by calculating theoretical assembly indices for objects made from inorganic building blocks. A fundamental misunderstanding made by the authors is that the threshold is a requirement of the theory, rather than experimental observation. This means that exploration of inorganic assembly indices similarly requires an experimental observation, correlated with the theoretical calculations. Then and only then can the exploration of complex inorganic molecules be done using AT and the threshold for living systems, as expressed with such building blocks, be determined. Since Hazen et al.[1] present no experimental measurements of assembly theory, their analysis is not falsifiable.

q-bio.OT

Assembly Theory Explains and Quantifies the Emergence of Selection and Evolution

Since the time of Darwin, scientists have struggled to reconcile the evolution of biological forms in a universe determined by fixed laws. These laws underpin the origin of life, evolution, human culture and technology, as set by the boundary conditions of the universe, however these laws cannot predict the emergence of these things. By contrast evolutionary theory works in the opposite direction, indicating how selection can explain why some things exist and not others. To understand how open-ended forms can emerge in a forward-process from physics that does not include their design, a new approach to understand the non-biological to biological transition is necessary. Herein, we present a new theory, Assembly Theory (AT), which explains and quantifies the emergence of selection and evolution. In AT, the complexity of an individual observable object is measured by its Assembly Index (a), defined as the minimal number of steps needed to construct the object from basic building blocks. Combining a with the copy number defines a new quantity called Assembly which quantifies the amount of selection required to produce a given ensemble of objects. We investigate the internal structure and properties of assembly space and quantify the dynamics of undirected exploratory processes as compared to the directed processes that emerge from selection. The implementation of assembly theory allows the emergence of selection in physical systems to be quantified at any scale as the transition from undirected-discovery dynamics to a selected process within the assembly space. This yields a mechanism for the onset of selection and evolution and a formal approach to defining life. Because the assembly of an object is easily calculable and measurable it is possible to quantify a lower limit on the amount of selection and memory required to produce complexity uniquely linked to biology in the universe.

physics.bio-ph

Clone Swarms: Learning to Predict and Control Multi-Robot Systems by Imitation

In this paper, we propose SwarmNet -- a neural network architecture that can learn to predict and imitate the behavior of an observed swarm of agents in a centralized manner. Tested on artificially generated swarm motion data, the network achieves high levels of prediction accuracy and imitation authenticity. We compare our model to previous approaches for modelling interaction systems and show how modifying components of other models gradually approaches the performance of ours. Finally, we also discuss an extension of SwarmNet that can deal with nondeterministic, noisy, and uncertain environments, as often found in robotics applications.

cs.NE

Beyond COVID-19: Network science and sustainable exit strategies

On May $28^{th}$ and $29^{th}$, a two day workshop was held virtually, facilitated by the Beyond Center at ASU and Moogsoft Inc. The aim was to bring together leading scientists with an interest in Network Science and Epidemiology to attempt to inform public policy in response to the COVID-19 pandemic. Epidemics are at their core a process that progresses dynamically upon a network, and are a key area of study in Network Science. In the course of the workshop a wide survey of the state of the subject was conducted. We summarize in this paper a series of perspectives of the subject, and where the authors believe fruitful areas for future research are to be found.

physics.soc-ph

Formalizing Falsification for Theories of Consciousness Across Computational Hierarchies

The scientific study of consciousness is currently undergoing a critical transition in the form of a rapidly evolving scientific debate regarding whether or not currently proposed theories can be assessed for their scientific validity. At the forefront of this debate is Integrated Information Theory (IIT), widely regarded as the preeminent theory of consciousness because of its quantification of consciousness in terms a scalar mathematical measure called $Φ$ that is, in principle, measurable. Epistemological issues in the form of the "unfolding argument" have provided a refutation of IIT by demonstrating how it permits functionally identical systems to have differences in their predicted consciousness. The implication is that IIT and any other proposed theory based on a system's causal structure may already be falsified even in the absence of experimental refutation. However, so far the arguments surrounding the issue of falsification of theories of consciousness are too abstract to readily determine the scope of their validity. Here, we make these abstract arguments concrete by providing a simple example of functionally equivalent machines realizable with table-top electronics that take the form of isomorphic digital circuits with and without feedback. This allows us to explicitly demonstrate the different levels of abstraction at which a theory of consciousness can be assessed. Within this computational hierarchy, we show how IIT is simultaneously falsified at the finite-state automaton (FSA) level and unfalsifiable at the combinatorial state automaton (CSA) level. We use this example to illustrate a more general set of criteria for theories of consciousness: to avoid being unfalsifiable or already falsified scientific theories of consciousness must be invariant with respect to changes that leave the inference procedure fixed at a given level in a computational hierarchy.

cs.AI

A Flexible Bayesian Framework for Assessing Habitability with Joint Observational and Model Constraints

The catalog of stellar evolution tracks discussed in our previous work is meant to help characterize exoplanet host-stars of interest for follow-up observations with future missions like JWST. However, the utility of the catalog has been predicated on the assumption that we would precisely know the age of the particular host-star in question; in reality, it is unlikely that we will be able to accurately estimate the age of a given system. Stellar age is relatively straightforward to calculate for stellar clusters, but it is difficult to accurately measure the age of an individual star to high precision. Unfortunately, this is the kind of information we should consider as we attempt to constrain the long-term habitability potential of a given planetary system of interest. This is ultimately why we must rely on predictions of accurate stellar evolution models, as well a consideration of what we can observably measure (stellar mass, composition, orbital radius of an exoplanet) in order to create a statistical framework wherein we can identify the best candidate systems for follow-up characterization. In this paper we discuss a statistical approach to constrain long-term planetary habitability by evaluating the likelihood that at a given time of observation, a star would have a planet in the 2 Gy continuously habitable zone (CHZ2). Additionally, we will discuss how we can use existing observational data (i.e. data assembled in the Hypatia catalog and the Kepler exoplanet host star database) for a robust comparison to the catalog of theoretical stellar models.

astro-ph.EP

Integrated Information Theory and Isomorphic Feed-Forward Philosophical Zombies

Any theory amenable to scientific inquiry must have testable consequences. This minimal criterion is uniquely challenging for the study of consciousness, as we do not know if it is possible to confirm via observation from the outside whether or not a physical system knows what it feels like to have an inside - a challenge referred to as the "hard problem" of consciousness. To arrive at a theory of consciousness, the hard problem has motivated the development of phenomenological approaches that adopt assumptions of what properties consciousness has based on first-hand experience and, from these, derive the physical processes that give rise to these properties. A leading theory adopting this approach is Integrated Information Theory (IIT), which assumes our subjective experience is a "unified whole", subsequently yielding a requirement for physical feedback as a necessary condition for consciousness. Here, we develop a mathematical framework to assess the validity of this assumption by testing it in the context of isomorphic physical systems with and without feedback. The isomorphism allows us to isolate changes in $Φ$ without affecting the size or functionality of the original system. Indeed, we show that the only mathematical difference between a "conscious" system with $Φ>0$ and an isomorphic "philosophical zombies" with $Φ=0$ is a permutation of the binary labels used to internally represent functional states. This implies $Φ$ is sensitive to functionally arbitrary aspects of a particular labeling scheme, with no clear justification in terms of phenomenological differences. In light of this, we argue any quantitative theory of consciousness, including IIT, should be invariant under isomorphisms if it is to avoid the existence of isomorphic philosophical zombies and the epistemological problems they pose.

cs.IT

Quantifying the pathways to life using assembly spaces

We have developed the concept of pathway assembly to explore the amount of extrinsic information required to build an object. To quantify this information in an agnostic way, we present a method to determine the amount of pathway assembly information contained within such an object by deconstructing the object into its irreducible parts, and then evaluating the minimum number of steps to reconstruct the object along any pathway. The mathematical formalisation of this approach uses an assembly space. By finding the minimal number of steps contained in the route by which the objects can be assembled within that space, we can compare how much information (I) is gained from knowing this pathway assembly index (PA) according to I_PA=log (|N|)/(|N_PA |) where, for an end product with PA=x, N is the set of objects possible that can be created from the same irreducible parts within x steps regardless of PA, and NPA is the subset of those objects with the precise pathway assembly index PA=x. Applying this formalism to objects formed in 1D, 2D and 3D space allows us to identify objects in the world or wider Universe that have high assembly numbers. We propose that objects with PA greater than a threshold are important because these are uniquely identifiable as those that must have been produced by biological or technological processes, rather than the assembly occurring via unbiased random processes alone. We think this approach is needed to help identify the new physical and chemical laws needed to understand what life is, by quantifying what life does.

cs.AI

A unified formal framework for developmental andevolutionary change in gene regulatory network models

The two most fundamental processes describing change in biology, development and evolu-tion, occur over drastically different timescales, difficult to reconcile within a unified framework. Development involves temporal sequences of cell states controlled by hierarchies of regulatory structures. It occurs over the lifetime of a single individual, and is associated to the gene expression level change of a given genotype. Evolution, by contrast entails genotypic change through the acquisition/loss of genes and changes in the network topology of interactions among genes. It involves the emergence of new, environmentally selected phenotypes over the lifetimes of many individuals. Here we present a model of regulatory network evolution that accounts for both timescales. We extend the framework of Boolean models of gene regulatory networks (GRN)-currently only applicable to describing development to include evolutionary processes. As opposed to one-to-one maps to specific attractors, we identify the phenotypes of the cells as the relevant macrostates of the GRN. A phenotype may now correspond to multiple attractors, and its formal definition no longer requires a fixed size for the genotype. This opens the possibility for a quantitative study of the phenotypic change of a genotype, which is itself changing over evolutionary timescales. We show how the realization of specific phenotypes can be controlled by gene duplication events (used here as an archetypal evolutionary event able to change the genotype), and how successive events of gene duplication lead to new regulatory structures via selection. At the same time, we show that our generalized framework does not inhibit network controllability and the possibility for network control theory to describe epigenetic signaling during development.

physics.bio-ph

Exoplanet Biosignatures: A Review of Remotely Detectable Signs of Life

In the coming years and decades, advanced space- and ground-based observatories will allow an unprecedented opportunity to probe the atmospheres and surfaces of potentially habitable exoplanets for signatures of life. Life on Earth, through its gaseous products and reflectance and scattering properties, has left its fingerprint on the spectrum of our planet. Aided by the universality of the laws of physics and chemistry, we turn to Earth's biosphere, both in the present and through geologic time, for analog signatures that will aid in the search for life elsewhere. Considering the insights gained from modern and ancient Earth, and the broader array of hypothetical exoplanet possibilities, we have compiled a state-of-the-art overview of our current understanding of potential exoplanet biosignatures including gaseous, surface, and temporal biosignatures. We additionally survey biogenic spectral features that are well-known in the specialist literature but have not yet been robustly vetted in the context of exoplanet biosignatures. We briefly review advances in assessing biosignature plausibility, including novel methods for determining chemical disequilibrium from remotely obtainable data and assessment tools for determining the minimum biomass required for a given atmospheric signature. We focus particularly on advances made since the seminal review by Des Marais et al. (2002). The purpose of this work is not to propose new biosignatures strategies, a goal left to companion papers in this series, but to review the current literature, draw meaningful connections between seemingly disparate areas, and clear the way for a path forward.

astro-ph.EP

Logic and connectivity jointly determine criticality in biological gene regulatory networks

The complex dynamics of gene expression in living cells can be well-approximated using Boolean networks. The average sensitivity is a natural measure of stability in these systems: values below one indicate typically stable dynamics associated with an ordered phase, whereas values above one indicate chaotic dynamics. This yields a theoretically motivated adaptive advantage to being near the critical value of one, at the boundary between order and chaos. Here, we measure average sensitivity for 66 publicly available Boolean network models describing the function of gene regulatory circuits across diverse living processes. We find the average sensitivity values for these networks are clustered around unity, indicating they are near critical. In many types of random networks, mean connectivity and the average activity bias of the logic functions have been found to be the most important network properties in determining average sensitivity, and by extension a network's criticality. Surprisingly, many of these gene regulatory networks achieve the near-critical state with and far from that predicted for critical systems: randomized networks sharing the local causal structure and local logic of biological networks better reproduce their critical behavior than controlling for macroscale properties such as and alone. This suggests the local properties of genes interacting within regulatory networks are selected to collectively be near-critical, and this non-local property of gene regulatory network dynamics cannot be predicted using the density of interactions alone.

q-bio.MN

How causal analysis can reveal autonomy in models of biological systems

Standard techniques for studying biological systems largely focus on their dynamical, or, more recently, their informational properties, usually taking either a reductionist or holistic perspective. Yet, studying only individual system elements or the dynamics of the system as a whole disregards the organisational structure of the system - whether there are subsets of elements with joint causes or effects, and whether the system is strongly integrated or composed of several loosely interacting components. Integrated information theory (IIT), offers a theoretical framework to (1) investigate the compositional cause-effect structure of a system, and to (2) identify causal borders of highly integrated elements comprising local maxima of intrinsic cause-effect power. Here we apply this comprehensive causal analysis to a Boolean network model of the fission yeast (Schizosaccharomyces pombe) cell-cycle. We demonstrate that this biological model features a non-trivial causal architecture, whose discovery may provide insights about the real cell cycle that could not be gained from holistic or reductionist approaches. We also show how some specific properties of this underlying causal architecture relate to the biological notion of autonomy. Ultimately, we suggest that analysing the causal organisation of a system, including key features like intrinsic control and stable causal borders, should prove relevant for distinguishing life from non-life, and thus could also illuminate the origin of life problem.

q-bio.QM