SearcharxivSearch

arXiv subjects

Nick S. Jones

Publications and source records attributed to Nick S. Jones.

At least 19 recordsLinked to original sources

Interpretable model-free inference of parametric variation across time-series data through large-scale feature extraction

Here we address the problem of estimating the dimensionality and nature of parametric variation in an unknown generative process directly from time-series data, without specifying or fitting a model. In particular we suppose that inter-instance variation in collections of time series is caused by parametric variation in the generating model. We hypothesize that, given a sufficiently large library of time-series features, low-dimensional parametric variation will manifest as low-dimensional structure in feature space, enabling interpretable estimators of the underlying degrees of freedom to be constructed. We test our hypothesis using a library of over 7000 diverse and interpretable time-series statistics and thirteen simulated systems with known parametric variation, spanning linear stochastic processes, nonlinear oscillators, and chaotic dynamics. Our unsupervised, data-driven approach often reconstructs the underlying parametric variation across this extensive range of simulated dynamical systems while also yielding interpretable estimators for each underlying dimension. Applied to the movement dynamics of 1143 fruit flies, we use this method to extract biologically meaningful components corresponding to sex and circadian rhythmicity. Our results pave the way for much-needed data-driven methods to bridge the gap between interpretable theoretical understanding of dynamics and the large and complex datasets that characterize modern scientific problems.

physics.data-an

Limits of message passing for node classification: How class-bottlenecks restrict signal-to-noise ratio

Message passing neural networks (MPNNs) are powerful models for node classification but suffer from performance limitations under heterophily (low same-class connectivity) and structural bottlenecks in the graph. We provide a unifying statistical framework exposing the relationship between heterophily and bottlenecks through the signal-to-noise ratio (SNR) of MPNN representations. The SNR decomposes model performance into feature-dependent parameters and feature-independent sensitivities. We prove that the sensitivity to class-wise signals is bounded by higher-order homophily -- a generalisation of classical homophily to multi-hop neighbourhoods -- and show that low higher-order homophily manifests locally as the interaction between structural bottlenecks and class labels (class-bottlenecks). Through analysis of graph ensembles, we provide a further quantitative decomposition of bottlenecking into underreaching (lack of depth implying signals cannot arrive) and oversquashing (lack of breadth implying signals arriving on fewer paths) with closed-form expressions. We prove that optimal graph structures for maximising higher-order homophily are disjoint unions of single-class and two-class-bipartite clusters. This yields BRIDGE, a graph ensemble-based rewiring algorithm that achieves near-perfect classification accuracy across all homophily regimes on synthetic benchmarks and significant improvements on real-world benchmarks, by eliminating the ``mid-homophily pitfall'' where MPNNs typically struggle, surpassing current standard rewiring techniques from the literature. Our framework, whose code we make available for public use, provides both diagnostic tools for assessing MPNN performance, and simple yet effective methods for enhancing performance through principled graph modification.

cs.LG

Geodesic Length Distribution in Sparse Network Ensembles

A key task in the study of networked systems is to derive local and global properties that impact connectivity, synchronizability, and robustness; computing shortest paths or geodesics yields measures of network connectivity that can explain such phenomena. We derive an analytic distribution of geodesic lengths on the giant component in the supercritical regime -- when the giant component exists -- or on small components in the subcritical regime, of any sparse (and possibly directed) network with conditionally independent edges, in the infinite-size limit. We provide specific results for widely used network models like stochastic block models, dot product graphs, random geometric graphs, and sparse graphons. The survival function of the geodesic length distribution possesses a simple closed-form expression which is asymptotically tight for finite lengths, has a natural interpretation of traversing independent geodesics in the network, and delivers novel insight into the aforementioned network families.

cs.SI

Influencing dynamics on social networks without knowledge of network microstructure

Social network based information campaigns can be used for promoting beneficial health behaviours and mitigating polarisation (e.g. regarding climate change or vaccines). Network-based intervention strategies typically rely on full knowledge of network structure. It is largely not possible or desirable to obtain population-level social network data due to availability and privacy issues. It is easier to obtain information about individuals' attributes (e.g. age, income), which are jointly informative of an individual's opinions and their social network position. We investigate strategies for influencing the system state in a statistical mechanics based model of opinion formation. Using synthetic and data based examples we illustrate the advantages of implementing coarse-grained influence strategies on Ising models with modular structure in the presence of external fields. Our work provides a scalable methodology for influencing Ising systems on large graphs and the first exploration of the Ising influence problem in the presence of ambient (social) fields. By exploiting the observation that strong ambient fields can simplify control of networked dynamics, our findings open the possibility of efficiently computing and implementing public information campaigns using insights from social network theory without costly or invasive levels of data collection.

physics.soc-ph

What would it take to build a thermodynamically reversible Universal Turing machine? Computational and thermodynamic constraints in a molecular design

We outline the construction of a molecular system that could, in principle, implement a thermodynamically reversible Universal Turing Machine (UTM). By proposing a concrete-albeit idealised-design and operational protocol, we reveal fundamental challenges that arise when attempting to implement arbitrary computations reversibly. Firstly, the requirements of thermodynamic reversibility inevitably lead to an intricate design. Secondly, thermodynamically reversible UTMs, unlike simpler devices, must also be logically reversible. Finally, implementing multiple distinct computations in parallel is necessary to take the cost of external control per computation to zero, but this approach is complicated the distinct halting times of different computations.

cond-mat.stat-mech

Modularity maximisation for graphons

Networks are a widely-used tool to investigate the large-scale connectivity structure in complex systems and graphons have been proposed as an infinite size limit of dense networks. The detection of communities or other meso-scale structures is a prominent topic in network science as it allows the identification of functional building blocks in complex systems. When such building blocks may be present in graphons is an open question. In this paper, we define a graphon-modularity and demonstrate that it can be maximised to detect communities in graphons. We then investigate specific synthetic graphons and show that they may show a wide range of different community structures. We also reformulate the graphon-modularity maximisation as a continuous optimisation problem and so prove the optimal community structure or lack thereof for some graphons, something that is usually not possible for networks. Furthermore, we demonstrate that estimating a graphon from network data as an intermediate step can improve the detection of communities, in comparison with exclusively maximising the modularity of the network. While the choice of graphon-estimator may strongly influence the accord between the community structure of a network and its estimated graphon, we find that there is a substantial overlap if an appropriate estimator is used. Our study demonstrates that community detection for graphons is possible and may serve as a privacy-preserving way to cluster network data.

stat.CO

Inference of a universal social scale and segregation measures using social connectivity kernels

How people connect with one another is a fundamental question in the social sciences, and the resulting social networks can have a profound impact on our daily lives. Blau offered a powerful explanation: people connect with one another based on their positions in a social space. Yet a principled measure of social distance, allowing comparison within and between societies, remains elusive. We use the connectivity kernel of conditionally-independent edge models to develop a family of segregation statistics with desirable properties: they offer an intuitive and universal characteristic scale on social space (facilitating comparison across datasets and societies), are applicable to multivariate and mixed node attributes, and capture segregation at the level of individuals, pairs of individuals, and society as a whole. We show that the segregation statistics can induce a metric on Blau space (a space spanned by the attributes of the members of society) and provide maps of two societies. Under a Bayesian paradigm, we infer the parameters of the connectivity kernel from eleven ego-network datasets collected in four surveys in the United Kingdom and United States. The importance of different dimensions of Blau space is similar across time and location, suggesting a macroscopically stable social fabric. Physical separation and age differences have the most significant impact on segregation within friendship networks with implications for intergenerational mixing and isolation in later stages of life.

cs.SI

Survival of the densest accounts for the expansion of mitochondrial mutations in ageing

The expansion of deleted mitochondrial DNA (mtDNA) molecules has been linked to ageing, particularly in skeletal muscle fibres; its mechanism has remained unclear for three decades. Previous accounts assigned a replicative advantage to the deletions, but there is evidence that cells can, instead, selectively remove defective mtDNA. We present a spatial model that, without a replicative advantage, but instead through a combination of enhanced density for mutants and noise, produces a wave of expanding mutations with wave speed consistent with experimental data, unlike a standard model based on replicative advantage. We provide a formula that predicts that the wave speed drops with copy number, in agreement with experimental data. Crucially, our model yields travelling waves of mutants even if mutants are preferentially eliminated. Justified by this exemplar of how noise, density and spatial structure affect muscle ageing, we introduce the mechanism of stochastic survival of the densest, an alternative to replicative advantage, that may underpin other phenomena, like the evolution of altruism.

q-bio.PE

Democratizing University Research

We detail an experimental programme we have been testing in our university. Our Advanced Hackspace, attempts to give all members of the university, from students to technicians, free access to the means to develop their own interdisciplinary research ideas, with resources including access to specialized fellows and biological and chemical hacklabs. We assess the aspects of our programme that led to our community being one of the largest collectives in our university and critically examine the successes and failures of our trial programmes. We supply metrics for assessing progress and outline challenges. We conclude with future directions that advance interdisciplinary research empowerment for all university members.

physics.ed-ph

Inference and Influence of Large-Scale Social Networks Using Snapshot Population Behaviour without Network Data

Population behaviours, such as voting and vaccination, depend on social networks. Social networks can differ depending on behaviour type and are typically hidden. However, we do often have large-scale behavioural data, albeit only snapshots taken at one timepoint. We present a method that jointly infers large-scale network structure and a networked model of human behaviour using only snapshot population behavioural data. This exploits the simplicity of a few parameter, geometric socio-demographic network model and a spin based model of behaviour. We illustrate, for the EU Referendum and two London Mayoral elections, how the model offers both prediction and the interpretation of our homophilic inclinations. Beyond offering the extraction of behaviour specific network structure from large-scale behavioural datasets, our approach yields a crude calculus linking inequalities and social preferences to behavioural outcomes. We give examples of potential network sensitive policies: how changes to income inequality, a social temperature and homophilic preferences might have reduced polarisation in a recent election.

cs.SI

Community detection in networks without observing edges

We develop a Bayesian hierarchical model to identify communities in networks for which we do not observe the edges directly, but instead observe a series of interdependent signals for each of the nodes. Fitting the model provides an end-to-end community detection algorithm that does not extract information as a sequence of point estimates but propagates uncertainties from the raw data to the community labels. Our approach naturally supports multiscale community detection as well as the selection of an optimal scale using model comparison. We study the properties of the algorithm using synthetic data and apply it to daily returns of constituents of the S&P100 index as well as climate data from US cities.

cs.SI

Mitochondrial network state scales mtDNA genetic dynamics

Mitochondrial DNA (mtDNA) mutations cause severe congenital diseases but may also be associated with healthy aging. MtDNA is stochastically replicated and degraded, and exists within organelles which undergo dynamic fusion and fission. The role of the resulting mitochondrial networks in the time evolution of the cellular proportion of mutated mtDNA molecules (heteroplasmy), and cell-to-cell variability in heteroplasmy (heteroplasmy variance), remains incompletely understood. Heteroplasmy variance is particularly important since it modulates the number of pathological cells in a tissue. Here, we provide the first wide-reaching theoretical framework which bridges mitochondrial network and genetic states. We show that, under a range of conditions, the (genetic) rate of increase in heteroplasmy variance and de novo mutation are proportionally modulated by the (physical) fraction of unfused mitochondria, independently of the absolute fission-fusion rate. In the context of selective fusion, we show that intermediate fusion/fission ratios are optimal for the clearance of mtDNA mutants. Our findings imply that modulating network state, mitophagy rate and copy number to slow down heteroplasmy dynamics when mean heteroplasmy is low could have therapeutic advantages for mitochondrial disease and healthy aging.

q-bio.SC

Biochemical Szilard engines for memory-limited inference

By developing and leveraging an explicit molecular realisation of a measurement-and-feedback-powered Szilard engine, we investigate the extraction of work from complex environments by minimal machines with finite capacity for memory and decision-making. Living systems perform inference to exploit complex structure, or correlations, in their environment, but the physical limits and underlying cost/benefit trade-offs involved in doing so remain unclear. To probe these questions, we consider a minimal model for a structured environment - a correlated sequence of molecules - and explore mechanisms based on extended Szilard engines for extracting the work stored in these non-equilibrium correlations. We consider systems limited to a single bit of memory making binary 'choices' at each step. We demonstrate that increasingly complex environments allow increasingly sophisticated inference strategies to extract more energy than simpler alternatives, and argue that optimal design of such machines should also consider the energy reserves required to ensure robustness against fluctuations due to mistakes.

cond-mat.stat-mech

CompEngine: a self-organizing, living library of time-series data

Modern biomedical applications often involve time-series data, from high-throughput phenotyping of model organisms, through to individual disease diagnosis and treatment using biomedical data streams. Data and tools for time-series analysis are developed and applied across the sciences and in industry, but meaningful cross-disciplinary interactions are limited by the challenge of identifying fruitful connections. Here we introduce the web platform, CompEngine, a self-organizing, living library of time-series data that lowers the barrier to forming meaningful interdisciplinary connections between time series. Using a canonical feature-based representation, CompEngine places all time series in a common space, regardless of their origin, allowing users to upload their data and immediately explore interdisciplinary connections to other data with similar properties, and be alerted when similar data is uploaded in the future. In contrast to conventional databases, which are organized by assigned metadata, CompEngine incentivizes data sharing by automatically connecting experimental and theoretical scientists across disciplines based on the empirical structure of their data. CompEngine's growing library of interdisciplinary time-series data also facilitates comprehensively characterization of algorithm performance across diverse types of data, and can be used to empirically motivate the development of new time-series analysis algorithms.

cs.DB

Mitochondrial heterogeneity

Cell-to-cell heterogeneity drives a range of (patho)physiologically important phenomena, such as cell fate and chemotherapeutic resistance. The role of metabolism, and particularly mitochondria, is increasingly being recognised as an important explanatory factor in cell-to-cell heterogeneity. Most eukaryotic cells possess a population of mitochondria, in the sense that mitochondrial DNA (mtDNA) is held in multiple copies per cell, where the sequence of each molecule can vary. Hence intra-cellular mitochondrial heterogeneity is possible, which can induce inter-cellular mitochondrial heterogeneity, and may drive aspects of cellular noise. In this review, we discuss sources of mitochondrial heterogeneity (variations between mitochondria in the same cell, and mitochondrial variations between supposedly identical cells) from both genetic and non-genetic perspectives, and mitochondrial genotype-phenotype links. We discuss the apparent homeostasis of mtDNA copy number, the observation of pervasive intra-cellular mtDNA mutation (we term `microheteroplasmy') and developments in the understanding of inter-cellular mtDNA mutation (`macroheteroplasmy'). We point to the relationship between mitochondrial supercomplexes, cristal structure, pH and cardiolipin as a potential amplifier of the mitochondrial genotype-phenotype link. We also discuss mitochondrial membrane potential and networks as sources of mitochondrial heterogeneity, and their influence upon the mitochondrial genome. Finally, we revisit the idea of mitochondrial complementation as a means of dampening mitochondrial genotype-phenotype links in light of recent experimental developments. The diverse sources of mitochondrial heterogeneity, as well as their increasingly recognised role in contributing to cellular heterogeneity, highlights the need for future single-cell mitochondrial measurements in the context of cellular noise studies.

q-bio.SC

Large algebraic connectivity fluctuations in spatial network ensembles imply a predictive advantage from node location information

A Random Geometric Graph (RGG) ensemble is defined by the disordered distribution of its node locations. We investigate how this randomness drives sample-to-sample fluctuations in the dynamical properties of these graphs. We study the distributional properties of the algebraic connectivity which is informative of diffusion and synchronization timescales in graphs. We use numerical simulations to provide the first characterisation of the algebraic connectivity distribution for RGG ensembles. We find that the algebraic connectivity can show fluctuations relative to its mean on the order of $30 \%$, even for relatively large RGG ensembles ($N=10^5$). We explore the factors driving these fluctuations for RGG ensembles with different choices of dimensionality, boundary conditions and node distributions. Within a given ensemble, the algebraic connectivity can covary with the minimum degree and can also be affected by the presence of density inhomogeneities in the nodal distribution. We also derive a closed-form expression for the expected algebraic connectivity for RGGs with periodic boundary conditions for general dimension.

physics.soc-ph

Co-occurrence simplicial complexes in mathematics: identifying the holes of knowledge

In the last years complex networks tools contributed to provide insights on the structure of research, through the study of collaboration, citation and co-occurrence networks. The network approach focuses on pairwise relationships, often compressing multidimensional data structures and inevitably losing information. In this paper we propose for the first time a simplicial complex approach to word co-occurrences, providing a natural framework for the study of higher-order relations in the space of scientific knowledge. Using topological methods we explore the conceptual landscape of mathematical research, focusing on homological holes, regions with low connectivity in the simplicial structure. We find that homological holes are ubiquitous, which suggests that they capture some essential feature of research practice in mathematics. Holes die when a subset of their concepts appear in the same article, hence their death may be a sign of the creation of new knowledge, as we show with some examples. We find a positive relation between the dimension of a hole and the time it takes to be closed: larger holes may represent potential for important advances in the field because they separate conceptually distant areas. We also show that authors' conceptual entropy is positively related with their contribution to homological holes, suggesting that polymaths tend to be on the frontier of research.

physics.soc-ph

Designing the Optimal Bit: Balancing Energetic Cost, Speed and Reliability

We consider the technologically relevant costs of operating a reliable bit that can be erased rapidly. We find that both erasing and reliability times are non-monotonic in the underlying friction, leading to a trade-off between erasing speed and bit reliability. Fast erasure is possible at the expense of low reliability at moderate friction, and high reliability comes at the expense of slow erasure in the underdamped and overdamped limits. Within a given class of bit parameters and control strategies, we define "optimal" designs of bits that meet the desired reliability and erasing time requirements with the lowest operational work cost. We find that optimal designs always saturate the bound on the erasing time requirement, but can exceed the required reliability time if critically damped. The non-trivial geometry of the reliability and erasing time-scales allows us to exclude large regions of parameter space as sub-optimal. We find that optimal designs are either critically damped or close to critical damping under the erasing procedure.

cond-mat.stat-mech