SearcharxivSearch

arXiv subjects

Rudy Arthur

Publications and source records attributed to Rudy Arthur.

At least 19 recordsLinked to original sources

Community Detection with the Canonical Ensemble

Network community detection is usually considered as an unsupervised learning problem. Given a network, the aim is to partition it using some general purpose algorithm. In this paper we instead treat community detection as a hypothesis testing problem. Given a network, we examine the evidence for specific community structure in the observed network compared to a null model. To do this we define an appropriate test statistic, analogous to a z-score, and several null models derived from maximising entropy under different constraints in the canonical ensemble. We demonstrate the application of this method on real and synthetic data and contrast our method to Bayesian approaches based on the stochastic block model. We demonstrate that this method gives definitive answers to concrete questions, which can be more useful to analysts than the output of a generic algorithm.

cs.SI

Measuring ESG Risk in Supply Networks

Environmental, Social and Governance (ESG) rating is a way for investors to prioritise investments in companies with good corporate behaviour. However, ESG ratings are vulnerable to greenwashing in a number of ways. In this paper we study the effect that trade with badly rated companies has on a target company's own rating. To do this we introduce a measurement framework, generalising PageRank and Alpha Centrality, which allows tuning of aggregation and path counting approaches to resist greenwashing and reflect the rater's opinions and preferences for harm accumulation. These metrics allow updating of the target's ESG rating, identification of influential neighbours and assessment of vulnerability of the target to bad behaviour in their supply network. We study these metrics on synthetic ESG interaction networks as well as a real inter-company network and the international trade network.

cs.SI

Adaptation by Cumulative Selection

Biological systems like long-lived clonal organisms, holobionts and clades challenge traditional evolutionary thinking since they adapt without populations or reproduction. This paper aims to provide an overarching theoretical framework which encompasses standard Darwinian evolution as well as other processes of adaptation. This framework is cumulative selection and I provide a general `recipe' for it to occur. Lewontin's recipe for evolution by natural selection is shown to be a particular example of cumulative selection, but not the only one. Similarly, reproduction, inheritance and populations are just one way to perform cumulative selection. I discuss several other examples of cumulative selection including clonal organisms, dioecious populations, Gaia and neural networks.

q-bio.PE

Life on the Edge: Using Planetary Context to Enhance Biosignatures and Avoid False Positives

We use a probability theory framework to discuss the search for biosignatures. This perspective allows us to analyse the potential for different biosignatures to provide convincing evidence of extraterrestrial life and to formalise frameworks for accumulating evidence. Analysing biosignatures as a function of planetary context motivates the introduction of 'peribiosignatures', biosignatures observed where life is unlikely. We argue, based on prior work in Gaia theory, that habitability itself is an example of a peribiosignature. Finally, we discuss the implications of context dependence on observational strategy, suggesting that searching the edges of the habitable zone rather than the middle might be more likely to provide convincing evidence of life.

astro-ph.EP

Valeriepieris Circles Reveal City and Regional Boundaries in England and Wales

We propose a new method of determining regional and city boundaries based on the Valeriepieris circle, the smallest circle containing a given fraction of the data. By varying the fraction in the circle we can map complex spatial data to a simple model of concentric rings which we then fit to determine natural density cutoffs. We apply this method to population, occupation, economic and transport data from England and Wales, finding that the regions determined by this method affirm well known social facts such as the disproportionate wealth of London or the relative isolation of the North East and South West of England. We then show how different data sets give us different views of the same cities, providing insight into their development and dynamics.

stat.AP

Exploring Network Structure with the Density of States

Community detection, as well as the identification of other structures like core periphery and disassortative patterns, is an important topic in network analysis. While most methods seek to find the best partition of the network according to some criteria, there is a body of results that suggest that a single network can have many good but distinct partitions. In this paper we introduce the density of states as a tool for studying the space of all possible network partitions. We demonstrate how to use the well known Wang-Landau method to compute a network's density of states. We show that, even using modularity to measure quality, the density of states can still rule out spurious structure in random networks and overcome resolution limits. We demonstrate how these methods can be used to find `building blocks', groups of nodes which are consistently found together in detected communities. This suggests an approach to partitioning based on exploration of the network's structure landscape rather than optimisation.

cs.SI

Constraints on Meso-Scale Structure in Complex Networks

A key topic in network science is the detection of intermediate or meso-scale structures. Community, core-periphery, disassortative and other partitions allow us to understand the organisation and function of large networks. In this work we study under what conditions certain common meso-scale structures are detectable using the idea of block modularity. We find that the configuration model imposes strong restrictions on core-periphery and related structures in directed networks. We derive inequalities expressing when such structures can be detected under the configuration model. Nestedness is closely related to core-periphery and is similarly restricted to only be detectable under certain conditions. We show that these conditions are a generalisation of the resolution limit to structures other than assortative communities. We show how block modularity is related to the degree corrected Stochastic Block Model and that optimisation of one can be made equivalent to the other in general. Finally, we discuss these issues in inferential versus descriptive approaches to meso-scale structure detection.

cs.SI

The Language of Weather: Social Media Reactions to Weather Accounting for Climatic and Linguistic Baselines

This study explores how different weather conditions influence public sentiment on social media, focusing on Twitter data from the UK. By considering climate and linguistic baselines, we improve the accuracy of weather-related sentiment analysis. Our findings show that emotional responses to weather are complex, influenced by combinations of weather variables and regional language differences. The results highlight the importance of context-sensitive methods for better understanding public mood in response to weather, which can enhance impact-based forecasting and risk communication in the context of climate change.

cs.HC

What doesn't kill Gaia makes her stronger

Life on Earth has experienced numerous upheavals over its approximately 4 billion year history. In previous work we have discussed how interruptions to stability lead, on average, to increases in habitability over time, a tendency we called Entropic Gaia. Here we continue this exploration, working with the Tangled Nature Model of co-evolution, to understand how the evolutionary history of life is shaped by periods of acute environmental stress. We find that while these periods of stress pose a risk of complete extinction, they also create opportunities for evolutionary exploration which would otherwise be impossible, leading to more populous and stable states among the survivors than in alternative histories without a stress period. We also study how the duration, repetition and number of refugia into which life escapes during the perturbation affects the final outcome. The model results are discussed in relation to both Earth history and the search for alien life.

q-bio.PE

Correlation and Autocorrelation of Data on Complex Networks

Networks where each node has one or more associated numerical values are common in applications. This work studies how summary statistics used for the analysis of spatial data can be applied to non-spatial networks for the purposes of exploratory data analysis. We focus primarily on Moran-type statistics and discuss measures of global autocorrelation, local autocorrelation and global correlation. We introduce null models based on fixing edges and permuting the data or fixing the data and permuting the edges. We demonstrate the use of these statistics on real and synthetic node-valued networks.

cs.SI

A General Method for Resampling Autocorrelated Spatial Data

Comparing spatial data sets is a ubiquitous task in data analysis, however the presence of spatial autocorrelation means that standard estimates of variance will be wrong and tend to over-estimate the statistical significance of correlations and other observations. While there are a number of existing approaches to this problem, none are ideal, requiring detailed analytical calculations, which are hard to generalise or detailed knowledge of the data generating process, which may not be available. In this work we propose a resampling approach based on Tobler's Law. By resampling the data with fixed spatial autocorrelation, measured by Moran's I, we generate a more realistic null model. Testing on real and synthetic data, we find that, as long as the spatial autocorrelation is not too strong, this approach works just as well as if we knew the data generating process.

stat.AP

A Critical Analysis of the What3Words Geocoding Algorithm

What3Words is a geocoding application that uses triples of words instead of alphanumeric coordinates to identify locations. What3Words has grown rapidly in popularity over the past few years and is used in logistical applications worldwide, including by emergency services. What3Words has also attracted criticism for being less reliable than claimed, in particular that the chance of confusing one address with another is high. This paper investigates these claims and shows that the What3Words algorithm for assigning addresses to grid boxes creates many pairs of confusable addresses, some of which are quite close together. The implications of this for the use of What3Words in critical or emergency situations is discussed.

cs.HC

Valeriepieris Circles for Spatial Data Analysis

The Valeriepieris circle is the smallest circle that can be draw on the globe containing half of the world's population. The Valeriepieris (VP) circle acts as a spatial median, effectively splitting spatial data into two halves in a unique way. In this paper the idea of the VP circle is generalised and a fast algorithm and corresponding software package to compute it are described. The VP circle is compared to other measures of centre and dispersion for population distributions and is shown to reflect expected differences between countries and changes over time. By studying the VP circle as a function of the included population fraction a new way of representing population distributions is constructed, as well as a mathematical model of its expected behaviour. Finally a measure of population `centralisation' is constructed which measures the tendency of a territory to be dominated by a single population centre or to have a more even distribution of population.

stat.AP

CIDER: Context sensitive sentiment analysis for short-form text

Researchers commonly perform sentiment analysis on large collections of short texts like tweets, Reddit posts or newspaper headlines that are all focused on a specific topic, theme or event. Usually, general-purpose sentiment analysis methods are used. These perform well on average but miss the variation in meaning that happens across different contexts, for example, the word "active" has a very different intention and valence in the phrase "active lifestyle" versus "active volcano". This work presents a new approach, CIDER (Context Informed Dictionary and sEmantic Reasoner), which performs context-sensitive linguistic analysis, where the valence of sentiment-laden terms is inferred from the whole corpus before being used to score the individual texts. In this paper, we detail the CIDER algorithm and demonstrate that it outperforms state-of-the-art generalist unsupervised sentiment analysis techniques on a large collection of tweets about the weather. CIDER is also applicable to alternative (non-sentiment) linguistic scales. A case study on gender in the UK is presented, with the identification of highly gendered and sentiment-laden days. We have made our implementation of CIDER available as a Python package: https://pypi.org/project/ciderpolarity/.

cs.CL

Does Gaia Play Dice? : Simple Models of non-Darwinian Selection

In this paper we introduce some simple models, based on rolling dice, to explore mechanisms proposed to explain planetary habitability. The idea is to study these selection mechanisms in an analytically tractable setting, isolating their consequences from other details which can confound or obscure their effect in more realistic models. We find that while the observable of interest, the face value shown on the die, `improves' over time in all models, for two of the more popular ideas, Selection by Survival and Sequential Selection, this is down to sampling effects. A modified version of Sequential Selection, Sequential Selection with Memory, implies a statistical tendency for systems to improve over time. We discuss the implications of this and its relationship to the ideas of the `Inhabitance Paradox' and the `Gaian bottleneck'.

q-bio.PE

A Gaian Habitable Zone

When searching for inhabited exoplanets, understanding the boundaries of the habitable zone around the parent star is key. If life can strongly influence its global environment, then we would expect the boundaries of the habitable zone to be influenced by the presence of life. Here using a simple abstract model of `tangled-ecology' where life can influence a global parameter, labelled as temperature, we investigate the boundaries of the habitable zone of our model system. As with other models of life-climate interactions, the species act to regulate the temperature. However, the system can also experience `punctuations', where the system's state jumps between different equilibria. Despite this, an ensemble of systems still tends to sustain or even improve conditions for life on average, a feature we call Entropic Gaia. The mechanism behind this is sequential selection with memory which is discussed in detail. With this modelling framework we investigate questions about how Gaia can affect and ultimately extend the habitable zone to what we call the Gaian habitable zone. This generates concrete predictions for the size of the habitable zone around stars, suggests directions for future work on the simulation of exoplanets and provides insight into the Gaian bottleneck hypothesis and the habitability/inhabitance paradox.

astro-ph.EP

Discovering Block Structure in Networks

A generalization of modularity, called block modularity, is defined. This is a quality function which evaluates a label assignment against an arbitrary block pattern. Therefore, unlike standard modularity or its variants, arbitrary network structures can be compared and an optimal block matrix can be determined. Some simple algorithms for optimising block modularity are described and applied on networks with planted structure. In many cases the planted structure is recovered. Cases where it is not are analysed and it is found that strong degree-correlations explain the planted structure so that the discovered pattern is more `surprising' than the planted one under the configuration model. Some well studied networks are analysed with this new method, which is found to automatically deconstruct the network in a very useful way for creating a summary of its key features.

physics.soc-ph

Studying the UK Job Market During the COVID-19 Crisis with Online Job Ads

The COVID-19 global pandemic and the lockdown policies enacted to mitigate it have had profound effects on the labour market. Understanding these effects requires us to obtain and analyse data in as close to real time as possible, especially as rules change rapidly and local lockdowns are enacted. In this work we study the UK labour market by analysing data from the online job board Reed.co.uk. Using topic modelling and geo-inference methods we are able to break down the data by sector and geography. We also study how the salary, contract type and mode of work have changed since the COVID-19 crisis hit the UK in March. Overall, vacancies were down by 60 to 70\% in the first weeks of lockdown. By the end of the year numbers had recovered somewhat, but the total job ad deficit is measured to be over 40\%. Broken down by sector, vacancies for hospitality and graduate jobs are greatly reduced, while there were more care work and nursing vacancies during lockdown. Differences by geography are less significant than between sectors, though there is some indication that local lockdowns stall recovery and less badly hit areas may have experienced a smaller reduction in vacancies. There are also small but significant changes in the salary distribution and number of full time and permanent jobs. In addition to these results, this work presents an open methodology that enables a rapid and detailed survey of the job market in these unsettled conditions and we describe a web application \url{jobtrender.com} that allows others to query this data set.

cs.CY