SearcharxivSearch

arXiv subjects

Daniel John Lawson

Publications and source records attributed to Daniel John Lawson.

9 recordsLinked to original sources

Valid Conformal Prediction for Dynamic GNNs

Dynamic graphs provide a flexible data abstraction for modelling many sorts of real-world systems, such as transport, trade, and social networks. Graph neural networks (GNNs) are powerful tools allowing for different kinds of prediction and inference on these systems, but getting a handle on uncertainty, especially in dynamic settings, is a challenging problem. In this work we propose to use a dynamic graph representation known in the tensor literature as the unfolding, to achieve valid prediction sets via conformal prediction. This representation, a simple graph, can be input to any standard GNN and does not require any modification to existing GNN architectures or conformal prediction routines. One of our key contributions is a careful mathematical consideration of the different inference scenarios which can arise in a dynamic graph modelling context. For a range of practically relevant cases, we obtain valid prediction sets with almost no assumptions, even dispensing with exchangeability. In a more challenging scenario, which we call the semi-inductive regime, we achieve valid prediction under stronger assumptions, akin to stationarity. We provide real data examples demonstrating validity, showing improved accuracy over baselines, and sign-posting different failure modes which can occur when those assumptions are violated.

stat.ML

A Simple and Powerful Framework for Stable Dynamic Network Embedding

In this paper, we address the problem of dynamic network embedding, that is, representing the nodes of a dynamic network as evolving vectors within a low-dimensional space. While the field of static network embedding is wide and established, the field of dynamic network embedding is comparatively in its infancy. We propose that a wide class of established static network embedding methods can be used to produce interpretable and powerful dynamic network embeddings when they are applied to the dilated unfolded adjacency matrix. We provide a theoretical guarantee that, regardless of embedding dimension, these unfolded methods will produce stable embeddings, meaning that nodes with identical latent behaviour will be exchangeable, regardless of their position in time or space. We additionally define a hypothesis testing framework which can be used to evaluate the quality of a dynamic network embedding by testing for planted structure in simulated networks. Using this, we demonstrate that, even in trivial cases, unstable methods are often either conservative or encode incorrect structure. In contrast, we demonstrate that our suite of stable unfolded methods are not only more interpretable but also more powerful in comparison to their unstable counterparts.

cs.SI

Meta-analysis of mid-p-values: some new results based on the convex order

The mid-p-value is a proposed improvement on the ordinary p-value for the case where the test statistic is partially or completely discrete. In this case, the ordinary p-value is conservative, meaning that its null distribution is larger than a uniform distribution on the unit interval, in the usual stochastic order. The mid-p-value is not conservative. However, its null distribution is dominated by the uniform distribution in a different stochastic order, called the convex order. The property leads us to discover some new finite-sample and asymptotic bounds on functions of mid-p-values, which can be used to combine results from different hypothesis tests conservatively, yet more powerfully, using mid-p-values rather than p-values. Our methodology is demonstrated on real data from a cyber-security application.

math.ST

Posterior predictive p-values and the convex order

Posterior predictive p-values are a common approach to Bayesian model-checking. This article analyses their frequency behaviour, that is, their distribution when the parameters and the data are drawn from the prior and the model respectively. We show that the family of possible distributions is exactly described as the distributions that are less variable than uniform on [0,1], in the convex order. In general, p-values with such a property are not conservative, and we illustrate how the theoretical worst-case error rate for false rejection can occur in practice. We describe how to correct the p-values to recover conservatism in several common scenarios, for example, when interpreting a single p-value or when combining multiple p-values into an overall score of significance. We also handle the case where the p-value is estimated from posterior samples obtained from techniques such as Markov Chain or Sequential Monte Carlo. Our results place posterior predictive p-values in a much clearer theoretical framework, allowing them to be used with more assurance.

math.ST

A general decision framework for structuring computation using Data Directional Scaling to process massive similarity matrices

As datasets grow it becomes infeasible to process them completely with a desired model. For giant datasets, we frame the order in which computation is performed as a decision problem. The order is designed so that partial computations are of value and early stopping yields useful results. Our approach comprises two related tools: a decision framework to choose the order to perform computations, and an emulation framework to enable estimation of the unevaluated computations. The approach is applied to the problem of computing similarity matrices, for which the cost of computation grows quadratically with the number of objects. Reasoning about similarities before they are observed introduces difficulties as there is no natural space and hence comparisons are difficult. We solve this by introducing a computationally convenient form of multidimensional scaling we call `data directional scaling'. High quality estimation is possible with massively reduced computation from the naive approach, and can be scaled to very large matrices. The approach is applied to the practical problem of assessing genetic similarity in population genetics. The use of statistical reasoning in decision making for large scale problems promises to be an important tool in applying statistical methodology to Big Data.

stat.CO

Apparent strength conceals instability in a model for the collapse of historical states

An explanation for the political processes leading to the sudden collapse of empires and states would be useful for understanding both historical and contemporary political events. We seek a general description of state collapse spanning eras and cultures, from small kingdoms to continental empires, drawing on a suitably diverse range of historical sources. Our aim is to provide an accessible verbal hypothesis that bridges the gap between mathematical and social methodology. We use game-theory to determine whether factions within a state will accept the political status quo, or wish to better their circumstances through costly rebellion. In lieu of precise data we verify our model using sensitivity analysis. We find that a small amount of dissatisfaction is typically harmless, but can trigger sudden collapse when there is a sufficient buildup of political inequality. Contrary to intuition, a state is predicted to be least stable when its leadership is at the height of its political power and thus most able to exert its influence through external warfare, lavish expense or autocratic decree.

physics.soc-ph

Populations in statistical genetic modelling and inference

What is a population? This review considers how a population may be defined in terms of understanding the structure of the underlying genetics of the individuals involved. The main approach is to consider statistically identifiable groups of randomly mating individuals, which is well defined in theory for any type of (sexual) organism. We discuss generative models using drift, admixture and spatial structure, and the ancestral recombination graph. These are contrasted with statistical models for inference, principle component analysis and other `non-parametric' methods. The relationships between these approaches are explored with both simulated and real-data examples. The state-of-the-art practical software tools are discussed and contrasted. We conclude that populations are a useful theoretical construct that can be well defined in theory and often approximately exist in practice.

q-bio.PE

Understanding clustering in type space using field theoretic techniques

The birth/death process with mutation describes the evolution of a population, and displays rich dynamics including clustering and fluctuations. We discuss an analytical `field-theoretical' approach to the birth/death process, using a simple dimensional analysis argument to describe evolution as a `Super-Brownian Motion' in the infinite population limit. The field theory technique provides corrections to this for large but finite population, and an exact description at arbitrary population size. This allows a characterisation of the difference between the evolution of a phenotype, for which strong local clustering is observed, and a genotype for which distributions are more dispersed. We describe the approach with sufficient detail for non-specialists.

q-bio.PE

Neutral Evolution as Diffusion in phenotype space: reproduction with mutation but without selection

The process of `Evolutionary Diffusion', i.e. reproduction with local mutation but without selection in a biological population, resembles standard Diffusion in many ways. However, Evolutionary Diffusion allows the formation of local peaks with a characteristic width that undergo drift, even in the infinite population limit. We analytically calculate the mean peak width and the effective random walk step size, and obtain the distribution of the peak width which has a power law tail. We find that independent local mutations act as a diffusion of interacting particles with increased stepsize.

q-bio.PE