SearcharxivSearch

arXiv subjects

Henry Wynn

Publications and source records attributed to Henry Wynn.

14 recordsLinked to original sources

Redundancy analysis using lcm-filtrations: networks, system signature and sensitivity evaluation

We introduce the lcm-filtration and stepwise filtration, comparing their performance across various scenarios in terms of computational complexity, efficiency, and redundancy. The lcm-filtration often involves identical steps or ideals, leading to unnecessary computations. To address this, we analyse how stepwise filtration can effectively compute only the non-identical steps, offering a more efficient approach. We compare these filtrations in applications to networks, system signatures, and sensitivity analysis.

cs.DM

Binary De Bruijn Processes

Binary time series data are very common in many applications, and are typically modelled independently via a Bernoulli process with a single probability of success. However, the probability of a success can be dependent on the outcome successes of past events. Presented here is a novel approach for modelling binary time series data called a binary de Bruijn process which takes into account temporal correlation. The structure is derived from de Bruijn Graphs - a directed graph, where given a set of symbols, V, and a 'word' length, m, the nodes of the graph consist of all possible sequences of V of length m. De Bruijn Graphs are equivalent to mth order Markov chains, where the 'word' length controls the number of states that each individual state is dependent on. This increases correlation over a wider area. To quantify how clustered a sequence generated from a de Bruijn process is, the run lengths of letters are observed along with run length properties. Inference is also presented along with two application examples: precipitation data and the Oxford and Cambridge boat race.

stat.ME

Modelling Correlated Bernoulli Data Part II: Inference

Binary data are highly common in many applications, however it is usually modelled with the assumption that the data are independently and identically distributed. This is typically not the case in many real-world examples and such the probability of a success can be dependent on the outcome successes of past events. The de Bruijn process (DBP) was introduced in Kimpton et al. [2022]. This is a correlated Bernoulli process which can be used to model binary data with known correlation. The correlation structures are included through the use of de Bruijn graphs, giving an extension to Markov chains. Given the DBP and an observed sequence of binary data, we present a method of inference using Bayes' factors. Results are applied to the Oxford and Cambridge annual boat race.

stat.ME

Lattice Conditional Independence Models and Hibi Ideals

Lattice Conditional Independence models are a class of models developed first for the Gaussian case in which a distributive lattice classifies all the conditional independence statements. The main result is that these models can equivalently be described via a transitive directed acyclic graph (TDAG) in which, as is normal for causal models, the conditional independence is in terms of conditioning on ancestors in the graph. We demonstrate that a parallel stream of research in algebra, the theory of Hibi ideals, not only maps directly to the LCI models but gives a vehicle to generalise the theory from the linear Gaussian case. Given a distributive lattice (i) each conditional independence statement is associated with a Hibi relation defined on the lattice, (ii) the directed graph is given by chains in the lattice which correspond to chains of conditional independence, (iii) the elimination ideal of product terms in the chains gives the Hibi ideal and (iv) the TDAG can be recovered from a special bipartite graph constructed via the Alexander dual of the Hibi ideal. It is briefly demonstrated that there are natural applications to statistical log-linear models, time series, and Shannon information flow.

math.ST

Comparing district heating options under uncertainty using stochastic ordering

District heating is a network of pipes through which heat is delivered from a centralised source. It is expected to play an important role in the decarbonisation of the energy sector in the coming years. In district heating, heat is traditionally generated through fossil fuels, often with combined heat and power (CHP) units. However, increasingly, waste heat is being used as a low carbon alternative, either directly or, for low temperature sources, via a heat pump. The design of district heating often has competing objectives: the need for inexpensive energy and meeting low carbon targets. In addition, the planning of district heating schemes is subject to multiple sources of uncertainty such as variability in heat demand and energy prices. This paper proposes a decision support tool to analyse and compare system designs for district heating under uncertainty using stochastic ordering (dominance). Contrary to traditional uncertainty metrics that provide statistical summaries and impose total ordering, stochastic ordering is a partial ordering and operates with full probability distributions. In our analysis, we apply the orderings in the mean and dispersion to the waste heat recovery problem in Brunswick, Germany.

stat.AP

Majorisation as a theory for uncertainty

Majorisation, also called rearrangement inequalities, yields a type of stochastic ordering in which two or more distributions can be compared. In this paper we argue that majorisation is a good candidate as a theory for uncertainty. We present operations that can be applied to study uncertainty in a range of settings and demonstrate our approach to assessing uncertainty with examples from well known distributions and from applications of climate projections and energy systems.

math.ST

A mean first passage time genome rearrangement distance

This paper introduces a new way to define a genome rearrangement distance, using the concept of mean first passage time from probability theory. Crucially, this distance estimate provides a genuine metric on genome space. We develop the theory and introduce a link to a graph-based zeta function. The approach is very general and can be applied to a wide variety of group-theoretic models of genome evolution.

q-bio.PE

The role of low temperature waste heat recovery in achieving 2050 goals: a policy positioning paper

Urban waste heat recovery, in which low temperature heat from urban sources is recovered for use in a district heat network, has a great deal of potential in helping to achieve 2050 climate goals. For example, heat from data centres, metro systems, public sector buildings and waste water treatment plants could be used to supply ten percent of Europe's heat demand. Despite this, at present, urban waste heat recovery is not widespread and is an immature technology. To help achieve greater uptake, three policy recommendations are made. First, policy raising awareness of waste heat recovery and creating a legal framework is suggested. Second, it is recommended that pilot projects are promoted to help demonstrate technical and economic feasibility. Finally, a pilot credit facility is proposed aimed at bridging the gap between potential investors and heat recovery projects.

econ.GN

The Scenario Culture

Scenario Analysis is a risk assessment tool that aims to evaluate the impact of a small number of distinct plausible future scenarios. In this paper, we provide an overview of important aspects of Scenario Analysis including when it is appropriate, the design of scenarios, uncertainty and encouraging creativity. Each of these issues is discussed in the context of climate, energy and legal scenarios.

stat.OT

Bregman divergences based on optimal design criteria and simplicial measures of dispersion

In previous work the authors defined the k-th order simplicial distance between probability distributions which arises naturally from a measure of dispersion based on the squared volume of random simplices of dimension k. This theory is embedded in the wider theory of divergences and distances between distributions which includes Kullback-Leibler, Jensen-Shannon, Jeffreys-Bregman divergence and Bhattacharyya distance. A general construction is given based on defining a directional derivative of a function $ϕ$ from one distribution to the other whose concavity or strict concavity influences the properties of the resulting divergence. For the normal distribution these divergences can be expressed as matrix formula for the (multivariate) means and covariances. Optimal experimental design criteria contribute a range of functionals applied to non-negative, or positive definite, information matrices. Not all can distinguish normal distributions but sufficient conditions are given. The k-th order simplicial distance is revisited from this aspect and the results are used to test empirically the identity of means and covariances.

math.ST

Group kernels for Gaussian process metamodels with categorical inputs

Gaussian processes (GP) are widely used as a metamodel for emulating time-consuming computer codes. We focus on problems involving categorical inputs, with a potentially large number L of levels (typically several tens), partitioned in G << L groups of various sizes. Parsimonious covariance functions, or kernels, can then be defined by block covariance matrices T with constant covariances between pairs of blocks and within blocks. We study the positive definiteness of such matrices to encourage their practical use. The hierarchical group/level structure, equivalent to a nested Bayesian linear model, provides a parameterization of valid block matrices T. The same model can then be used when the assumption within blocks is relaxed, giving a flexible parametric family of valid covariance matrices with constant covariances between pairs of blocks. The positive definiteness of T is equivalent to the positive definiteness of a smaller matrix of size G, obtained by averaging each block. The model is applied to a problem in nuclear waste analysis, where one of the categorical inputs is atomic number, which has more than 90 levels.

math.ST

An extended Generalised Variance, with Applications

We consider a measure $ψ$ k of dispersion which extends the notion of Wilk's generalised variance, or entropy, for a d-dimensional distribution, and is based on the mean squared volume of simplices of dimension k $\le$ d formed by k + 1 independent copies. We show how $ψ$ k can be expressed in terms of the eigenvalues of the covariance matrix of the distribution, also when a n-point sample is used for its estimation, and prove its concavity when raised at a suitable power. Some properties of entropy-maximising distributions are derived, including a necessary and sufficient condition for optimality. Finally, we show how this measure of dispersion can be used for the design of optimal experiments, with equivalence to A and D-optimal design for k = 1 and k = d respectively. Simple illustrative examples are presented.

math.ST

Minimal average degree aberration and the state polytope for experimental designs

For a particular experimental design, there is interest in finding which polynomial models can be identified in the usual regression set up. The algebraic methods based on Groebner bases provide a systematic way of doing this. The algebraic method does not in general produce all estimable models but it can be shown that it yields models which have minimal average degree in a well-defined sense and in both a weighted and unweighted version. This provides an alternative measure to that based on "aberration" and moreover is applicable to any experimental design. A simple algorithm is given and bounds are derived for the criteria, which may be used to give asymptotic Nyquist-like estimability rates as model and sample sizes increase.

stat.ME

Nonlinear Matroid Optimization and Experimental Design

We study the problem of optimizing nonlinear objective functions over matroids presented by oracles or explicitly. Such functions can be interpreted as the balancing of multi-criteria optimization. We provide a combinatorial polynomial time algorithm for arbitrary oracle-presented matroids, that makes repeated use of matroid intersection, and an algebraic algorithm for vectorial matroids. Our work is partly motivated by applications to minimum-aberration model-fitting in experimental design in statistics, which we discuss and demonstrate in detail.

math.CO