SearcharxivSearch

arXiv subjects

Jack Noonan

Publications and source records attributed to Jack Noonan.

15 recordsLinked to original sources

Learning from data with structured missingness

Missing data are an unavoidable complication in many machine learning tasks. When data are `missing at random' there exist a range of tools and techniques to deal with the issue. However, as machine learning studies become more ambitious, and seek to learn from ever-larger volumes of heterogeneous data, an increasingly encountered problem arises in which missing values exhibit an association or structure, either explicitly or implicitly. Such `structured missingness' raises a range of challenges that have not yet been systematically addressed, and presents a fundamental hindrance to machine learning at scale. Here, we outline the current literature and propose a set of grand challenges in learning from data with structured missingness.

stat.ML

Improving exploration strategies in large dimensions and rate of convergence of global random search algorithms

We consider global optimization problems, where the feasible region $\X$ is a compact subset of $\mathbb{R}^d$ with $d \geq 10$. For these problems, we demonstrate the following. First: the actual convergence of global random search algorithms is much slower than that given by the classical estimates, based on the asymptotic properties of random points. Second: the usually recommended space exploration schemes are inefficient in the non-asymptotic regime. Specifically, (a) uniform sampling on entire~$\X$ is much less efficient than uniform sampling on a suitable subset of $\X$, and (b) the effect of replacement of random points by low-discrepancy sequences is negligible.

math.OC

An integrated approach to test for missing not at random

Missing data can lead to inefficiencies and biases in analyses, in particular when data are missing not at random (MNAR). It is thus vital to understand and correctly identify the missing data mechanism. Recovering missing values through a follow up sample allows researchers to conduct hypothesis tests for MNAR, which are not possible when using only the original incomplete data. Investigating how properties of these tests are affected by the follow up sample design is little explored in the literature. Our results provide comprehensive insight into the properties of one such test, based on the commonly used selection model framework. We determine conditions for recovery samples that allow the test to be applied appropriately and effectively, i.e. with known Type I error rates and optimized with respect to power. We thus provide an integrated framework for testing for the presence of MNAR and designing follow up samples in an efficient cost-effective way. The performance of our methodology is evaluated through simulation studies as well as on a real data sample.

stat.ME

Covering of high-dimensional sets

Let $(\mathcal{X},ρ)$ be a metric space and $λ$ be a Borel measure on this space defined on the $σ$-algebra generated by open subsets of $\mathcal{X}$; this measure $λ$ defines volumes of Borel subsets of $\mathcal{X}$. The principal case is where $\mathcal{X} = \mathbb{R}^d$, $ρ$ is the Euclidean metric, and $λ$ is the Lebesgue measure. In this article, we are not going to pay much attention to the case of small dimensions $d$ as the problem of construction of good covering schemes for small $d$ can be attacked by the brute-force optimization algorithms. On the contrary, for medium or large dimensions (say, $d\geq 10$), there is little chance of getting anything sensible without understanding the main issues related to construction of efficient covering designs.

math.OC

Efficient quantization and weak covering of high dimensional cubes

Let $\mathbb{Z}_n = \{Z_1, \ldots, Z_n\}$ be a design; that is, a collection of $n$ points $Z_j \in [-1,1]^d$. We study the quality of quantization of $[-1,1]^d$ by the points of $\mathbb{Z}_n$ and the problem of quality of coverage of $[-1,1]^d$ by ${\cal B}_d(\mathbb{Z}_n,r)$, the union of balls centred at $Z_j \in \mathbb{Z}_n$. We concentrate on the cases where the dimension $d$ is not small ($d\geq 5$) and $n$ is not too large, $n \leq 2^d$. We define the design ${\mathbb{D}_{n,δ}}$ as a $2^{d-1}$ design defined on vertices of the cube $[-δ,δ]^d$, $0\leq δ\leq 1$. For this design, we derive a closed-form expression for the quantization error and very accurate approximations for {the coverage area} vol$([-1,1]^d \cap {\cal B}_d(\mathbb{Z}_n,r))$. We provide results of a large-scale numerical investigation confirming the accuracy of the developed approximations and the efficiency of the designs ${\mathbb{D}_{n,δ}}$.

math.ST

Random and quasi-random designs in group testing

For large classes of group testing problems, we derive lower bounds for the probability that all significant items are uniquely identified using specially constructed random designs. These bounds allow us to optimize parameters of the randomization schemes. We also suggest and numerically justify a procedure of constructing designs with better separability properties than pure random designs. We illustrate theoretical considerations with a large simulation-based study. This study indicates, in particular, that in the case of the common binary group testing, the suggested families of designs have better separability than the popular designs constructed from disjunct matrices. We also derive several asymptotic expansions and discuss the situations when the resulting approximations achieve high accuracy.

math.ST

Online change-point detection for a transient change

We consider a popular online change-point problem of detecting a transient change in distributions of i.i.d. random variables. For this change-point problem, several change-point procedures are formulated and some advanced results for a particular procedure are surveyed. Some new approximations for the average run length to false alarm are offered and the power of these procedures for detecting a transient change in mean of a sequence of normal random variables is compared.

math.ST

Non-lattice covering and quanitization of high dimensional sets

The main problem considered in this paper is construction and theoretical study of efficient $n$-point coverings of a $d$-dimensional cube $[-1,1]^d$. Targeted values of $d$ are between 5 and 50; $n$ can be in hundreds or thousands and the designs (collections of points) are nested. This paper is a continuation of our paper \cite{us}, where we have theoretically investigated several simple schemes and numerically studied many more. In this paper, we extend the theoretical constructions of \cite{us} for studying the designs which were found to be superior to the ones theoretically investigated in \cite{us}. We also extend our constructions for new construction schemes which provide even better coverings (in the class of nested designs) than the ones numerically found in \cite{us}. In view of a close connection of the problem of quantization to the problem of covering, we extend our theoretical approximations and practical recommendations to the problem of construction of efficient quantization designs in a cube $[-1,1]^d$. In the last section, we discuss the problems of covering and quantization in a $d$-dimensional simplex; practical significance of this problem has been communicated to the authors by Professor Michael Vrahatis, a co-editor of the present volume.

math.ST

Generic probabilistic modelling and non-homogeneity issues for the UK epidemic of COVID-19

Coronavirus COVID-19 spreads through the population mostly based on social contact. To gauge the potential for widespread contagion, to cope with associated uncertainty and to inform its mitigation, more accurate and robust modelling is centrally important for policy making. We provide a flexible modelling approach that increases the accuracy with which insights can be made. We use this to analyse different scenarios relevant to the COVID-19 situation in the UK. We present a stochastic model that captures the inherently probabilistic nature of contagion between population members. The computational nature of our model means that spatial constraints (e.g., communities and regions), the susceptibility of different age groups and other factors such as medical pre-histories can be incorporated with ease. We analyse different possible scenarios of the COVID-19 situation in the UK. Our model is robust to small changes in the parameters and is flexible in being able to deal with different scenarios. This approach goes beyond the convention of representing the spread of an epidemic through a fixed cycle of susceptibility, infection and recovery (SIR). It is important to emphasise that standard SIR-type models, unlike our model, are not flexible enough and are also not stochastic and hence should be used with extreme caution. Our model allows both heterogeneity and inherent uncertainty to be incorporated. Due to the scarcity of verified data, we draw insights by calibrating our model using parameters from other relevant sources, including agreement on average (mean field) with parameters in SIR-based models.

stat.AP

Comparison of different exit scenarios from the lock-down for COVID-19 epidemic in the UK and assessing uncertainty of the predictions

We model further development of the COVID-19 epidemic in the UK given the current data and assuming different scenarios of handling the epidemic. In this research, we further extend the stochastic model suggested in \cite{us} and incorporate in it all available to us knowledge about parameters characterising the behaviour of the virus and the illness induced by it. The models we use are flexible, comprehensive, fast to run and allow us to incorporate the following: -time-dependent strategies of handling the epidemic; -spatial heterogeneity of the population and heterogeneity of development of epidemic in different areas; -special characteristics of particular groups of people, especially people with specific medical pre-histories and elderly. Standard epidemiological models such as SIR and many of its modifications are not flexible enough and hence are not precise enough in the studies that requires the use of the features above. Decision-makers get serious benefits from using better and more flexible models as they can avoid of nuanced lock-downs, better plan the exit strategy based on local population data, different stages of the epidemic in different areas, making specific recommendations to specific groups of people; all this resulting in a lesser impact on economy, improved forecasts of regional demand upon NHS allowing for intelligent resource allocation.

q-bio.PE

Covering of high-dimensional cubes and quantization

As the main problem, we consider covering of a $d$-dimensional cube by $n$ balls with reasonably large $d$ (10 or more) and reasonably small $n$, like $n=100$ or $n=1000$. We do not require the full coverage but only 90\% or 95\% coverage. We establish that efficient covering schemes have several important properties which are not seen in small dimensions and in asymptotical considerations, for very large $n$. One of these properties can be termed `do not try to cover the vertices' as the vertices of the cube and their close neighbourhoods are very hard to cover and for large $d$ there are far too many of them. We clearly demonstrate that, contrary to a common belief, placing balls at points which form a low-discrepancy sequence in the cube, makes for a very inefficient covering scheme. For a family of random coverings, we are able to provide very accurate approximations to the coverage probability. We then extend our results to the problems of coverage of a cube by smaller cubes and quantization, the latter being also referred to as facility location. Along with theoretical considerations and derivation of approximations, we discuss results of a large-scale numerical investigation.

math.ST

Approximations for the boundary crossing probabilities of moving sums of random variables

In this paper we study approximations for the boundary crossing probabilities of moving sums of i.i.d. normal r.v. We approximate a discrete time problem with a continuous time problem allowing us to apply established theory for stationary Gaussian processes. By then subsequently correcting approximations for discrete time, we show that the developed approximations are very accurate even for small window length. Also, they have high accuracy when the original r.v. are not exactly normal and when the weights in the moving window are not all equal. We then provide accurate and simple approximations for ARL, the average run length until crossing the boundary.

math.ST

Approximations for the boundary crossing probabilities of moving sums of normal random variables

In this paper we study approximations for boundary crossing probabilities for the moving sums of i.i.d. normal random variables. We propose approximating a discrete time problem with a continuous time problem allowing us to apply developed theory for stationary Gaussian processes and to consider a number of approximations (some well known and some not). We bring particular attention to the strong performance of a newly developed approximation that corrects the use of continuous time results in a discrete time setting. Results of extensive numerical comparisons are reported. These results show that the developed approximation is very accurate even for small window length.

math.ST

Approximating Shepp's constants for the Slepian process

Slepian process $S(t)$ is a stationary Gaussian process with zero mean and covariance $ E S(t)S(t')=\max\{0,1-|t-t'|\}\, . $ For any $T>0$ and $h>0$, define $F_T(h ) = {\rm Pr}\left\{\max_{t \in [0,T]} S(t) < h \right\} $ and the constants $Λ(h) = -\lim_{T \to \infty} \frac1T \log F_T(h)$ and $λ(h)=\exp\{-Λ(h) \}$; we will call them `Shepp's constants'. The aim of the paper is construction of accurate approximations for $F_T(h)$ and hence for the Shepp's constants. We demonstrate that at least some of the approximations are extremely accurate.

math.PR

First passage time for Slepian process with linear barrier

In this paper we extend results of L.A. Shepp by finding explicit formulas for the first passage probability $F_{a,b}(T\, |\, x)={\rm Pr}(S(t) 0$, where $S(t)$ is a Gaussian process with mean 0 and covariance $\mathbb{E} S(t)S(t')=\max\{0,1-|t-t'|\}\,.$ We then extend the results to the case of piecewise-linear barriers and outline applications to change-point detection problems. Previously, explicit formulas for $F_{a,b}(T\, |\, x)$ were known only for the cases $b=0$ (constant barrier) or $T\leq 1$ (short interval).

math.PR