SearcharxivSearch

arXiv subjects

Ben O'Neill

Publications and source records attributed to Ben O'Neill.

At least 19 recordsLinked to original sources

Quantile-stratified sampling for multivariate normal simulations and other multivariate distributions

In this paper we show how to extend quantile-stratified sampling to produce simulations from various multivariate distributions. These simulations have desirable space-filling and coverage properties relative to simulation using IID sampling. We examine the coverage performance of these simulations against IID sampling by looking at plots of ordered log-density values from the simulations.

stat.ME

A Data-Informed Local Subspaces Method for Error-Bounded Lossy Compression of Large-Scale Scientific Datasets

The growing volume of scientific simulation data presents a significant challenge for storage and transfer. Error-bounded lossy compression has emerged as a critical solution for mitigating these challenges, providing a means to reduce data size while ensuring that reconstructed data remains valid for scientific analysis. In this paper, we present a data-driven scientific data compressor, called Discontinuous Data-informed Local Subspaces (Discontinuous DLS), to improve compression-to-error ratios over data-agnostic compressors. This error-bounded compressor leverages localized spatial and temporal subspaces, informed by the underlying data structure, to enhance compression efficiency and preserve key features. The presented technique is flexible and applicable to a wide range of scientific data, including fluid dynamics, environmental simulations, and other high-dimensional, time-dependent datasets. We describe the core principles of the method and demonstrate its ability to significantly reduce storage requirements without compromising critical data fidelity. The technique is implemented in a distributed computing environment using MPI, and its performance is evaluated against state-of-the-art error-bounded compression methods in terms of compression ratio and reconstruction accuracy. This study highlights discontinuous DLS as a promising approach for large-scale scientific data compression in high-performance computing environments, providing a robust solution for managing the growing data demands of modern scientific simulations.

cs.DC

Directional Gaussian hypergeometric beta distributions and their uses in contaminated binary sampling

We examine the Gaussian hypergeometric beta distribution and look at the effect of having an additional term in the density kernel relative to the standard beta distribution. We reparameterise and classify this distribution into left and right directional variants using parameters that give a simple and symmetrical representation of the directional push/pull from this additional term in the density kernel. We examine the properties of the directional variants and their uses in contaminated binary sampling using Bayesian inference. We find that the Gaussian hypergeometric beta distribution arises as the appropriate posterior distribution for inference in certain kinds of contaminated binary models and that the directional parameterisation aids in representation of the resulting Bayesian models. We derive a broad range of properties and computational methods for the directional variants of the distribution.

math.ST

One-dimensional quantile-stratified sampling and its application in statistical simulations

In this paper we examine quantile-stratified samples from a known univariate probability distribution, with stratification occurring over a partition of the quantile regions in the distribution. We examine some general properties of this sampling method and we contrast it with standard IID sampling to highlight its similarities and differences. We examine the applications of this sampling method to various statistical simulations including importance sampling. We conduct simulation analysis to compare the performance of standard importance sampling against the quantile-stratified importance sampling to see how they each perform on a range of functions.

stat.ME

Unifying design-based and model-based sampling theory -- some suggestions to clear the cobwebs

This paper gives a holistic overview of both the design-based and model-based paradigms for sampling theory. Both methods are presented within a unified framework with a simple consistent notation, and the differences in the two paradigms are explained within this common framework. We examine the different definitions of the "population variance" within the two paradigms and examine the use of Bessel's correction for a population variance. We critique some messy aspects of the presentation of the design-based paradigm and implore readers to avoid the standard presentation of this framework in favour of a more explicit presentation that includes explicit conditioning in probability statements. We also discuss a number of confusions that arise from the standard presentation of the design-based paradigm and argue that Bessel's correction should be applied to the population variance.

stat.ME

The distribution of order statistics under sampling without replacement

This paper examines the distribution of order statistics taken from simple-random-sampling without replacement (SRSWOR) from a finite population with values 1,...,N. This distribution is a shifted version of the beta-binomial distribution, parameterised in a particular way. We derive the distribution and show how it relates to the distribution of order statistics under IID sampling from a uniform distribution over the unit interval. We examine properties of the distribution, including moments and asymptotic results. We also generalise the distribution to sampling without replacement of order statistics from an arbitrary finite population. We examine the properties of the order statistics for inference about an unknown population size (called the German tank problem) and we derive relevant estimation results based on observation of an arbitrary set of order statistics. We also introduce an algorithm that simulates sampling without replacement of order statistics from an arbitrary finite population without having to generate the entire sample.

math.ST

Three Distributions in the Extended Occupancy Problem

The classical and extended occupancy distributions are useful for examining the number of occupied bins in problems involving random allocation of balls to bins. We examine the extended occupancy problem by framing it as a Markov chain and deriving the spectral decomposition of the transition probability matrix. We look at three distributions of interest that arise from the problem, all involving the noncentral Stirling numbers of the second kind. These distributions give a useful generalisation to the binomial and negative-binomial distributions. We examine how these distributions relate to one another, and we derive recursive properties and mixture properties that characterise the distributions.

math.PR

An exposition of possibility and probability

This paper considers the notion of possible events which are insignificant in probabilistic analysis (i.e. events that have zero probability). The paper discusses the method of modal logic based on "possible worlds" and discusses a mathematical framework for the concepts of possibility, impossibility and certainty that are sometimes (incorrectly) thought to be defined with respect to probability. The relationship between possibility and probability is explored for general probability spaces and for refinements of these spaces conditional on other events, with particular focus on the properties of events having zero probability. We derive conditions under which possibility and significance diverge and conditions under which they can be reconciled as equivalent ideas within certain contexts. We also apply this analysis to discuss issues in possibility and probability in the multinomial model.

stat.OT

Computing highest density regions for continuous univariate distributions with known probability functions

We examine the problem of computing the highest density region (HDR) in a computational context where the user has access to a density function and quantile function for the distribution (e.g., in the statistical language R). We examine several common classes of continuous univariate distributions based on the shape of the density function; this includes monotone densities, quasi-concave and quasi-convex densities, and general multimodal densities. In each case we show how the user can compute the HDR from the quantile and density functions by framing the problem as a nonlinear optimisation problem. We implement these methods in R to obtain general functions to compute HDRs for classes of distributions, and for commonly used families of distributions. We compare our method to existing R packages for computing HDRs and we show that our method performs favourably in terms of both accuracy and average speed.

stat.CO

Smallest covering regions and highest density regions for discrete distributions

This paper examines the problem of computing a canonical smallest covering region for an arbitrary discrete probability distribution. This optimisation problem is similar to the classical 0-1 knapsack problem, but it involves optimisation over a set that may be countably infinite, raising a computational challenge that makes the problem non-trivial. To solve the problem we present theorems giving useful conditions for an optimising region and we develop an iterative one-at-a-time computational method to compute a canonical smallest covering region. We show how this can be programmed in pseudo-code and we examine the performance of our method. We compare this algorithm with other algorithms available in statistical computation packages to compute HDRs. We find that our method is the only one that accurately computes HDRs for arbitrary discrete distributions.

stat.CO

Binomial Prediction Using the Frequent Outcome Approach

Within the context of the binomial model, we analyse sequences of values that are almost-uniform and we discuss a prediction method called the frequent outcome approach, in which the outcome that has occurred the most in the observed trials is the most likely to occur again. Using this prediction method we derive probability statements for the prior probability of correct prediction, conditional on the underlying parameter value in the binomial model. We show that this prediction method converges to a level of accuracy that is equivalent to ideal prediction based on knowledge of the model parameter.

math.ST

An examination of the negative occupancy distribution and the coupon-collector distribution

We examine the negative occupancy distribution and the coupon-collector distribution, both of which arise as distributions relating to hitting times in the extended occupancy problem. These distributions constitute a full solution to a generalised version of the coupon collector problem, by describing the behaviour of the number of items we need to collect to obtain a full collection or a partial collection of any size. We examine the properties of these distributions and show how they can be computed and approximated. We give some practical guidance on the feasibility of computing large blocks of values from the distributions, and when approximation is required.

math.PR

An examination of the spillage distribution

We examine a family of discrete probability distributions that describes the "spillage number" in the extended balls-in-bins model. The spillage number is defined as the number of balls that occupy their bins minus the total number of occupied bins. This probability distribution can be characterised as a normed version of the expansion of the noncentral Stirling numbers of the second kind in terms of the central Stirling numbers of the second kind. Alternatively it can be derived in a natural way from the extended balls-in-bins model. We derive the generating functions for this distribution and important moments of the distribution. We also derive an algorithm for recursive computation of the mass values for the distribution. Finally, we examine the asymptotic behaviour of the spillage distribution and the performance of an approximation to the distribution.

math.PR

A generalised matching distribution for the problem of coincidences

This paper examines the classical matching distribution arising in the "problem of coincidences". We generalise the classical matching distribution with a preliminary round of allocation where items are correctly matched with some fixed probability, and remaining non-matched items are allocated using simple random sampling without replacement. Our generalised matching distribution is a convolution of the classical matching distribution and the binomial distribution. We examine the properties of this latter distribution and show how its probability functions can be computes. We also show how to use the distribution for matching tests and inferences of matching ability.

stat.OT

Transformation and simulation for a generalised queuing problem using a G/G/n/G/+ queuing model

We examine a generalised queuing model which we call the G/G/n/G/+ model, which encompasses the G/G/n and G/G/n/s models as special cases. Our model accommodates useful generalisations in user behaviour and limitations on the facilities for the queuing process. Give a set of inputs for the users and facilities, we develop a recursive algorithm that computes all aspects of the queuing process, including the waiting-times, use-times and unserved-times for each user. We also show how the queue can be represented graphically in a "queuing plot". We use our algorithm to undertake simulation analysis to determine the distribution of queuing outputs given specified distributions for the inputs, and we show how this can be used to optimise the number of facilities. We conduct some simple simulations to illustrate the method using standard queuing models. Our method is implemented in various queuing functions in the utilities package in R.

stat.CO

Mathematical properties and finite-population correction for the Wilson score interval

In this paper we examine the properties of the Wilson score interval, used for inferences for an unknown binomial proportion parameter. We examine monotonicity and consistency properties of the interval and we generalise it to give two alternative forms for inferences undertaken in a finite population. We discuss the nature of the "finite population correction" in these generalised intervals and examine their monotonicity and consistency properties. This analysis gives the appropriate confidence interval for an unknown population proportion or unknown unsampled proportion in a finite or infinite population. We implement the generalised confidence interval forms in a user-friendly function in R.

math.ST

Gaussian ARMA models in the ts.extend package

This paper introduces and describes the R package ts.extend, which adds probability functions for stationary Gaussian ARMA models and some related utility functions for time-series. We show how to use the package to compute the density and distributions functions for models in this class, and generate random vectors from this model. The package allows the user to use marginal or conditional models using a simple syntax for conditioning variables and marginalised elements. This allows users to simulate time-series vectors from any stationary Gaussian ARMA model, even if some elements are conditional values or omitted values. We also show how to use the package to compute the spectral intensity of a time-series vector and implement the permutation-spectrum test for a time-series vector to detect the presence of a periodic signal.

stat.CO

The Permutation-Spectrum Test: Identifying Periodic Signals using the Maximum Fourier Intensity

This paper examines the problem of testing whether a discrete time-series vector contains a periodic signal or is merely noise. To do this we examine the stochastic behaviour of the maximum intensity of the observed time-series vector and formulate a simple hypothesis test that rejects the null hypothesis of exchangeability if the maximum intensity spike in the Fourier domain is "too big" relative to its null distribution. This comparison is undertaken by simulating the null distribution of the maximum intensity using random permutations of the time-series vector. We show that this test has a p-value that is uniformly distributed for an exchangeable time-series vector, and that the p-value increases when there is a periodic signal present in the observed vector. We compare our test to Fisher's spectrum test, which assumes normality of the underlying noise terms. We show that our test is more robust than this test, and accommodates noise vectors with fat tails.

stat.CO