SearcharxivSearch

arXiv subjects

Rose Baker

Publications and source records attributed to Rose Baker.

13 recordsLinked to original sources

Estimating accurate covariance matrices on fitted model parameters

The accurate computation of the covariance matrix of fitted model parameters is a somewhat neglected task in Statistics. Algorithms are given for computing accurate covariance matrices derived from computing the Hessian matrix by numerical differentiation, and also for the covariance matrix of the posterior distribution of model parameters. Evaluations on two datasets where the Hessian could be computed analytically show that the numerical differentiation algorithm is very accurate.

stat.CO

Discrete distributions from a Markov chain

A discrete-time stochastic process derived from a model of basketball is used to generalize any discrete distribution. The generalized distributions can have one or two more parameters than the parent distribution. Those derived from binomial, Poisson and negative binomial distributions can be underdispersed or overdispersed. The mean can be simply expressed in terms of model parameters, thus making inference for the mean straightforward. Probabilities can be quickly computed, enabling likelihood-based inference. Random number generation is also straightforward. The properties of some of the new distributions are described and their use is illustrated with examples.

stat.AP

Reactive Social distancing in a SIR model of epidemics such as COVID-19

A model of reactive social distancing in epidemics is proposed, in which the infection rate changes with the number infected. The final-size equation for the total number that the epidemic will infect can be derived analytically, as can the peak infection proportion. This model could assist planners during for example the COVID-19 pandemic.

q-bio.PE

Outliers in meta-analysis: an asymmetric trimmed-mean approach

The adaptive asymmetric trimmed mean is a known way of estimating central location, usually in conjunction with the bootstrap. It is here modified and applied to meta-analysis, as a way of dealing with outlying results by down-weighting the corresponding studies. This requires a modified bootstrap and a method of down-weighting studies, as opposed to removing single observations. This methodology is shown in analysis of some well-travelled datasets to down-weight outliers in agreement with other methods, and Monte-Carlo studies show that it does does not appreciably down-weight studies when outliers are absent. Conceptually simple, it does not make parametric assumptions about the outliers.

stat.AP

Creating new distributions using integration and summation by parts

Methods for generating new distributions from old can be thought of as techniques for simplifying integrals used in reverse. Hence integrating a probability density function (pdf) by parts provides a new way of modifying distributions; the resulting pdfs are integrals that sometimes require computation as special functions. Summation by parts can be used similarly for discrete distributions. The general methodology is given, with some examples of distribution classes and of specific distributions, and fits to data.

math.ST

A new measure of treatment effect for random-effects meta-analysis of comparative binary outcome data

Comparative binary outcome data are of fundamental interest in statistics and are often pooled in meta-analyses. Here we examine the simplest case where for each study there are two patient groups and a binary event of interest, giving rise to a series of $2 \times 2$ tables. A variety of measures of treatment effect are then available and are conventionally used in meta-analyses, such as the odds ratio, the risk ratio and the risk difference. Here we propose a new type of measure of treatment effect for this type of data that is very easily interpretable by lay audiences. We give the rationale for the new measure and we present three contrasting methods for computing its within-study variance so that it can be used in conventional meta-analyses. We then develop three alternative methods for random-effects meta-analysis that use our measure and we apply our methodolgy to some real examples. We conclude that our new measure is a fully viable alternative to existing measures. It has the advantage that its interpretation is especially simple and direct, so that its meaning can be more readily understood by those with little or no formal statistical training. This may be especially valuable when presenting `plain language summaries', such as those used by Cochrane.

stat.ME

A flexible and computationally tractable discrete distribution derived from a stationary renewal process

A class of discrete distributions can be derived from stationary renewal processes. They have the useful property that the mean is a simple function of the model parameters. Thus regressions of the distribution mean on covariates can be carried out and marginal effects of covariates calculated. Probabilities can be easily computed in closed form for only two such distributions, when the event interarrival times in the renewal process follow either a gamma or an inverse Gaussian distribution. The gamma-based distribution has more attractive properties and is described and fitted to data. The inverse-Gaussian based distribution is also briefly discussed.

stat.ME

A new generalization of the beta distribution

The beta distribution is the best-known distribution for modelling doubly-bounded data, \eg percentage data or probabilities. A new generalization of the beta distribution is proposed, which uses a cubic transformation of the beta random variable. The new distribution is label-invariant like the beta distribution and has rational expressions for the moments. This facilitates its use in mean regression. The properties are discussed, and two examples of fitting to data are given. A modification is also explored in which the Jacobian of the transformation is omitted. This gives rise to messier expressions for the moments but better modal behaviour. In addition, the Jacobian alone gives rise to a general quadratic distribution that is of interest. The new distributions allow good fitting of unimodal data that fit poorly to the beta distribution, and could also be useful as prior distributions.

stat.ME

Event count distributions from renewal processes: fast computation of probabilities

Discrete distributions derived from renewal processes, ie distributions of the number of events by some time t are beginning to be used in econometrics and health sciences. A new fast method is presented for computation of the probabilities for these distributions. We calculate the count probabilities by repeatedly convolving the discretized distribution, and then correct them using Richardson extrapolation. When just one probability is required, a second algorithm is described, an adaptation of De Pril's method, in which the computation time does not depend on the ordinality, so that even high-order probabilities can be rapidly found. Any survival distribution can be used to model the inter-arrival times, which gives a rich class of models with great flexibility for modelling both underdispersed and overdispersed data. This work could pave the way for the routine use of these distributions as an additional tool for modelling event count data. An empirical example using fertility data illustrates the use of the method and was fully implemented using an R package Countr developed by the authors and available from the Comprehensive R Archive Network.

stat.ME

Copulas from Order Statistics

A new class of copulas based on order statistics was introduced by Baker (2008). Here, further properties of the bivariate and multivariate copulas are described, such as that of likelihood ratio dominance (LRD), and further bivariate copulas are introduced that generalize the earlier work. One of the new copulas is an integral of a product of Bessel functions of imaginary argument, and can attain the Fr\'echet bound. The use of these copulas for fitting data is described, and illustrated with examples. It was found empirically that the multivariate copulas previously proposed are not flexible enough to be generally useful in data fitting, and further development is needed in this area.

stat.ME

Application of some new heavy-tailed survival distributions

Some new survival distributions are introduced based on a generalised exponential function. This class of distributions includes heavy-tailed generalisations of exponential, Weibull and gamma distributions. Properties of the distributions are described, and R code is available for computation of pdf, quantiles, inverse quantiles, random numbers, etc. A use of these distributions for robust inference is suggested, and this is exemplified with a Monte-Carlo study.

stat.ME

Properties and Applications of some Distributions derived from Frullani's integral

Frullani's integral dates from 1821, but a probabilistic interpretation of it has never been made. In this paper, Frullani's integral formula is shown to result from mixing a lifetime distribution by allowing the logarithm of the scale factor to be uniformly distributed over a finite range. This gives a class of long-tailed distributions related to slash distributions, where the pdf is simply expressed in terms of the survival function of the `parent' distribution. The resulting survival distributions have all moments finite, and can exhibit the bimodal hazard functions sometimes seen in practice. A distribution of this type analogous to the t-distribution is derived, the corresponding multivariate distributions are given, and two skewed versions of this distribution are derived. The use of the mixed distributions for inference is exemplified by fitting them to several datasets. It is expected that there will be many applications, in health, reliability, telecommunications, finance, etc.

stat.ME

A new distribution for robust least squares

A new distribution is introduced, which we call the twin-t distribution. This distribution is heavy-tailed like the t distribution, but closer to normality in the central part of the curve. Its properties are described, e.g. the pdf, the distribution function, moments, and random number generation. This distribution could have many applications, but here we focus on its use as an aid to robustness. We give examples of its application in robust regression and in curve fitting. Extensions such as skew and multivariate twin-t distributions, and a twin of

stat.ME