SearcharxivSearch

arXiv subjects

Paul Kabaila

Publications and source records attributed to Paul Kabaila.

At least 19 recordsLinked to original sources

Adaptive-precision computation of custom Gauss quadrature for statistical applications

An $n$-point Gauss quadrature rule approximates the weighted integral of a function by a weighted sum of $n$ evaluations of this function and is exact for polynomials of degree at most $2n-1$. Such rules can be highly accurate with relatively few evaluations of this function. For weight functions associated with classical orthogonal polynomials of a continuous variable (such as Legendre, Hermite and Laguerre), these quadrature rules are readily available. We suppose that this is not the case, so that these rules must be custom-made. We present CustomGaussQuadrature, a Julia package that implements two approaches for computing the Gauss quadrature nodes and weights: the moment determinants method and the Stieltjes procedure. The principal contribution is an implementation of the ill-conditioned moment determinants method with adaptively chosen working precision. Comparisons between computations at increasing precisions provide practical error indicators used to target absolute and relative errors in the Gauss rule nodes and weights, respectively, of about $10^{-16}$. The package also adapts the working precision and number of auxiliary quadrature nodes used for the Stieltjes procedure. Numerical examples using scaled chi and Weibull probability density functions as weight functions show close agreement between the results obtained by the two approaches. The package is intended particularly for statistical applications in which a single custom Gauss rule is computed once and then reused to approximate many weighted integrals having the same weight function but different computationally expensive integrands.

stat.CO

Custom-made Gauss quadrature for statisticians

The theory and computational methods for custom-made Gauss quadrature have been described in Gautschi's 2004 monograph. Gautschi has also provided Fortran and MATLAB code for the implementation and illustration of these methods. We have written an R package, implemented in the high-precision arithmetic provided by the R package Rmpfr, that uses a moment-based method via moment determinants to compute a Gauss quadrature rule, with up to 33 nodes, provided that the moments can be computed to arbitrary precision using the standard mathematical functions provided by the Rmpfr package. Our hope is that the provision of our free R package and the numerical results that we present will encourage other statisticians to also consider the custom-made construction of Gauss quadrature rules.

stat.CO

Confidence intervals in general regression models that utilize uncertain prior information

We consider a general regression model, without a scale parameter. Our aim is to construct a confidence interval for a scalar parameter of interest $θ$ that utilizes the uncertain prior information that a distinct scalar parameter $τ$ takes the specified value $t$. This confidence interval should have good coverage properties. It should also have scaled expected length, where the scaling is with respect to the usual confidence interval, that (a) is substantially less than 1 when the prior information is correct, (b) has a maximum value that is not too large and (c) is close to 1 when the data and prior information are highly discordant. The asymptotic joint distribution of the maximum likelihood estimators $θ$ and $τ$ is similar to the joint distributions of these estimators in the particular case of a linear regression with normally distributed errors having known variance. This similarity is used to construct a confidence interval with the desired properties by using the confidence interval, computed using the R package ciuupi, that utilizes the uncertain prior information in this particular linear regression case. An important practical application of this confidence interval is to a quantal bioassay carried out to compare two similar compounds. In this context, the uncertain prior information is that the hypothesis of "parallelism" holds. We provide extensive numerical results that illustrate the properties of this confidence interval in this context.

stat.ME

Computation of the expected value of a function of a chi-distributed random variable

We consider the problem of numerically evaluating the expected value of a smooth bounded function of a chi-distributed random variable, divided by the square root of the number of degrees of freedom. This problem arises in the contexts of simultaneous inference, the selection and ranking of populations and in the evaluation of multivariate t probabilities. It also arises in the assessment of the coverage probability and expected volume properties of the some non-standard confidence regions. We use a transformation put forward by Mori, followed by the application of the trapezoidal rule. This rule has the remarkable property that, for suitable integrands, it is exponentially convergent. We use it to create a nested sequence of quadrature rules, for the estimation of the approximation error, so that previous evaluations of the integrand are not wasted. The application of the trapezoidal rule requires the approximation of an infinite sum by a finite sum. We provide a new easily computed upper bound on the error of this approximation. Our overall conclusion is that this method is a very suitable candidate for the computation of the coverage and expected volume properties of non-standard confidence regions.

stat.CO

Confidence intervals centred on bootstrap smoothed estimators: an impossibility result

Recently, Kabaila and Wijethunga assessed the performance of a confidence interval centred on a bootstrap smoothed estimator, with width proportional to an estimator of Efron's delta method approximation to the standard deviation of this estimator. They used a testbed situation consisting of two nested linear regression models, with error variance assumed known, and model selection using a preliminary hypothesis test. This assessment was in terms of coverage and scaled expected length, where the scaling is with respect to the expected length of the usual confidence interval with the same minimum coverage probability. They found that this confidence interval has scaled expected length that (a) has a maximum value that may be much greater than 1 and (b) is greater than a number slightly less than 1 when the simpler model is correct. We therefore ask the following question. For a confidence interval, centred on the bootstrap smoothed estimator, does there exist a formula for its data-based width such that, in this testbed situation, it has the desired minimum coverage and scaled expected length that (a) has a maximum value that is not too much larger than 1 and (b) is substantially less than 1 when the simpler model is correct? Using a recent decision-theoretic performance bound due to Kabaila and Kong, it is shown that the answer to this question is `no' for a wide range of scenarios.

math.ST

Finite sample properties of the Buckland-Burnham-Augustin confidence interval centered on a model averaged estimator

We consider the confidence interval centered on a frequentist model averaged estimator that was proposed by Buckland, Burnham & Augustin (1997). In the context of a simple testbed situation involving two linear regression models, we derive exact expressions for the confidence interval and then for the coverage and scaled expected length of the confidence interval. We use these measures to explore the exact finite sample performance of the Buckland-Burnham-Augustin confidence interval. We also explore the limiting asymptotic case (as the residual degrees of freedom increases) and compare our results for this case to those obtained for the asymptotic coverage of the confidence interval by Hjort & Claeskens (2003).

stat.ME

On confidence intervals centered on bootstrap smoothed estimators

Bootstrap smoothed (bagged) estimators have been proposed as an improvement on estimators found after preliminary data-based model selection. Efron, 2014, derived a widely applicable formula for a delta method approximation to the standard deviation of the bootstrap smoothed estimator. He also considered a confidence interval centered on the bootstrap smoothed estimator, with width proportional to the estimate of this standard deviation. Kabaila and Wijethunga, 2019, assessed the performance of this confidence interval in the scenario of two nested linear regression models, the full model and a simpler model, for the case of known error variance and preliminary model selection using a hypothesis test. They found that the performance of this confidence interval was not substantially better than the usual confidence interval based on the full model, with the same minimum coverage. We extend this assessment to the case of unknown error variance by deriving a computationally convenient exact formula for the ideal (i.e. in the limit as the number of bootstrap replications diverges to infinity) delta method approximation to the standard deviation of the bootstrap smoothed estimator. Our results show that, unlike the known error variance case, there are circumstances in which this confidence interval has attractive properties.

stat.ME

Confidence intervals centered on bootstrap smoothed estimators

Bootstrap smoothed (bagged) parameter estimators have been proposed as an improvement on estimators found after preliminary data-based model selection. The key result of Efron (2014) is a very convenient and widely applicable formula for a delta method approximation to the standard deviation of the bootstrap smoothed estimator. This approximation provides an easily computed guide to the accuracy of this estimator. In addition, Efron (2014) proposed a confidence interval centered on the bootstrap smoothed estimator, with width proportional to the estimate of this approximation to the standard deviation. We evaluate this confidence interval in the scenario of two nested linear regression models, the full model and a simpler model, and a preliminary test of the null hypothesis that the simpler model is correct. We derive computationally convenient expressions for the ideal bootstrap smoothed estimator and the coverage probability and expected length of this confidence interval. In terms of coverage probability, this confidence interval outperforms the post-model-selection confidence interval with the same nominal coverage and based on the same preliminary test. We also compare the performance of confidence interval centered on the bootstrap smoothed estimator, in terms of expected length, to the usual confidence interval, with the same minimum coverage probablility, based on the full model.

stat.ME

Admissibility of the usual confidence set for the mean of a univariate or bivariate normal population: The unknown-variance case

In the Gaussian linear regression model (with unknown mean and variance), we show that the standard confidence set for one or two regression coefficients is admissible in the sense of Joshi (1969). This solves a long-standing open problem in mathematical statistics, and this has important implications on the performance of modern inference procedures post-model-selection or post-shrinkage, particularly in situations where the number of parameters is larger than the sample size. As a technical contribution of independent interest, we introduce a new class of conjugate priors for the Gaussian location-scale model.

math.ST

The effect of a Durbin-Watson pretest on confidence intervals in regression

Consider a linear regression model and suppose that our aim is to find a confidence interval for a specified linear combination of the regression parameters. In practice, it is common to perform a Durbin-Watson pretest of the null hypothesis of zero first-order autocorrelation of the random errors against the alternative hypothesis of positive first-order autocorrelation. If this null hypothesis is accepted then the confidence interval centred on the Ordinary Least Squares estimator is used; otherwise the confidence interval centred on the Feasible Generalized Least Squares estimator is used. We provide new tools for the computation, for any given design matrix and parameter of interest, of graphs of the coverage probability functions of the confidence interval resulting from this two-stage procedure and the confidence interval that is always centred on the Feasible Generalized Least Squares estimator. These graphs are used to choose the better confidence interval, prior to any examination of the observed response vector.

stat.ME

Optimized recentered confidence spheres for the multivariate normal mean

Casella and Hwang, 1983, JASA, introduced a broad class of recentered confidence spheres for the mean $\boldsymbolθ$ of a multivariate normal distribution with covariance matrix $σ^2 \boldsymbol{I}$, for $σ^2$ known. Both the center and radius functions of these confidence spheres are flexible functions of the data. For the particular case of confidence spheres centered on the positive-part James-Stein estimator and with radius determined by empirical Bayes considerations, they show numerically that these confidence spheres have the desired minimum coverage probability $1-α$ and dominate the usual confidence sphere in terms of scaled volume. We shift the focus from the scaled volume to the scaled expected volume of the recentered confidence sphere. Since both the coverage probability and the scaled expected volume are functions of the Euclidean norm of $\boldsymbolθ$, it is feasible to optimize the performance of the recentered confidence sphere by numerically computing both the center and radius functions so as to optimize some clearly specified criterion. We suppose that we have uncertain prior information that $\boldsymbolθ= \boldsymbol{0}$. This motivates us to determine the center and radius functions of the confidence sphere by numerical minimization of the scaled expected volume of the confidence sphere at $\boldsymbolθ= \boldsymbol{0}$, subject to the constraints that (a) the coverage probability never falls below $1-α$ and (b) the radius never exceeds the radius of the standard $1-α$ confidence sphere. Our results show that, by focusing on this clearly specified criterion, significant gains in performance (in terms of this criterion) can be achieved. We also present analogous results for the much more difficult case that $σ^2$ is unknown.

stat.ME

The large sample coverage probability of confidence intervals in general regression models after a preliminary hypothesis test

We derive a computationally convenient formula for the large sample coverage probability of a confidence interval for a scalar parameter of interest following a preliminary hypothesis test that a specified vector parameter takes a given value in a general regression model. Previously, this large sample coverage probability could only be estimated by simulation. Our formula only requires the evaluation, by numerical integration, of either a double or triple integral, irrespective of the dimension of this specified vector parameter. We illustrate the application of this formula to a confidence interval for the log odds ratio of myocardial infarction when the exposure is recent oral contraceptive use, following a preliminary test that two specified interactions in a logistic regression model are zero. For this real-life data, we compare this large sample coverage probability with the actual coverage probability of this confidence interval, obtained by simulation.

stat.ME

Confidence Intervals that Utilize Uncertain Prior Information about Exogeneity in Panel Data

Consider panel data modelled by a linear random intercept model that includes a time-varying covariate. Suppose that we have uncertain prior information that this covariate is exogenous. We present a new confidence interval for the slope parameter that utilizes this uncertain prior information. This interval has minimum coverage probability very close to its nominal coverage. Let the scaled expected length of this new confidence interval be its expected length divided by the expected length of the confidence interval, with the same minimum coverage, constructed using the fixed effects model. This new interval has scaled expected length that (a) is substantially less than 1 when the prior information is correct, (b) has a maximum value that is not too much larger than 1 and (c) is close to 1 when the data strongly contradict the prior information. We illustrate the properties of this new interval using an airfare data set.

stat.ME

Upper bounds on the minimum coverage probability of model averaged tail area confidence intervals in regression

Frequentist model averaging has been proposed as a method for incorporating "model uncertainty" into confidence interval construction. Such proposals have been of particular interest in the environmental and ecological statistics communities. A promising method of this type is the model averaged tail area (MATA) confidence interval put forward by Turek and Fletcher, 2012. The performance of this interval depends greatly on the data-based model weights on which it is based. A computationally convenient formula for the coverage probability of this interval is provided by Kabaila, Welsh and Abeysekera, 2016, in the simple scenario of two nested linear regression models. We consider the more complicated scenario that there are many (32,768 in the example considered) linear regression models obtained as follows. For each of a specified set of components of the regression parameter vector, we either set the component to zero or let it vary freely. We provide an easily-computed upper bound on the minimum coverage probability of the MATA confidence interval. This upper bound provides evidence against the use of a model weight based on the Bayesian Information Criterion (BIC).

stat.ME

The Performance of the Turek-Fletcher Model Averaged Confidence Interval

We consider the model averaged tail area (MATA) confidence interval proposed by Turek and Fletcher, CSDA, 2012, in the simple situation in which we average over two nested linear regression models. We prove that the MATA for any reasonable weight function belongs to the class of confidence intervals defined by Kabaila and Giri, JSPI, 2009. Each confidence interval in this class is specified by two functions b and s. Kabaila and Giri show how to compute these functions so as to optimize these intervals in terms of satisfying the coverage constraint and minimizing the expected length for the simpler model, while ensuring that the expected length has desirable properties for the full model. These Kabaila and Giri "optimized" intervals provide an upper bound on the performance of the MATA for an arbitrary weight function. This fact is used to evaluate the MATA for a broad class of weights based on exponentiating a criterion related to Mallows' C_P. Our results show that, while far from ideal, this MATA performs surprisingly well, provided that we choose a member of this class that does not put too much weight on the simpler model.

stat.ME

Conditional assessment of the impact of a Hausman pretest on confidence intervals

We assess the impact of a Hausman pretest, applied to panel data, on a confidence interval for the slope, conditional on the observed values of the time-varying covariate. This assessment has the advantages that it (a) relates to the values of this covariate at hand, (b) is valid irrespective of how this covariate is generated, (c) uses finite sample results and (d) results in an assessment that is determined by the values of this covariate and only 2 unknown parameters. Our conditional analysis shows that the confidence interval constructed after a Hausman pretest should not be used.

math.ST

The impact of a Hausman pretest on the coverage probability and expected length of confidence intervals

In the analysis of clustered and longitudinal data, which includes a covariate that varies both between and within clusters (e.g. time-varying covariate in longitudinal data), a Hausman pretest is commonly used to decide whether subsequent inference is made using the linear random intercept model or the fixed effects model. We assess the effect of this pretest on the coverage probability and expected length of a confidence interval for the slope parameter. Our results show that for the small levels of significance of the Hausman pretest commonly used in applications, the minimum coverage probability of this confidence interval can be far below nominal. Furthermore, the expected length of this confidence interval is, on average, larger than the expected length of a confidence interval for the slope parameter based on the fixed effects model with the same minimum coverage.

stat.ME