Searcharxiv⌕ Search

arXiv subjects

Luke A. Prendergast

Publications and source records attributed to Luke A. Prendergast.

At least 19 recordsLinked to original sources

The dangers of using three-number summaries to estimate unknown standard deviations: sensitivity analyses and some possible improvements incorporating shape

In recent years, there has been much progress toward the development of methods for converting three- and five-number summary statistics (i.e. minimum, maximum, median, and quartiles) to means and standard deviations (SDs). This is commonly done in the meta-analysis setting, where some studies report means and SDs, while other report quantile summaries. However, we show that three-number summaries, which are the most common, do not contain enough information to reliably estimate SDs. We show that very poor estimates can result, which may invalidate any inference and provide details of a sensitivity analysis that can allow researchers to have greater confidence in their results, or highlight potential sources of bias. We further explore whether nominating additional information can provide enough information regarding the unknown data shape to improve SD estimations, and in doing so introduce a new estimator using the scaled Beta distribution. Simulations and a real data example are used to highlight the advantages and disadvantages of this approach. A Web application is also provided to help researchers perform sensitivity analyses.

stat.ME↗

rquest: An R package for hypothesis tests and confidence intervals for quantiles and summary measures based on quantiles

Sample quantiles, such as the median, are often better suited than the sample mean for summarising location characteristics of a data set. Similarly, linear combinations of sample quantiles and ratios of such linear combinations, e.g. the interquartile range and quantile-based skewness measures, are often used to quantify characteristics such as spread and skew. While often reported, it is uncommon to accompany quantile estimates with confidence intervals or standard errors. The rquest package provides a simple way to conduct hypothesis tests and derive confidence intervals for quantiles, linear combinations of quantiles, ratios of dependent linear combinations (e.g., Bowley's measure of skewness) and differences and ratios of all of the above for comparisons between independent samples. Many commonly used measures based on quantiles are included, although it is also very simple for users to define their own. Additionally, quantile-based measures of inequality are also considered. The methods are based on recent research showing that reliable distribution-free confidence intervals can be obtained, even for moderate sample sizes. Several examples are provided herein.

stat.ME↗

Confidence intervals for median absolute deviations

The median absolute deviation (MAD) is a robust measure of scale that is simple to implement and easy to interpret. Motivated by this, we introduce interval estimators of the MAD to make reliable inferences for dispersion for a single population and ratios and differences of MADs for comparing two populations. Our simulation results show that the coverage probabilities of the intervals are very close to the nominal coverage for a variety of distributions. We have used partial influence functions to investigate the robustness properties of the difference and ratios of independent MADs.

math.ST↗

Slice Weighted Average Regression

It has previously been shown that ordinary least squares can be used to estimate the coefficients of the single-index model under only mild conditions. However, the estimator is non-robust leading to poor estimates for some models. In this paper we propose a new sliced least-squares estimator that utilizes ideas from Sliced Inverse Regression. Slices with problematic observations that contribute to high variability in the estimator can easily be down-weighted to robustify the procedure. The estimator is simple to implement and can result in vast improvements for some models when compared to the usual least-squares approach. While the estimator was initially conceived with the single-index model in mind, we also show that multiple directions can be obtained, therefore providing another notable advantage of using slicing with least squares. Several simulation studies and a real data example are included, as well as some comparisons with some other recent methods.

stat.ME↗

A note on switching eigenvalues under small perturbations

Sensitivity of eigenvectors and eigenvalues of symmetric matrix estimates to the removal of a single observation have been well documented in the literature. However, a complicating factor can exist in that the rank of the eigenvalues may change due to the removal of an observation, and with that so too does the perceived importance of the corresponding eigenvector. We refer to this problem as "switching of eigenvalues". Since there is not enough information in the new eigenvalues post observation removal to indicate that this has happened, how do we know that this switching has occurred? In this paper, we show that approximations to the eigenvalues can be used to help determine when switching may have occurred. We then discuss possible actions researchers can take based on this knowledge, for example making better choices when it comes to deciding how many principal components should be retained and adjustments to approximate influence diagnostics that perform poorly when switching has occurred. Our results are easily applied to any eigenvalue problem involving symmetric matrix estimators. We highlight our approach with application to a real data example.

stat.ME↗

Extending the coefficient of variation for measuring heterogeneity following a meta-regression

Meta-regression is often used to form hypotheses about what is associated with heterogeneity in a meta-analysis and to estimate the extent to which effects can vary between cohorts and other distinguishing factors. However, study-level variables, called moderators, that are available and used in the meta-regression analysis will rarely explain all of the heterogeneity. Therefore, measuring and trying to understand residual heterogeneity is still important in a meta-regression, although it is not clear how some heterogeneity measures should be used in the meta-regression context. The coefficient of variation, and its variants, are useful measures of relative heterogeneity. We consider these measures in the context of meta-regression which allows researchers to investigate heterogeneity at different levels of the moderator and also average relative heterogeneity overall. We also provide CIs for the measures and our simulation studies show that these intervals have good coverage properties. We recommend that these measures and corresponding intervals could provide useful insights into moderators that may be contributing to the presence of heterogeneity in a meta-analysis and lead to a better understanding of estimated mean effects.

stat.ME↗

On choosing optimal response transformations for dimension reduction

It has previously been shown that response transformations can be very effective in improving dimension reduction outcomes for a continuous response. The choice of transformation used can make a big difference in the visualization of the response versus the dimension reduced regressors. In this article, we provide an automated approach for choosing parameters of transformation functions to seek optimal results. A criterion based on an influence measure between dimension reduction spaces is utilized for choosing the optimal parameter value of the transformation. Since influence measures can be time-consuming for large data sets, two efficient criteria are also provided. Given that a different transformation may be suitable for each direction required to form the subspace, we also employ an iterative approach to choosing optimal parameter values. Several simulation studies and a real data example highlight the effectiveness of the proposed methods.

stat.ME↗

Robust analogs to the Coefficient of Variation

The coefficient of variation (CV) is commonly used to measure relative dispersion. However, since it is based on the sample mean and standard deviation, outliers can adversely affect the CV. Additionally, for skewed distributions the mean and standard deviation do not have natural interpretations and, consequently, neither does the CV. Here we investigate the extent to which quantile-based measures of relative dispersion can provide appropriate summary information as an alternative to the CV. In particular, we investigate two measures, the first being the interquartile range (in lieu of the standard deviation), divided by the median (in lieu of the mean), and the second being the median absolute deviation (MAD), divided by the median, as robust estimators of relative dispersion. In addition to comparing the influence functions of the competing estimators and their asymptotic biases and variances, we compare interval estimators using simulation studies to assess coverage.

math.ST↗

An efficient estimator of the parameters of the Generalized Lambda Distribution

Estimation of the four generalized lambda distribution parameters is not straightforward, and available estimators that perform best have large computation times. In this paper, we introduce a simple two-step estimator of the parameters that is comparatively very quick to compute and performs well when compared with other methods. This computational efficiency makes the use of bootstrapping to obtain interval estimators for the parameters possible. Simulations are used to assess the performance of the new estimators and applications to several data sets are included.

stat.ME↗

Insights and inference for the proportion below the relative poverty line

We examine a commonly used relative poverty measure called the headcount ratio ($H_p$), defined to be the proportion of incomes falling below the relative poverty line, which is defined to be a fraction $p$ of the median income. We do this by considering this concept for theoretical income populations, and its potential for determining actual changes following transfer of incomes from the wealthy to those whose incomes fall below the relative poverty line. In the process we derive and evaluate the performance of large sample confidence intervals for $H_p$. Finally, we illustrate the estimators on real income data sets.

stat.ME↗

Mean skewness measures

Skewness measures can be used to measure the level of asymmetry of a distribution. Given the prevalence of statistical methods that assume underlying symmetry, and also the desire for symmetry in order to make meaningful judgements for common summary measures (e.g. the sample mean), reliably quantifying asymmetry is an important problem. There are several measures, among them generalizations of Bowley's well known skewness coefficient, that use sample quartiles and other quantile-based measures. The main drawbacks of many measures is that they are either limited to quartiles and do not take into account more extreme tail behavior, or that they require one to choose other quantiles (i.e. choose a value for $p$ different from 0.25) in place of the quartiles. Our objective is to (i) average the skewness measures over all $p$ and (ii) provide interval estimators for the new measure with good coverage properties. Our simulation results show that the interval estimators perform very well for all distributions considered.

math.ST↗

Influence functions for Linear Discriminant Analysis: Sensitivity analysis and efficient influence diagnostics

Whilst influence functions for linear discriminant analysis (LDA) have been found for a single discriminant when dealing with two groups, until now these have not been derived in the setting of a general number of groups. In this paper we explore the relationship between Sliced Inverse Regression (SIR) and LDA, and exploit this relationship to develop influence functions for LDA from those already derived for SIR. These influence functions can be used to understand robustness properties of LDA and also to detect influential observations in practice. We illustrate the usefulness of these via their application to a real data set.

math.ST↗

Interval estimators for inequality measures using grouped data

Income inequality measures are often used as an indication of economic health. How to obtain reliable confidence intervals for these measures based on sampled data has been studied extensively in recent years. To preserve confidentiality, income data is often made available in summary form only (i.e. histograms, frequencies between quintiles, etc.). In this paper, we show that good coverage can be achieved for bootstrap and Wald-type intervals for quantile-based measures when only grouped (binned) data are available. These coverages are typically superior to those that we have been able to achieve for intervals for popular measures such as the Gini index in this grouped data setting. To facilitate the bootstrapping, we use the Generalized Lambda Distribution and also a linear interpolation approximation method to approximate the underlying density. The latter is possible when groups means are available. We also apply our methods to real data sets.

stat.AP↗

Interval estimators for ratios of independent quantiles and interquantile ranges

Recent research has shown that interval estimators with good coverage properties are achievable for some functions of quantiles, even when sample sizes are not large. Motivated by this, we consider interval estimators for the ratios of independent quantiles and interquantile ranges that will be useful when comparing location and scale for two samples. Simulations show that the intervals have excellent coverage properties for a wide range of distributions, including those that are heavily skewed. Examples are also considered that highlight the usefulness of using these approaches to compare location and scale.

math.ST↗

Decomposing the Quantile Ratio Index with applications to Australian income and wealth data

The quantile ratio index introduced by Prendergast and Staudte 2017 is a simple and effective measure of relative inequality for income data that is resistant to outliers. It measures the average relative distance of a randomly chosen income from its symmetric quantile. Another useful property of this index is investigated here: given a partition of the income distribution into a union of sets of symmetric quantiles, one can find the conditional inequality for each set as measured by the quantile ratio index and readily combine them in a weighted average to obtain the index for the entire population. When applied to data for various years, one can track how these contributions to inequality vary over time, as illustrated here for Australian Bureau of Statistics income and wealth data.

stat.ME↗

Confidence Intervals for Quantiles from Histograms and Other Grouped Data

Interval estimation of quantiles has been treated by many in the literature. However, to the best of our knowledge there has been no consideration for interval estimation when the data are available in grouped format. Motivated by this, we introduce several methods to obtain confidence intervals for quantiles when only grouped data is available. Our preferred method for interval estimation is to approximate the underlying density using the Generalized Lambda Distribution (GLD) to both estimate the quantiles and variance of the quantile estimators. We compare the GLD method with some other methods that we also introduce which are based on a frequency approximation approach and a linear interpolation approximation of the density. Our methods are strongly supported by simulations showing that excellent coverage can be achieved for a wide number of distributions. These distributions include highly-skewed distributions such as the log-normal, Dagum and Singh-Maddala distributions. We also apply our methods to real data and show that inference can be carried out on published outcomes that have been summarized only by a histogram. Our methods are therefore useful for a broad range of applications. We have also created a web application that can be used to conveniently calculate the estimators.

stat.AP↗

A Simple and Effective Inequality Measure

Ratios of quantiles are often computed for income distributions as rough measures of inequality, and inference for such ratios have recently become available. The special case when the quantiles are symmetrically chosen; that is, when the p/2 quantile is divided by the (1-p/2), is of special interest because the graph of such ratios, plotted as a function of p over the unit interval, yields an informative inequality curve. The area above the curve and less than the horizontal line at one is an easily interpretable coefficient of inequality. The advantages of these concepts over the traditional Lorenz curve and Gini coefficient are numerous: they are defined for all positive income distributions, they can be robustly estimated and distribution-free confidence intervals for the inequality coefficient are easily found. Moreover the inequality curves satisfy a median-based transference principle and are convex for many commonly assumed income distributions.

stat.ME↗

Quantile Versions of the Lorenz Curve

The classical Lorenz curve is often used to depict inequality in a population of incomes, and the associated Gini coefficient is relied upon to make comparisons between different countries and other groups. The sample estimates of these moment-based concepts are sensitive to outliers and so we investigate the extent to which quantile-based definitions can capture income inequality and lead to more robust procedures. Distribution-free estimates of the corresponding coefficients of inequality are obtained, as well as sample sizes required to estimate them to a given accuracy. Convexity, transference and robustness of the measures are examined and illustrated.

math.ST↗