SearcharxivSearch

arXiv subjects

Andreas Eberl

Publications and source records attributed to Andreas Eberl.

6 recordsLinked to original sources

Assessing Monotone Dependence: Area Under the Curve Meets Rank Correlation

The assessment of monotone dependence between random variables is a classical problem in statistics and a gamut of application domains. Consequently, researchers have sought measures of association that are invariant under strictly increasing transformations of the margins, with the extant literature being splintered. For continuous variables, symmetric rank correlation coefficients, such as Spearman's Rho and Kendall's Tau, have been studied at great length in the statistical literature. For dichotomous outcomes, the asymmetric area under the curve (AUC) measure is used to assess monotone dependence. We unify and complete thus far disconnected strands of literature, by establishing common population level theory, common estimators, and common tests that bridge continuous and dichotomous settings and apply to all linearly ordered outcomes. Originating in the biomedical literature, the C index provides a bridge between AUC, to which it reduces for a dichotomous outcome, and Kendall's Tau, to which it relates linearly under continuity. To establish the same kind of bridge between AUC and Spearman's Rho, we introduce asymmetric grade correlation, AGC$(X,Y)$, as the covariance of the mid distribution function transforms, or grades, of $X$ and $Y$, divided by the variance of the grade of $Y$. The coefficient of monotone association then is CMA$(X,Y) = \frac{1}{2} ($AGC$(X,Y) + 1)$. When $X$ and $Y$ are continuous, AGC is symmetric and equals Spearman's Rho. When $Y$ is dichotomous, CMA equals AUC. We establish central limit theorems for the sample versions of these measures, and we develop tests of DeLong type for their equality with a shared outcome $Y$. In case studies, we assess progress in data-driven weather prediction and evaluate methods of uncertainty quantification for large language models.

stat.ME

An easily verifiable dispersion order for discrete distributions

Dispersion is a fundamental concept in statistics, yet standard approaches - especially via stochastic orders - face limitations in the discrete setting. In particular, the classical dispersive order, well-established for continuous distributions, becomes overly restrictive for discrete random variables due to support inclusion requirements. To address this, we propose a novel weak dispersive order for discrete distributions. This order retains desirable properties while relaxing structural constraints, thereby broadening applicability. We further introduce a class of variability measures based on probability concentration, offering robust and interpretable alternatives that conform to classical axioms. Empirical illustrations highlight the practical relevance of this framework.

stat.ME

Defining Dispersion: A Fundamental Order for Univariate Discrete Distributions

The measurement of dispersion is one of the most fundamental and ubiquitous statistical concepts, in both applied and theoretical contexts. For dispersion measures, such as the standard deviation, to effectively capture the variability of a given distribution, they must, by definition, preserve some stochastic order of dispersion. The so-called dispersive order is the most basic order that serves as a foundation underneath the concept of dispersion measures. However, this order is incompatible with almost all discrete distributions, including lattice and most empirical distributions. As a result, popular measures may fail to accurately capture the dispersion of such distributions. In this paper, discrete adaptations of the dispersive order are defined and analyzed. They are shown to be a compromise between being equivalent to the original dispersive order on their joint area of applicability and other crucial properties. Moreover, they share many characteristic properties with the dispersive order, validating their role as a foundation for measuring discrete dispersion in a manner closely aligned with the continuous setting. Their behaviour on well-known families of lattice distribution is generally as expected when parameter differences are sufficiently large. Most popular dispersion measures preserve both discrete dispersive orders, rigorously ensuring that they are also meaningful in discrete settings. However, the interquantile range fails to preserve either discrete order, indicating that it is unsuitable for measuring the dispersion of discrete distributions.

stat.ME

Centre-free kurtosis orderings for asymmetric distributions

The concept of kurtosis is used to describe and compare theoretical and empirical distributions in a multitude of applications. In this connection, it is commonly applied to asymmetric distributions. However, there is no rigorous mathematical foundation establishing what is meant by kurtosis of an asymmetric distribution and what is required to measure it properly. All corresponding proposals in the literature centre the comparison with respect to kurtosis around some measure of central location. Since this either disregards critical amounts of information or is too restrictive, we instead revisit a canonical approach that has barely received any attention in the literature. It reveals the non-transitivity of kurtosis orderings due to an intrinsic entanglement of kurtosis and skewness as the underlying problem. This is circumvented by restricting attention to sets of distributions with equal skewness, on which the proposed kurtosis ordering is shown to be transitive. Moreover, we introduce a functional that preserves this order for arbitrary asymmetric distributions. As application, we examine the families of Weibull and sinh-arsinh distributions and show that the latter family exhibits a skewness-invariant kurtosis behaviour.

stat.ME

Stochastic orders and measures of skewness and dispersion based on expectiles

Recently, expectile-based measures of skewness akin to well-known quantile-based skewness measures have been introduced, and it has been shown that these measures possess quite promising properties (Eberl and Klar, 2021, 2020). However, it remained unanswered whether they preserve the convex transformation order of van Zwet, which is sometimes seen as a basic requirement for a measure of skewness. It is one of the aims of the present work to answer this question in the affirmative. These measures of skewness are scaled using interexpectile distances. We introduce orders of variability based on these quantities and show that the so-called weak expectile dispersive order is equivalent to the dilation order. Further, we analyze the statistical properties of empirical interexpectile ranges in some detail.

math.ST

Expectile based measures of skewness

In the literature, quite a few measures have been proposed for quantifying the deviation of a probability distribution from symmetry. The most popular of these skewness measures are based on the third centralized moment and on quantiles. However, there are major drawbacks in using these quantities. These include a strong emphasis on the distributional tails and a poor asymptotic behaviour for the (empirical) moment based measure as well as difficult statistical inference and strange behaviour for discrete distributions for quantile based measures. Therefore, in this paper, we introduce skewness measures based on or connected with expectiles. Since expectiles can be seen as smoothed versions of quantiles, they preserve the advantages over the moment based measure while not exhibiting most of the disadvantages of quantile based measures. We introduce corresponding empirical counterparts and derive asymptotic properties. Finally, we conduct a simulation study, comparing the newly introduced measures with established ones, and evaluating the performance of the respective estimators.

math.ST