SearcharxivSearch

arXiv subjects

Bernhard Klar

Publications and source records attributed to Bernhard Klar.

At least 19 recordsLinked to original sources

Circular Expectiles

In this work, we introduce circular expectiles as minimizers of an asymmetric circular loss function based on chord distance. In contrast to the linear expectile criterion, the resulting circular optimization problem is non-convex, so existence and uniqueness require a separate analysis. The construction extends linear expectiles to directional data while preserving the circular mean as the symmetric case corresponding to $\alpha=1/2$. We derive basic representations of the objective function and the associated identification function, and give a geometric interpretation that generalizes the corresponding representation for the circular mean. Furthermore, we prove the existence and uniqueness of the minimizers for distributions with positive density on the circle. The empirical circular expectile is defined by using the sample circular mean as reference direction for the induced linear order on the circle. We prove the uniqueness of the empirical expectile, as well as its consistency and finite-dimensional asymptotic normality. Finally, we indicate possible applications to circular measures of dispersion, skewness, and symmetry diagnostics.

stat.ME

An easily verifiable dispersion order for discrete distributions

Dispersion is a fundamental concept in statistics, yet standard approaches - especially via stochastic orders - face limitations in the discrete setting. In particular, the classical dispersive order, well-established for continuous distributions, becomes overly restrictive for discrete random variables due to support inclusion requirements. To address this, we propose a novel weak dispersive order for discrete distributions. This order retains desirable properties while relaxing structural constraints, thereby broadening applicability. We further introduce a class of variability measures based on probability concentration, offering robust and interpretable alternatives that conform to classical axioms. Empirical illustrations highlight the practical relevance of this framework.

stat.ME

Nonparametric estimation on the circle based on Fej\'er polynomials

This paper presents a comprehensive study of nonparametric estimation techniques on the circle using Fej\'er polynomials, which are analogues of Bernstein polynomials for periodic functions. Building upon Fej\'er's uniform approximation theorem, the paper introduces circular density and distribution function estimators based on Fej\'er kernels. It establishes their theoretical properties, including uniform strong consistency and asymptotic expansions. Since the estimation of the distribution function on the circle depends on the choice of the origin, we propose a data-dependent method to address this issue. The proposed methods are extended to account for measurement errors by incorporating classical and Berkson error models, adjusting the Fej\'er estimator to mitigate their effects. Simulation studies analyze the finite-sample performance of these estimators under various scenarios, including mixtures of circular distributions and measurement error models. An application to rainfall data demonstrates the practical application of the proposed estimators, demonstrating their robustness and effectiveness in the presence of rounding-induced Berkson errors.

stat.ME

On the application of Jammalamadaka-Jim\'enez Gamero-Meintanis test for circular regression model assessment

We study a circular-circular multiplicative regression model, characterized by an angular error distribution assumed to be wrapped Cauchy. We propose a specification procedure for this model, focusing on adapting a recently proposed goodness-of-fit test for circular distributions. We derive its limiting properties and study the power performance of the test through extensive simulations, including the adaptation of some other well-known goodness-of-fit tests for this type of data. To emphasize the practical relevance of our methodology, we apply it to several small real-world datasets and wind direction measurements in the Black Forest region of southwestern Germany, demonstrating the power and versatility of the presented approach.

stat.ME

Defining Dispersion: A Fundamental Order for Univariate Discrete Distributions

The measurement of dispersion is one of the most fundamental and ubiquitous statistical concepts, in both applied and theoretical contexts. For dispersion measures, such as the standard deviation, to effectively capture the variability of a given distribution, they must, by definition, preserve some stochastic order of dispersion. The so-called dispersive order is the most basic order that serves as a foundation underneath the concept of dispersion measures. However, this order is incompatible with almost all discrete distributions, including lattice and most empirical distributions. As a result, popular measures may fail to accurately capture the dispersion of such distributions. In this paper, discrete adaptations of the dispersive order are defined and analyzed. They are shown to be a compromise between being equivalent to the original dispersive order on their joint area of applicability and other crucial properties. Moreover, they share many characteristic properties with the dispersive order, validating their role as a foundation for measuring discrete dispersion in a manner closely aligned with the continuous setting. Their behaviour on well-known families of lattice distribution is generally as expected when parameter differences are sufficiently large. Most popular dispersion measures preserve both discrete dispersive orders, rigorously ensuring that they are also meaningful in discrete settings. However, the interquantile range fails to preserve either discrete order, indicating that it is unsuitable for measuring the dispersion of discrete distributions.

stat.ME

Robust performance metrics for imbalanced classification problems

We show that established performance metrics in binary classification, such as Matthews' correlation coefficient (MCC), Cohen's $\kappa$, the F-score or the Jaccard similarity coefficient are not robust to class imbalance in the sense that if the proportion of the minority class tends to $0$, the true positive rate (TPR) of the Bayes classifier under these metrics tends to $0$ as well. Thus, in imbalanced classification problems, these metrics favour classifiers which ignore the minority class. To alleviate this issue we introduce robustified modifications of the MCC, of Cohen's $\kappa$ and of the F-score with an additional tuning parameter which allows to adapt the amount of robustness against class imbalance. As theoretical guarantee we show that the Bayes-optimal classifier for these robustified performance metrics, when expressed in terms of the density ratio $f_1/f_0$ of the class-conditional densities $f_i$, has a threshold parameter which is upper-bounded in terms of the tuning parameters. Therefore, even in strongly imbalanced settings, the TPR associated to this classifier will be bounded away from $0$. We numerically illustrate the behaviour of the various performance metrics and the effect of the tuning parameters in simulations as well as on a credit default data set. We also discuss connections to the receiver operating characteristic and precision-recall curves, which provide an alternative perspective on the proposed notion of robustness, and give recommendations on how to combine their usage with performance metrics.

stat.ML

A Pareto tail plot without moment restrictions

We propose a mean functional which exists for any probability distributions, and which characterizes the Pareto distribution within the set of distributions with finite left endpoint. This is in sharp contrast to the mean excess plot which is not meaningful for distributions without existing mean, and which has a nonstandard behaviour if the mean is finite, but the second moment does not exist. The construction of the plot is based on the so called principle of a single huge jump, which differentiates between distributions with moderately heavy and super heavy tails. We present an estimator of the tail function based on $U$-statistics and study its large sample properties. The use of the new plot is illustrated by several loss datasets.

stat.ME

A gamma tail statistic and its asymptotics

Asmussen and Lehtomaa [Distinguishing log-concavity from heavy tails. Risks 5(10), 2017] introduced an interesting function $g$ which is able to distinguish between log-convex and log-concave tail behaviour of distributions, and proposed a randomized estimator for $g$. In this paper, we show that $g$ can also be seen as a tool to detect gamma distributions or distributions with gamma tail. We construct a more efficient estimator $\hat{g}_n$ based on $U$-statistics, propose several estimators of the (asymptotic) variance of $\hat{g}_n$, and study their performance by simulations. Finally, the methods are applied to several data sets of daily precipitation.

stat.ME

Lancaster correlation -- a new dependence measure linked to maximum correlation

We suggest novel correlation coefficients which equal the maximum correlation for a class of bivariate Lancaster distributions while being only slightly smaller than maximum correlation for a variety of further bivariate distributions. In contrast to maximum correlation, however, our correlation coefficients allow for rank and moment-based estimators which are simple to compute and have tractable asymptotic distributions. Confidence intervals resulting from these asymptotic approximations and the covariance bootstrap show good finite-sample coverage. In a simulation, the power of asymptotic as well as permutation tests for independence based on our correlation measures compares favorably with competing methods based on distance correlation or rank coefficients for functional dependence, among others. Moreover, for the bivariate normal distribution, our correlation coefficients equal the absolute value of the Pearson correlation, an attractive feature for practitioners which is not shared by various competitors. We illustrate the practical usefulness of our methods in applications to two real data sets.

stat.ME

Using Proxies to Improve Forecast Evaluation

Comparative evaluation of forecasts of statistical functionals relies on comparing averaged losses of competing forecasts after the realization of the quantity $Y$, on which the functional is based, has been observed. Motivated by high-frequency finance, in this paper we investigate how proxies $\tilde Y$ for $Y$ - say volatility proxies - which are observed together with $Y$ can be utilized to improve forecast comparisons. We extend previous results on robustness of loss functions for the mean to general moments and ratios of moments, and show in terms of the variance of differences of losses that using proxies will increase the power in comparative forecast tests. These results apply both to testing conditional as well as unconditional dominance. Finally, we numerically illustrate the theoretical results, both for simulated high-frequency data as well as for high-frequency log returns of several cryptocurrencies.

stat.ME

Centre-free kurtosis orderings for asymmetric distributions

The concept of kurtosis is used to describe and compare theoretical and empirical distributions in a multitude of applications. In this connection, it is commonly applied to asymmetric distributions. However, there is no rigorous mathematical foundation establishing what is meant by kurtosis of an asymmetric distribution and what is required to measure it properly. All corresponding proposals in the literature centre the comparison with respect to kurtosis around some measure of central location. Since this either disregards critical amounts of information or is too restrictive, we instead revisit a canonical approach that has barely received any attention in the literature. It reveals the non-transitivity of kurtosis orderings due to an intrinsic entanglement of kurtosis and skewness as the underlying problem. This is circumvented by restricting attention to sets of distributions with equal skewness, on which the proposed kurtosis ordering is shown to be transitive. Moreover, we introduce a functional that preserves this order for arbitrary asymmetric distributions. As application, we examine the families of Weibull and sinh-arsinh distributions and show that the latter family exhibits a skewness-invariant kurtosis behaviour.

stat.ME

Stochastic orders and measures of skewness and dispersion based on expectiles

Recently, expectile-based measures of skewness akin to well-known quantile-based skewness measures have been introduced, and it has been shown that these measures possess quite promising properties (Eberl and Klar, 2021, 2020). However, it remained unanswered whether they preserve the convex transformation order of van Zwet, which is sometimes seen as a basic requirement for a measure of skewness. It is one of the aims of the present work to answer this question in the affirmative. These measures of skewness are scaled using interexpectile distances. We introduce orders of variability based on these quantities and show that the so-called weak expectile dispersive order is equivalent to the dilation order. Further, we analyze the statistical properties of empirical interexpectile ranges in some detail.

math.ST

Cauchy or not Cauchy? New goodness-of-fit tests for the Cauchy distribution

We introduce a new characterization of the Cauchy distribution and propose a class of goodness-of-fit tests to the Cauchy family. The limit distribution is derived in a Hilbert space framework under the null hypothesis and under fixed alternatives. The new tests are consistent against a large class of alternatives. A comparative Monte Carlo simulation study shows that the test is competitive to the state of the art procedures, and we apply the tests to log-returns of cryptocurrencies.

math.ST

Smooth Distribution Function Estimation for Lifetime Distributions using Szasz-Mirakyan Operators

In this paper, we introduce a new smooth estimator for continuous distribution functions on the positive real half-line using Szasz-Mirakyan operators, similar to Bernstein's approximation theorem. We show that the proposed estimator outperforms the empirical distribution function in terms of asymptotic (integrated) mean-squared error, and generally compares favourably with other competitors in theoretical comparisons. Also, we conduct the simulations to demonstrate the finite sample performance of the proposed estimator.

math.ST

Minimum $L^q$-distance estimators for non-normalized parametric models

We propose and investigate a new estimation method for the parameters of models consisting of smooth density functions on the positive half axis. The procedure is based on a recently introduced characterization result for the respective probability distributions, and is to be classified as a minimum distance estimator, incorporating as a distance function the $L^q$-norm. Throughout, we deal rigorously with issues of existence and measurability of these implicitly defined estimators. Moreover, we provide consistency results in a common asymptotic setting, and compare our new method with classical estimators for the exponential-, the Rayleigh-, and the Burr Type XII distribution in Monte Carlo simulation studies. We also assess the performance of different estimators for non-normalized models in the context of an exponential-polynomial family.

math.ST

Expectile based measures of skewness

In the literature, quite a few measures have been proposed for quantifying the deviation of a probability distribution from symmetry. The most popular of these skewness measures are based on the third centralized moment and on quantiles. However, there are major drawbacks in using these quantities. These include a strong emphasis on the distributional tails and a poor asymptotic behaviour for the (empirical) moment based measure as well as difficult statistical inference and strange behaviour for discrete distributions for quantile based measures. Therefore, in this paper, we introduce skewness measures based on or connected with expectiles. Since expectiles can be seen as smoothed versions of quantiles, they preserve the advantages over the moment based measure while not exhibiting most of the disadvantages of quantile based measures. We introduce corresponding empirical counterparts and derive asymptotic properties. Finally, we conduct a simulation study, comparing the newly introduced measures with established ones, and evaluating the performance of the respective estimators.

math.ST

Weighted scoring rules and hypothesis testing

We discuss weighted scoring rules for forecast evaluation and their connection to hypothesis testing. First, a general construction principle for strictly locally proper weighted scoring rules based on conditional densities and scoring rules for probability forecasts is proposed. We show how likelihood-based weighted scoring rules from the literature fit into this framework, and also introduce a weighted version of the Hyvärinen score, which is a local scoring rule in the sense that it only depends on the forecast density and its derivatives at the observation, and does not require evaluation of integrals. Further, we discuss the relation to hypothesis testing. Using a weighted scoring rule introduces a censoring mechanism, in which the form of the density is irrelevant outside the region of interest. For the resulting testing problem with composite null - and alternative hypotheses, we construct optimal tests, and identify the associated weighted scoring rule. As a practical consequence, using a weighted scoring rule allows to decide in favor of a forecast which is superior to a competing forecast on a region of interest, even though it may be inferior outside this region. A simulation study and an application to financial time-series data illustrate these findings.

stat.ME

Expectile Asymptotics

We discuss in detail the asymptotic distribution of sample expectiles. First, we show uniform consistency under the assumption of a finite mean. In case of a finite second moment, we show that for expectiles other then the mean, only the additional assumption of continuity of the distribution function at the expectile implies asymptotic normality, otherwise, the limit is non-normal. For a continuous distribution function we show the uniform central limit theorem for the expectile process. If, in contrast, the distribution is heavy-tailed, and contained in the domain of attraction of a stable law with $1 < α< 2$, then we show that the expectile is also asymptotically stable distributed. Our findings are illustrated in a simulation section.

stat.ME