Searcharxiv⌕ Search

arXiv subjects

Estate V. Khmaladze

Publications and source records attributed to Estate V. Khmaladze.

8 recordsLinked to original sources

On the statistical analysis of grouped data: when Pearson $χ^2$ and other divisible statistics are not goodness-of-fit tests

Thousands of experiments are analyzed, and papers are published each year involving the statistical analysis of grouped data. While this area of statistics is often perceived -- somewhat naively -- as saturated, several misconceptions still affect everyday practice, and new frontiers have so far remained unexplored. Researchers must be aware of the limitations affecting their analyses and what new possibilities are at their hands. The article introduces a unifying approach to the analysis of divisible statistics -- that includes Pearson's $χ^2$, the likelihood ratio, and spectral statistics, as special cases -- when a statistician deals with a large number of bins/groups, thus leading to a large number of small or moderate frequencies. Performance of the tests is analyzed against the class of contiguous (local) alternatives. Perhaps the most surprising result here is that, in this `sparse' regime, most of the tests proposed in the literature can be modified to produce more powerful tests, and no single test based on a divisible statistic leads to a goodness-of-fit test. Distribution-free goodness-of-fit tests are also constructed.

stat.ME↗

Goodness-of-Fit Testing for Point Processes in Large Populations

Suppose we have an observed path from a point process counting event occurrences in a large population. Based on the observed path, we would like to test the null hypothesis that the conditional intensity of the point process belongs to a particular parametric family. We propose a novel approach to conducting such goodness-of-fit tests. The idea is to construct a unitary transformation of a natural parametric testing process such that it converges weakly to a ``standard'' target process, independent of the particular parametric form assumed under the null hypothesis. This transformation therefore paves the way for asymptotically distribution-free goodness-of-fit testing of parametric point processes. We demonstrate the good finite-sample performance of our approach through Monte Carlo simulations of Aalen-type survival processes, without and with censoring, mixture cure models, and software reliability models, and we illustrate its applicability with observed human lifetimes as well as real software failures.

math.ST↗

Distribution free testing for linear regression. Extension to general parametric regression

Recently a distribution free approach for testing parametric hypotheses based on unitary transformations has been suggested in \cite{Khm13, Khm16, Khm17} and further studied in \cite{Ngu17} and \cite{Rob19}. In this note we show that the transformation takes extremely simple form in distribution free testing of linear regression. Then we extend it to general parametric regression with vector-valued covariates.

stat.ME↗

Asymptotic hypothesis testing for the colour blind problem

In the classical two-sample problem, the conventional approach for testing distributions equality is based on the difference between the two marginal empirical distribution functions, whereas a test for independence is based on the contrast between the bivariate and the product of the marginal empirical distribution functions. In this article we consider the problem of testing independence and distributions equality when the observer is "colour blind" so he cannot distinguish the distribution which has generated each of the two measurements. Within a nonparametric framework, we propose an empirical process for this problem and find the linear statistic which is asymptotically optimal for testing the equality of the marginal distributions against a specific form of contiguous alternatives.

math.ST↗

Asymptotically distribution-free goodness-of-fit testing for tail copulas

Let $(X_1,Y_1),\ldots,(X_n,Y_n)$ be an i.i.d. sample from a bivariate distribution function that lies in the max-domain of attraction of an extreme value distribution. The asymptotic joint distribution of the standardized component-wise maxima $\bigvee_{i=1}^nX_i$ and $\bigvee_{i=1}^nY_i$ is then characterized by the marginal extreme value indices and the tail copula $R$. We propose a procedure for constructing asymptotically distribution-free goodness-of-fit tests for the tail copula $R$. The procedure is based on a transformation of a suitable empirical process derived from a semi-parametric estimator of $R$. The transformed empirical process converges weakly to a standard Wiener process, paving the way for a multitude of asymptotically distribution-free goodness-of-fit tests. We also extend our results to the $m$-variate ($m>2$) case. In a simulation study we show that the limit theorems provide good approximations for finite samples and that tests based on the transformed empirical process have high power.

math.ST↗

Differentiation of sets - The general case

In recent work by Khmaladze and Weil (2008) and by Einmahl and Khmaladze (2011), limit theorems were established for local empirical processes near the boundary of compact convex sets $K$ in $\R$. The limit processes were shown to live on the normal cylinder $Σ$ of $K$, respectively on a class of set-valued derivatives in $Σ$. The latter result was based on the concept of differentiation of sets at the boundary $\partial K$ of $K$, which was developed in Khmaladze (2007). Here, we extend the theory of set-valued derivatives to boundaries $\partial F$ of rather general closed sets $F\subset \R$, making use of a local Steiner formula for closed sets, established in Hug, Last and Weil (2004).

math.CA↗

Goodness-of-fit problem for errors in nonparametric regression: Distribution free approach

This paper discusses asymptotically distribution free tests for the classical goodness-of-fit hypothesis of an error distribution in nonparametric regression models. These tests are based on the same martingale transform of the residual empirical process as used in the one sample location model. This transformation eliminates extra randomization due to covariates but not due the errors, which is intrinsically present in the estimators of the regression function. Thus, tests based on the transformed process have, generally, better power. The results of this paper are applicable as soon as asymptotic uniform linearity of nonparametric residual empirical process is available. In particular they are applicable under the conditions stipulated in recent papers of Akritas and Van Keilegom and Müller, Schick and Wefelmeyer.

math.ST↗

Martingale transforms goodness-of-fit tests in regression models

This paper discusses two goodness-of-fit testing problems. The first problem pertains to fitting an error distribution to an assumed nonlinear parametric regression model, while the second pertains to fitting a parametric regression model when the error distribution is unknown. For the first problem the paper contains tests based on a certain martingale type transform of residual empirical processes. The advantage of this transform is that the corresponding tests are asymptotically distribution free. For the second problem the proposed asymptotically distribution free tests are based on innovation martingale transforms. A Monte Carlo study shows that the simulated level of the proposed tests is close to the asymptotic level for moderate sample sizes.

math.ST↗