SearcharxivSearch

arXiv subjects

Andrey Akinshin

Publications and source records attributed to Andrey Akinshin.

9 recordsLinked to original sources

Quantile-respectful density estimation based on the Harrell-Davis quantile estimator

Traditional density and quantile estimators are often inconsistent with each other. Their simultaneous usage may lead to inconsistent results. To address this issue, we propose a novel smooth density estimator that is naturally consistent with the Harrell-Davis quantile estimator. We also provide a jittering implementation to support discrete-continuous mixture distributions.

stat.ME

Weighted quantile estimators

In this paper, we consider a generic scheme that allows building weighted versions of various quantile estimators, such as traditional quantile estimators based on linear interpolation of two order statistics, the Harrell-Davis quantile estimator and its trimmed modification. The obtained weighted quantile estimators are especially useful in the problem of estimating a distribution at the tail of a time series using quantile exponential smoothing. The presented approach can also be applied to other problems, such as quantile estimation of weighted mixture distributions.

stat.ME

Finite-sample Rousseeuw-Croux scale estimators

The Rousseeuw-Croux $S_n$, $Q_n$ scale estimators and the median absolute deviation $\operatorname{MAD}_n$ can be used as consistent estimators for the standard deviation under normality. All of them are highly robust: the breakdown point of all three estimators is $50\%$. However, $S_n$ and $Q_n$ are much more efficient than\ $\operatorname{MAD}_n$: their asymptotic Gaussian efficiency values are $58\%$ and $82\%$ respectively compared to $37\%$ for\ $\operatorname{MAD}_n$. Although these values look impressive, they are only asymptotic values. The actual Gaussian efficiency of $S_n$ and $Q_n$ for small sample sizes is noticeable lower than in the asymptotic case. The original work by Rousseeuw and Croux (1993) provides only rough approximations of the finite-sample bias-correction factors for $S_n$, $Q_n$ and brief notes on their finite-sample efficiency values. In this paper, we perform extensive Monte-Carlo simulations in order to obtain refined values of the finite-sample properties of the Rousseeuw-Croux scale estimators. We present accurate values of the bias-correction factors and Gaussian efficiency for small samples ($n \leq 100$) and prediction equations for samples of larger sizes.

stat.ME

Quantile absolute deviation

The median absolute deviation (MAD) is a popular robust measure of statistical dispersion. However, when it is applied to non-parametric distributions (especially multimodal, discrete, or heavy-tailed), lots of statistical inference issues arise. Even when it is applied to distributions with slight deviations from normality and these issues are not actual, the Gaussian efficiency of the MAD is only 37% which is not always enough. In this paper, we introduce the quantile absolute deviation (QAD) as a generalization of the MAD. This measure of dispersion provides a flexible approach to analyzing properties of non-parametric distributions. It also allows controlling the trade-off between robustness and statistical efficiency. We use the trimmed Harrell-Davis median estimator based on the highest density interval of the given width as a complimentary median estimator that gives increased finite-sample Gaussian efficiency compared to the sample median and a breakdown point matched to the QAD. As a rule of thumb, we suggest using two new measures of dispersion called the standard QAD and the optimal QAD. They give 54% and 65% of Gaussian efficiency having breakdown points of 32% and 14% respectively.

stat.ME

Trimmed Harrell-Davis quantile estimator based on the highest density interval of the given width

Traditional quantile estimators that are based on one or two order statistics are a common way to estimate distribution quantiles based on the given samples. These estimators are robust, but their statistical efficiency is not always good enough. A more efficient alternative is the Harrell-Davis quantile estimator which uses a weighted sum of all order statistics. Whereas this approach provides more accurate estimations for the light-tailed distributions, it's not robust. To be able to customize the trade-off between statistical efficiency and robustness, we could consider a trimmed modification of the Harrell-Davis quantile estimator. In this approach, we discard order statistics with low weights according to the highest density interval of the beta distribution.

stat.ME

Finite-sample bias-correction factors for the median absolute deviation based on the Harrell-Davis quantile estimator and its trimmed modification

The median absolute deviation is a widely used robust measure of statistical dispersion. Using a scale constant, we can use it as an asymptotically consistent estimator for the standard deviation under normality. For finite samples, the scale constant should be corrected in order to obtain an unbiased estimator. The bias-correction factor depends on the sample size and the median estimator. When we use the traditional sample median, the factor values are well known, but this approach does not provide optimal statistical efficiency. In this paper, we present the bias-correction factors for the median absolute deviation based on the Harrell-Davis quantile estimator and its trimmed modification which allow us to achieve better statistical efficiency of the standard deviation estimations. The obtained estimators are especially useful for samples with a small number of elements.

stat.ME

Geometry of error amplification in solving Prony system with near-colliding nodes

We consider a reconstruction problem for ``spike-train'' signals $F$ of an a priori known form $F(x)=\sum_{j=1}^{d}a_{j}δ\left(x-x_{j}\right),$ from their moments $m_k(F)=\int x^kF(x)dx.$ We assume that the moments $m_k(F)$, $k=0,1,\ldots,2d-1$, are known with an absolute error not exceeding $ε> 0$. This problem is essentially equivalent to solving the Prony system $\sum_{j=1}^d a_jx_j^k=m_k(F), \ k=0,1,\ldots,2d-1.$ We study the ``geometry of error amplification'' in reconstruction of $F$ from $m_k(F),$ in situations where the nodes $x_1,\ldots,x_d$ near-collide, i.e. form a cluster of size $h \ll 1$. We show that in this case, error amplification is governed by certain algebraic varieties in the parameter space of signals $F$, which we call the ``Prony varieties''. Based on this we produce lower and upper bounds, of the same order, on the worst case reconstruction error. In addition we derive separate lower and upper bounds on the reconstruction of the amplitudes and the nodes. Finally we discuss how to use the geometry of the Prony varieties to improve the reconstruction accuracy given additional a priori information.

math.CA

Accuracy of reconstruction of spike-trains with two near-colliding nodes

We consider a signal reconstruction problem for signals $F$ of the form $ F(x)=\sum_{j=1}^{d}a_{j}δ\left(x-x_{j}\right),$ from their moments $m_k(F)=\int x^kF(x)dx.$ We assume $m_k(F)$ to be known for $k=0,1,\ldots,N,$ with an absolute error not exceeding $ε> 0$. We study the "geometry of error amplification" in reconstruction of $F$ from $m_k(F),$ in situations where two neighboring nodes $x_i$ and $x_{i+1}$ near-collide, i.e $x_{i+1}-x_i=h \ll 1$. We show that the error amplification is governed by certain algebraic curves $S_{F,i},$ in the parameter space of signals $F$, along which the first three moments $m_0,m_1,m_2$ remain constant.

math.CA

Accuracy of spike-train Fourier reconstruction for colliding nodes

We consider Fourier reconstruction problem for signals F, which are linear combinations of shifted delta-functions. We assume the Fourier transform of F to be known on the frequency interval [-N,N], with an absolute error not exceeding e > 0. We give an absolute lower bound (which is valid with any reconstruction method) for the "worst case" reconstruction error of F in situations where the nodes (i.e. the positions of the shifted delta-functions in F) are known to form an l elements cluster of a size h << 1. Using "decimation" reconstruction algorithm we provide an upper bound for the reconstruction error, essentially of the same form as the lower one. Roughly, our main result states that for N*h of order of (2l-1)-st root of e the worst case reconstruction error of the cluster nodes is of the same order as h, and hence the inside configuration of the cluster nodes (in the worst case scenario) cannot be reconstructed at all. On the other hand, decimation algorithm reconstructs F with the accuracy of order of 2l-st root of e.

math.CA