SearcharxivSearch

arXiv subjects

Philip T. Labo

Publications and source records attributed to Philip T. Labo.

3 recordsLinked to original sources

The Asymptotics of Wide Remedians

The remedian uses a $k\times b$ matrix to approximate the median of $n\leq b^{k}$ streaming input values by recursively replacing buffers of $b$ values with their medians, thereby ignoring its $200(\lceil b/2\rceil / b)^{k}%$ most extreme inputs. Rousseeuw & Bassett (1990) and Chao & Lin (1993); Chen & Chen (2005) study the remedian's distribution as $k\rightarrow\infty$ and as $k,b\rightarrow\infty$. The remedian's breakdown point vanishes as $k\rightarrow\infty$, but approaches $(1/2)^{k}$ as $b\rightarrow\infty$. We study the remedian's robust-regime distribution as $b\rightarrow\infty$, deriving a normal distribution for standardized (mean, median, remedian, remedian rank) as $b\rightarrow\infty$, thereby illuminating the remedian's accuracy in approximating the sample median. We derive the asymptotic efficiency of the remedian relative to the mean and the median. Finally, we discuss the estimation of more than one quantile at once, proposing an asymptotic distribution for the random vector that results when we apply remedian estimation in parallel to the components of i.i.d. random vectors.

math.ST

Using Exponential Histograms to Approximate the Quantiles of Heavy- and Light-Tailed Data

Exponential histograms, with bins of the form $\left\{ \left(\rho^{k-1},\rho^{k}\right]\right\} _{k\in\mathbb{Z}}$, for $\rho>1$, straightforwardly summarize the quantiles of streaming data sets (Masson et al. 2019). While they guarantee the relative accuracy of their estimates, they appear to use only $\log n$ values to summarize $n$ inputs. We study four aspects of exponential histograms -- size, accuracy, occupancy, and largest gap size -- when inputs are i.i.d. $\mathrm{Exp}\left(\lambda\right)$ or i.i.d. $\mathrm{Pareto}\left(\nu,\beta\right)$, taking $\mathrm{Exp}\left(\lambda\right)$ (or, $\mathrm{Pareto}\left(\nu,\beta\right)$) to represent all light- (or, heavy-) tailed distributions. We show that, in these settings, size grows like $\log n$ and takes on a Gumbel distribution as $n$ grows large. We bound the missing mass to the right of the histogram and the mass of its final bin and show that occupancy grows apace with size. Finally, we approximate the size of the largest number of consecutive, empty bins. Our study gives a deeper and broader view of this low-memory approach to quantile estimation.

math.ST

Rank Distributions for Independent Normals with a Single Outlier

Thurstone's latent-normal model, introduced a century ago to describe human preferences in psychometrics (1927), remains a cornerstone for modeling random rankings. Yet when the underlying normals differ in distribution, the joint law of ranks $R_{i}:=\sum_{j=1}^{n}\mathbf{1}_{X_{j}\leq X_{i}}$ is virtually unexplored. We study the simplest non-identically-distributed case: $n+1$ independent normals with $X_{0}\sim\mathcal{N}\left(\mu_{0},\,\sigma_{0}^{2}\right)$ and $X_{i}\sim\mathcal{N}\left(\mu,\,\sigma^{2}\right)$ for $1\leq i\leq n$. Here, $R_0 \mid X_0 \;\sim\; 1 + \mathrm{Binomial}\bigl(n,\;\Phi\bigl(\bigl(X_0 - \mu\bigr)\big/\sigma\bigr)\bigr)$, and the success probability $\Phi\bigl(\bigl(X_0 - \mu\bigr)\big/\sigma\bigr)$ is accurately modeled by a beta distribution. Exploiting beta-binomial conjugacy, we observe that $R_{0}-1$ follows a beta-binomial law, which then yields a precise approximation for the joint distribution of $\left(R_{0},R_{i_{1}},\ldots,R_{i_{m}}\right)$. We derive closed-form expressions for $\mathbb{E}R_{i}$, $\mathrm{Cov}\left(R_{i},R_{j}\right)$, and the limiting distributions of $\left(R_{0},R_{i_{1}},\ldots,R_{i_{m}}\right)$ as key parameters grow large or small.

math.ST