SearcharxivSearch

arXiv subjects

Davy Paindaveine

Publications and source records attributed to Davy Paindaveine.

At least 19 recordsLinked to original sources

M\"obius-Invariant Goodness-of-Fit Tests for the Spherical Cauchy Model

We introduce a class of goodness-of-fit tests for the spherical Cauchy model on the unit hypersphere. The proposed procedures exploit the invariance of the spherical Cauchy family under M\"obius transformations: after estimating the parameter by the sample M\"obius mean, the observations are transformed to approximate spherical uniformity, and a projection-based uniformity statistic is applied to the resulting sample. We show that, under mild conditions, the resulting tests are exactly distribution-free under the null hypothesis, so that exact critical values can be arbitrarily well approximated by Monte Carlo simulation. We study the M\"obius mean as a population functional, establish its existence and uniqueness under mild conditions, and prove equivariance and asymptotic linearity of its empirical counterpart, which coincides with the spherical Cauchy maximum likelihood estimator. We derive the asymptotic null distribution of the test statistic and show that it coincides with that of the underlying uniformity statistic after removing the degree-one spherical-harmonic component, which corresponds to the tangent space of the spherical Cauchy model. We establish consistency against fixed alternatives and characterize local powers through the spherical-harmonic decomposition of contiguous alternatives. Monte Carlo experiments demonstrate the finite-sample accuracy of the asymptotic approximations and the empirical power of the proposed tests. A real data example is treated.

math.ST

Stopping on the last success with unknown odds: asymptotic minimax optimality of the plug-in rule

We study the last-success problem for sequential Bernoulli trials in the homogeneous setting where $X_1,\ldots,X_n$ are i.i.d. $\mathrm{Bernoulli}(p)$ but the success probability $p\in(0,1)$ is unknown to the decision maker. When $p$ is known, Bruss' sum-the-odds theorem yields an optimal threshold rule with win probability $V_n(p)$; when it is not, the odds driving this threshold must be learned from the very sequence on which one is trying to stop, which turns the problem into a genuinely statistical decision problem over the class of $p$-blind rules---those depending on the data but not on $p$. Writing $W_n^\pi(p)$ for the win probability of such a rule, we show that, for any $p_0\in(0,\tfrac12)$, $$ \lim_{n\to\infty}\sqrt n\,\inf_\pi\sup_{p\in[p_0,1)}\bigl(V_n(p)-W_n^\pi(p)\bigr) = C_\star , $$ with $C_\star:=\tfrac12\sup_{u>0}u\Phi(-u)\approx0.085$ (here, $\Phi$ is the standard normal distribution function), and that this exact constant is attained by the natural plug-in odds rule, which is therefore asymptotically minimax optimal. The result is in fact local: at every transition point $1/k$ of the oracle threshold, the deficit admits an exact local minimax constant proportional to $\gamma_k=\tfrac{1}{\sqrt k}(1-\tfrac1k)^{k-3/2}$, and $C_\star$ is the largest of these, attained at $k=2$. We further quantify the cost of the natural sample-splitting alternative, show that the plug-in rule is asymptotically oracle-optimal in the sparse regime $p=p_n\to0$ with $np_n\to\infty$, and prove that no $p$-blind sequence of rules, even randomized, can converge to the oracle uniformly over $p\in(0,1)$, the obstruction being located in the critical window where $p$ is of order $1/n$.

math.PR

Win rates at first-passage times for biased simple random walks

We study the win rate $R_{N_d}/N_d$ of a biased simple random walk $S_n$ on $\mathbb{Z}$ at the first-passage time $N_d=\inf\{n\ge 0:S_n=d\}$, with $p=P[X_1=+1]\in[1/2,1)$. Using generating-function techniques and integral representations, we derive explicit formulas for the expectation and variance of $R_{N_d}/N_d$ along with monotonicity properties in the threshold $d$ and the bias $p$. We also provide closed-form expressions and use them to design unbiased coin-flipping estimators of $\pi$ based on first-passage sampling; the resulting schemes illustrate how biasing the coin can dramatically improve both approximation accuracy and computational cost.

math.PR

On the robustness of semi-discrete optimal transport

We derive the breakdown point for solutions of semi-discrete optimal transport problems, which characterizes the robustness of the multivariate quantiles based on optimal transport proposed in \cite{GS}. We do so under very mild assumptions: the absolutely continuous reference measure is only assumed to have a support that is \textcolor{mygreen}{convex}, whereas the target measure is a general discrete measure on a finite number, $n$ say, of atoms. The breakdown point depends on the target measure only through its probability weights (hence not on the location of the atoms) and involves the geometry of the reference measure through the \cite{Tuk1975} concept of halfspace depth. Remarkably, depending on this geometry, the breakdown point of the optimal transport median can be strictly smaller than the breakdown point of the univariate median or the breakdown point of the spatial median, namely~$\lceil n/2\rceil /2$. In the context of robust location estimation, our results provide a subtle insight on how to perform multivariate trimming when constructing trimmed means based on optimal transport.

math.PR

On a class of Sobolev tests for symmetry of directions, their detection thresholds, and asymptotic powers

We consider a class of symmetry hypothesis testing problems including testing isotropy on $\mathbb{R}^d$ and testing rotational symmetry on the hypersphere $\mathcal{S}^{d-1}$. For this class, we study the null and non-null behaviors of Sobolev tests, with emphasis on their consistency rates. Our main results show that: (i) Sobolev tests exhibit a detection threshold (see Bhattacharya, 2019, 2020) that does not only depend on the coefficients defining these tests; and (ii) tests with non-zero coefficients at odd (respectively, even) ranks only are blind to alternatives with angular functions whose $k$th-order derivatives at zero vanish for any $k$ odd (even). Our non-standard asymptotic results are illustrated with Monte Carlo exercises. A case study in astronomy applies the testing toolbox to evaluate the symmetry of orbits of long- and short-period comets.

math.ST

Revisiting the name variant of the two-children problem

Initially proposed by Martin Gardner in the 1950s, the famous two-children problem is often presented as a paradox in probability theory. A relatively recent variant of this paradox states that, while in a two-children family for which at least one child is a girl, the probability that the other child is a boy is $2/3$, this probability becomes $1/2$ if the first name of the girl is disclosed (provided that two sisters may not be given the same first name). We revisit this variant of the problem and show that, if one adopts a natural model for the way first names are given to girls, then the probability that the other child is a boy may take any value in $(0,2/3)$. By exploiting the concept of Schur-concavity, we study how this probability depends on model parameters.

math.PR

On the consistency of incomplete U-statistics under infinite second-order moments}

We derive a consistency result, in the $L_1$-sense, for incomplete U-statistics in the non-standard case where the kernel at hand has infinite second-order moments. Assuming that the kernel has finite moments of order $p(\geq 1)$, we obtain a bound on the $L_1$ distance between the incomplete U-statistic and its Dirac weak limit, which allows us to obtain, for any fixed $p$, an upper bound on the consistency rate. Our results hold for most classical sampling schemes that are used to obtain incomplete U-statistics.

math.ST

On optimal tests for rotational symmetry against new classes of hyperspherical distributions

Motivated by the central role played by rotationally symmetric distributions in directional statistics, we consider the problem of testing rotational symmetry on the hypersphere. We adopt a semiparametric approach and tackle problems where the location of the symmetry axis is either specified or unspecified. For each problem, we define two tests and study their asymptotic properties under very mild conditions. We introduce two new classes of directional distributions that extend the rotationally symmetric class and are of independent interest. We prove that each test is locally asymptotically maximin, in the Le Cam sense, for one kind of the alternatives given by the new classes of distributions, both for specified and unspecified symmetry axis. The tests, aimed to detect location-like and scatter-like alternatives, are combined into convenient hybrid tests that are consistent against both alternatives. We perform Monte Carlo experiments that illustrate the finite-sample performances of the proposed tests and their agreement with the asymptotic results. Finally, the practical relevance of our tests is illustrated on a real data application from astronomy. The R package rotasym implements the proposed tests and allows practitioners to reproduce the data application.

stat.ME

On the behavior of extreme $d$-dimensional spatial quantiles under minimal assumptions

"Spatial" or "geometric" quantiles are the only multivariate quantiles coping with both high-dimensional data and functional data, also in the framework of multiple-output quantile regression. This work studies spatial quantiles in the finite-dimensional case, where the spatial quantile $μ_{α,u}(P)$ of the distribution $P$ taking values in $\mathbb{R}^d $ is a point in $\mathbb{R}^d$ indexed by an order $α\in[0,1)$ and a direction $u$ in the unit sphere $\mathcal{S}^{d-1}$ of $\mathbb{R}^d$ --- or equivalently by a vector $αu$ in the open unit ball of $\mathbb{R}^d$. Recently, Girard and Stupfler (2017) proved that (i) the extreme quantiles $μ_{α,u}(P)$ obtained as $α\to 1$ exit all compact sets of $\mathbb{R}^d$ and that (ii) they do so in a direction converging to $u$. These results help understanding the nature of these quantiles: the first result is particularly striking as it holds even if $P$ has a bounded support, whereas the second one clarifies the delicate dependence of spatial quantiles on $u$. However, they were established under assumptions imposing that $P$ is non-atomic, so that it is unclear whether they hold for empirical probability measures. We improve on this by proving these results under much milder conditions, allowing for the sample case. This prevents using gradient condition arguments, which makes the proofs very challenging. We also weaken the well-known sufficient condition for uniqueness of finite-dimensional spatial quantiles.

math.ST

On the power of axial tests of uniformity on spheres

Testing uniformity on the $p$-dimensional unit sphere is arguably the most fundamental problem in directional statistics. In this paper, we consider this problem in the framework of axial data, that is, under the assumption that the $n$ observations at hand are randomly drawn from a distribution that charges antipodal regions equally. More precisely, we focus on axial, rotationally symmetric, alternatives and first address the problem under which the direction $θ$ of the corresponding symmetry axis is specified. In this setup, we obtain Le Cam optimal tests of uniformity, that are based on the sample covariance matrix (unlike their non-axial analogs, that are based on the sample average). For the more important unspecified-$θ$ problem, some classical tests are available in the literature, but virtually nothing is known on their non-null behavior. We therefore study the non-null behavior of the celebrated Bingham test and of other tests that exploit the single-spiked nature of the considered alternatives. We perform Monte Carlo exercises to investigate the finite-sample behavior of our tests and to show their agreement with our asymptotic results.

math.ST

Sign tests for weak principal directions

We consider inference on the first principal direction of a $p$-variate elliptical distribution. We do so in challenging double asymptotic scenarios for which this direction eventually fails to be identifiable. In order to achieve robustness not only with respect to such weak identifiability but also with respect to heavy tails, we focus on sign-based statistical procedures, that is, on procedures that involve the observations only through their direction from the center of the distribution. We actually consider the generic problem of testing the null hypothesis that the first principal direction coincides with a given direction of $\mathbb{R}^p$. We first focus on weak identifiability setups involving single spikes (that is, involving spectra for which the smallest eigenvalue has multiplicity $p-1$). We show that, irrespective of the degree of weak identifiability, such setups offer local alternatives for which the corresponding sequence of statistical experiments converges in the Le Cam sense. Interestingly, the limiting experiments depend on the degree of weak identifiability. We exploit this convergence result to build optimal sign tests for the problem considered. In classical asymptotic scenarios where the spectrum is fixed, these tests are shown to be asymptotically equivalent to the sign-based likelihood ratio tests available in the literature. Unlike the latter, however, the proposed sign tests are robust to arbitrarily weak identifiability. We show that our tests meet the asymptotic level constraint irrespective of the structure of the spectrum, hence also in possibly multi-spike setups. We fully characterize the non-null asymptotic distributions of the corresponding test statistics under weak identifiability, which allows us to quantify the corresponding local asymptotic powers.

math.ST

Preliminary test estimation in ULAN models

Preliminary test estimation, which is a natural procedure when it is suspected a priori that the parameter to be estimated might take value in a submodel of the model at hand, is a classical topic in estimation theory. In the present paper, we establish general results on the asymptotic behavior of preliminary test estimators. More precisely, we show that, in uniformly locally asymptotically normal (ULAN) models, a general asymptotic theory can be derived for preliminary test estimators based on estimators admitting generic Bahadur-type representations. This allows for a detailed comparison between classical estimators and preliminary test estimators in ULAN models. Our results, that, in standard linear regression models, are shown to reduce to some classical results, are also illustrated in more modern and involved setups, such as the multisample one where $m$ covariance matrices ${\pmbΣ}_1, \ldots, {\pmbΣ}_m$ are to be estimated when it is suspected that these matrices might be equal, might be proportional, or might share a common "scale". Simulation results confirm our theoretical findings.

math.ST

Inference for spherical location under high concentration

Motivated by the fact that circular or spherical data are often much concentrated around a location $\pmbθ$, we consider inference about $\pmbθ$ under "high concentration" asymptotic scenarios for which the probability of any fixed spherical cap centered at $\pmbθ$ converges to one as the sample size $n$ diverges to infinity. Rather than restricting to Fisher-von Mises-Langevin distributions, we consider a much broader, semiparametric, class of rotationally symmetric distributions indexed by the location parameter $\pmbθ$, a scalar concentration parameter $κ$ and a functional nuisance $f$. We determine the class of distributions for which high concentration is obtained as $κ$ diverges to infinity. For such distributions, we then consider inference (point estimation, confidence zone estimation, hypothesis testing) on $\pmbθ$ in asymptotic scenarios where $κ_n$ diverges to infinity at an arbitrary rate with the sample size $n$. Our asymptotic investigation reveals that, interestingly, optimal inference procedures on $\pmbθ$ show consistency rates that depend on $f$. Using asymptotics "à la Le Cam", we show that the spherical mean is, at any $f$, a parametrically super-efficient estimator of $\pmbθ$ and that the Watson and Wald tests for $\mathcal{H}_0:{\pmbθ}={\pmbθ}_0$ enjoy similar, non-standard, optimality properties. We illustrate our results through simulations and treat a real data example. On a technical point of view, our asymptotic derivations require challenging expansions of rotationally symmetric functionals for large arguments of $f$.

math.ST

From Halfspace M-depth to Multiple-output Expectile Regression

Despite the renewed interest in the Newey and Powell (1987) concept of expectiles in fields such as econometrics, risk management, and extreme value theory, expectile regression---or, more generally, M-quantile regression---unfortunately remains limited to single-output problems. To improve on this, we introduce hyperplane-valued multivariate M-quantiles that show strong advantages, for instance in terms of equivariance, over the various point-valued multivariate M-quantiles available in the literature. Like their competitors, our multivariate M-quantiles are directional in nature and provide centrality regions when all directions are considered. These regions define a new statistical depth, the halfspace M-depth, whose deepest point, in the expectile case, is the mean vector. Remarkably, the halfspace M-depth can alternatively be obtained by substituting, in the celebrated Tukey (1975) halfspace depth, M-quantile outlyingness for standard quantile outlyingness, which supports a posteriori the claim that our multivariate M-quantile concept is the natural one. We investigate thoroughly the properties of the proposed multivariate M-quantiles, of halfspace M-depth, and of the corresponding regions. Since our original motivation was to define multiple-output expectile regression methods, we further focus on the expectile case. We show in particular that expectile depth is smoother than the Tukey depth and enjoys interesting monotonicity properties that are extremely promising for computational purposes. Unlike their quantile analogs, the proposed multivariate expectiles also satisfy the coherency axioms of multivariate risk measures. Finally, we show that our multivariate expectiles indeed allow performing multiple-output expectile regression, which is illustrated on simulated and real data.

math.ST

Detecting the direction of a signal on high-dimensional spheres: Non-null and Le Cam optimality results

We consider one of the most important problems in directional statistics, namely the problem of testing the null hypothesis that the spike direction $θ$ of a Fisher-von Mises-Langevin distribution on the $p$-dimensional unit hypersphere is equal to a given direction $θ_0$. After a reduction through invariance arguments, we derive local asymptotic normality (LAN) results in a general high-dimensional framework where the dimension $p_n$ goes to infinity at an arbitrary rate with the sample size $n$, and where the concentration $κ_n$ behaves in a completely free way with $n$, which offers a spectrum of problems ranging from arbitrarily easy to arbitrarily challenging ones. We identify various asymptotic regimes, depending on the convergence/divergence properties of $(κ_n)$, that yield different contiguity rates and different limiting experiments. In each regime, we derive Le Cam optimal tests under specified $κ_n$ and we compute, from the Le Cam third lemma, asymptotic powers of the classical Watson test under contiguous alternatives. We further establish LAN results with respect to both spike direction and concentration, which allows us to discuss optimality also under unspecified $κ_n$. To investigate the non-null behavior of the Watson test outside the parametric framework above, we derive its local asymptotic powers through martingale CLTs in the broader, semiparametric, model of rotationally symmetric distributions. A Monte Carlo study shows that the finite-sample behaviors of the various tests remarkably agree with our asymptotic results.

math.ST

Testing for Principal Component Directions under Weak Identifiability

We consider the problem of testing, on the basis of a $p$-variate Gaussian random sample, the null hypothesis ${\cal H}_0: {\pmb θ}_1= {\pmb θ}_1^0$ against the alternative ${\cal H}_1: {\pmb θ}_1 \neq {\pmb θ}_1^0$, where ${\pmb θ}_1$ is the "first" eigenvector of the underlying covariance matrix and ${\pmb θ}_1^0$ is a fixed unit $p$-vector. In the classical setup where eigenvalues $λ_1>λ_2\geq \ldots\geq λ_p$ are fixed, the Anderson (1963) likelihood ratio test (LRT) and the Hallin, Paindaveine and Verdebout (2010) Le Cam optimal test for this problem are asymptotically equivalent under the null hypothesis, hence also under sequences of contiguous alternatives. We show that this equivalence does not survive asymptotic scenarios where $λ_{n1}/λ_{n2}=1+O(r_n)$ with $r_n=O(1/\sqrt{n})$. For such scenarios, the Le Cam optimal test still asymptotically meets the nominal level constraint, whereas the LRT severely overrejects the null hypothesis. Consequently, the former test should be favored over the latter one whenever the two largest sample eigenvalues are close to each other. By relying on the Le Cam's asymptotic theory of statistical experiments, we study the non-null and optimality properties of the Le Cam optimal test in the aforementioned asymptotic scenarios and show that the null robustness of this test is not obtained at the expense of power. Our asymptotic investigation is extensive in the sense that it allows $r_n$ to converge to zero at an arbitrary rate. While we restrict to single-spiked spectra of the form $λ_{n1}>λ_{n2}=\ldots=λ_{np}$ to make our results as striking as possible, we extend our results to the more general elliptical case. Finally, we present an illustrative real data example.

math.ST

Tyler shape depth

In many problems from multivariate analysis, the parameter of interest is a shape matrix, that is, a normalized version of the corresponding scatter or dispersion matrix. In this paper, we propose a depth concept for shape matrices that involves data points only through their directions from the center of the distribution. We use the terminology Tyler shape depth since the resulting estimator of shape, namely the deepest shape matrix, is the median-based counterpart of the M-estimator of shape of Tyler (1987). Beyond estimation, shape depth, like its Tyler antecedent, also allows hypothesis testing on shape. Its main benefit, however, lies in the ranking of shape matrices it provides, whose practical relevance is illustrated in principal component analysis and in shape-based outlier detection. We study the invariance, quasi-concavity and continuity properties of Tyler shape depth, the topological and boundedness properties of the corresponding depth regions, existence of a deepest shape matrix and prove Fisher consistency in the elliptical case. Finally, we derive a Glivenko-Cantelli-type result and establish almost sure consistency of the deepest shape matrix estimator.

math.ST

Distance-based Depths for Directional Data

Directional data are constrained to lie on the unit sphere of~$\mathbb{R}^q$ for some~$q\geq 2$. To address the lack of a natural ordering for such data, depth functions have been defined on spheres. However, the depths available either lack flexibility or are so computationally expensive that they can only be used for very small dimensions~$q$. In this work, we improve on this by introducing a class of distance-based depths for directional data. Irrespective of the distance adopted, these depths can easily be computed in high dimensions too. We derive the main structural properties of the proposed depths and study how they depend on the distance used. We discuss the asymptotic and robustness properties of the corresponding deepest points. We show the practical relevance of the proposed depths in two applications, related to (i) spherical location estimation and (ii) supervised classification. For both problems, we show through simulation studies that distance-based depths have strong advantages over their competitors.

math.ST