SearcharxivSearch

arXiv subjects

Marc Hallin

Publications and source records attributed to Marc Hallin.

At least 19 recordsLinked to original sources

Quantiles for multivariate distribution: a short survey of population concepts

Quantiles are among the most fundamental concepts in Probability and Statistics, from descriptive to inferential. However, quantile functions are well-defined and well-understood in the context of one-dimensional probability distributions, where their definition as the inverse of distribution functions---a definition which is intimately related to the canonical ordering of the real line. Starting with dimension $d=2$, such a canonical ordering is no longer available in~${\mathbb R}^d$; nor is it available in nonlinear manifolds such as hyperspheres, tori, or polyspheres (vectors of directional variables). This results in the absence of an obvious and widely accepted definition of quantiles. The need to extend the concept of quantile beyond the classical univariate context nevertheless has sparked a large body of literature, giving rise to a variety of more or less satisfactory quantile concepts; the recent years have been particularly active in this respect. This short review is an attempt to complement and update Serfling's 25-year-old survey (Serfling 2008) by summarizing and unifying some of these concepts in light of recent contributions.

math.ST

Maximal Ancillarity, Semiparametric Efficiency, and the Elimination of Nuisances

Restricting statistical experiments via nuisance-ancillary $\sigma$-fields yields nuisance-free experiments. However, a moot point with ancillarity is that maximal ancillary $\sigma$-fields are typically not unique. There are exceptions, though, among which the limiting experiments in a locally asymptotically normal (LAN) context. Building on this, we address the maximal ancillarity uniqueness problem by adopting a H\'ajek-Le Cam asymptotic perspective and define the concept of sequences of locally asymptotically maximal nuisance-ancillary $\sigma$-fields. We then show that any semiparametrically efficient procedure admits versions that are measurable with respect to such $\sigma$-fields while enjoying strict finite-sample nuisance-ancillarity, hence eliminating the nuisance without the hassle of estimating it. This is in sharp contrast with classical tangent space projections, which also achieve semiparametric efficiency but only enjoy asymptotic nuisance-ancillarity -- at the price, moreover, of adequately estimating the nuisance. When the nuisance is the density of some noise or innovation driving the data-generating process of a LAN experiment, we show that a sequence of locally asymptotically maximal nuisance-ancillary $\sigma$-fields is generated by the so-called center-outward residual ranks and signs based on measure transportation results. Restricting local experiments to such $\sigma$-fields yields sequences of finite-sample nuisance-free (here, distribution-free) restrictions of the original local LAN experiments that nevertheless achieve the semiparametric efficiency bounds of the original ones.

math.ST

Nonparametric Vector Quantile Autoregression

Prediction is a key issue in time series analysis. Just as classical mean regression models, classical autoregressive methods, yielding L$^2$ point-predictions, provide rather poor predictive summaries; a much more informative approach is based on quantile (auto)regression, where the whole distribution of future observations conditional on the past is consistently recovered. Since their introduction by Koenker and Xiao in 2006, autoregressive quantile autoregression methods have become a popular and successful alternative to the traditional L$^2$ ones. Due to the lack of a widely accepted concept of multivariate quantiles, however, quantile autoregression methods so far have been limited to univariate time series. Building upon recent measure-transportation-based concepts of multivariate quantiles, we develop here a nonparametric vector quantile autoregressive approach to the analysis and prediction of (nonlinear as well as linear) multivariate time series.

math.ST

A model-free test of the time-reversibility of climate change processes

Time-reversibility is a crucial feature of many time series models, while time-irreversibility is the rule rather than the exception in real-life data. Testing the null hypothesis of time-reversibilty, therefore, should be an important step preliminary to the identification and estimation of most traditional time-series models. Existing procedures, however, mostly consist of testing necessary but not sufficient conditions, leading to under-rejection, or sufficient but non-necessary ones, which leads to over-rejection. Moreover, they generally are model-besed. In contrast, the copula spectrum studied by Goto et al. ($\textit{Ann. Statist.}$ 2022, $\textbf{50}$: 3563--3591) allows for a model-free necessary and sufficient time-reversibility condition. A test based on this copula-spectrum-based characterization has been proposed by authors. This paper illustrates the performance of this test, with an illustration in the analysis of climatic data.

stat.ME

Quantiles and Quantile Regression on Riemannian Manifolds: a measure-transportation-based approach

Increased attention has been given recently to the statistical analysis of variables with values on nonlinear manifolds. A natural but nontrivial problem in that context is the definition of quantile concepts. We are proposing a solution for compact Riemannian manifolds without boundaries; typical examples are polyspheres, hyperspheres, and toro\"{\i}dal manifolds equipped with their Riemannian metrics. Our concept of quantile function comes along with a concept of distribution function and, in the empirical case, ranks and signs. The absence of a canonical ordering is offset by resorting to the data-driven ordering induced by optimal transports. Theoretical properties, such as the uniform convergence of the empirical distribution and conditional (and unconditional) quantile functions and distribution-freeness of ranks and signs, are established. Statistical inference applications, from goodness-of-fit to distribution-free rank-based testing, are without number. Of particular importance is the case of quantile regression with directional or toro\"{\i}dal multiple output, which is given special attention in this paper. Extensive simulations are carried out to illustrate these novel concepts.

math.ST

The Dynamic, the Static, and the Weak: Factor models and the analysis of high-dimensional time series

Several fundamental and closely interconnected issues related to factor models are reviewed and discussed: dynamic versus static loadings, rate-strong versus rate-weak factors, the concept of weakly common component recently introduced by Gersing et al. (2023), the irrelevance of cross-sectional ordering and the assumption of cross-sectional exchangeability, the impact of undetected strong factors, and the problem of combining common and idiosyncratic forecasts. Conclusions all point to the advantages of the General Dynamic Factor Model approach of Forni et al. (2000) over the widely used Static Approximate Factor Model introduced by Chamberlain and Rothschild (1983).

econ.EM

Consistent Distribution Free Affine Invariant Tests for the Validity of Independent Component Models

We propose a family of tests of the validity of the assumptions underlying independent component analysis methods. The tests are formulated as L2-type procedures based on characteristic functions and involve weights; a proper choice of these weights and the estimation method for the mixing matrix yields consistent and affine-invariant tests. Due to the complexity of the asymptotic null distribution of the resulting test statistics, implementation is based on permutational and resampling strategies. This leads to distribution-free procedures regardless of whether these procedures are performed on the estimated independent components themselves or the componentwise ranks of their components. A Monte Carlo study involving various estimation methods for the mixing matrix, various weights, and a competing test based on distance covariance is conducted under the null hypothesis as well as under alternatives. A real-data application demonstrates the practical utility and effectiveness of the method.

stat.ME

Multivariate Quantiles: Geometric and Measure-Transportation-Based Contours

Quantiles are a fundamental concept in probability and theoretical statistics and a daily tool in their applications. While the univariate concept of quantiles is quite clear and well understood, its multivariate extension is more problematic. After half a century of continued efforts and many proposals, two concepts, essentially, are emerging: the so-called (relabeled) geometric quantiles, extending the characterization of univariate quantiles as minimizers of an L1 loss function involving the check functions, and the more recent center-outward quantiles based on measure transportation ideas. These two concepts yield distinct families of quantile regions and quantile contours. Our objective here is to present a comparison of their main theoretical properties and a numerical investigation of their differences.

math.ST

Dynamic Factor Models: a Genealogy

Dynamic factor models have been developed out of the need of analyzing and forecasting time series in increasingly high dimensions. While mathematical statisticians faced with inference problems in high-dimensional observation spaces were focusing on the so-called spiked-model-asymptotics, econometricians adopted an entirely and considerably more effective asymptotic approach, rooted in the factor models originally considered in psychometrics. The so-called dynamic factor model methods, in two decades, has grown into a wide and successful body of techniques that are widely used in central banks, financial institutions, economic and statistical institutes. The objective of this chapter is not an extensive survey of the topic but a sketch of its historical growth, with emphasis on the various assumptions and interpretations, and a family tree of its main variants.

econ.EM

On Bounded Completeness and the $L_1$-Denseness of Likelihood Ratios

The classical concept of bounded completeness and its relation to sufficiency and ancillarity play a fundamental role in unbiased estimation, unbiased testing, and the validity of inference in the presence of nuisance parameters. In this short note, we provide a direct proof of a little-known result by \cite{Far62} on a characterization of bounded completeness based on an $L^1$ denseness property of the linear span of likelihood ratios. As an application, we show that an experiment with infinite-dimensional observation space is boundedly complete iff suitably chosen restricted subexperiments with finite-dimensional observation spaces are.

math.ST

Monotone Measure-Preserving Maps in Hilbert Spaces: Existence, Uniqueness, and Stability

The contribution of this work is twofold. The first part deals with a Hilbert-space version of McCann's celebrated result on the existence and uniqueness of monotone measure-preserving maps: given two probability measures $\rm P$ and $\rm Q$ on a separable Hilbert space $\mathcal{H}$ where $\rm P$ does not give mass to "small sets" (namely, Lipschitz hypersurfaces), we show, without imposing any moment assumptions, that there exists a gradient of convex function $\nabla\psi$ pushing ${\rm P} $ forward to ${\rm Q}$. In case $\mathcal{H}$ is infinite-dimensional, ${\rm P}$-a.s. uniqueness is not guaranteed, though. If, however, ${\rm Q}$ is boundedly supported (a natural assumption in several statistical applications), then this gradient is ${\rm P}$ a.s. unique. In the second part of the paper, we establish stability results for transport maps in the sense of uniform convergence over compact "regularity sets". As a consequence, we obtain a central limit theorem for the fluctuations of the optimal quadratic transport cost in a separable Hilbert space.

math.PR

Nonparametric Measure-Transportation-Based Methods for Directional Data

This paper proposes various nonparametric tools based on measure transportation for directional data. We use optimal transports to define new notions of distribution and quantile functions on the hypersphere, with meaningful quantile contours and regions and closed-form formulas under the classical assumption of rotational symmetry. The empirical versions of our distribution functions enjoy the expected Glivenko-Cantelli property of traditional distribution functions. They provide fully distribution-free concepts of ranks and signs and define data-driven systems of (curvilinear) parallels and (hyper)meridians. Based on this, we also construct a universally consistent test of uniformity and a class of fully distribution-free and universally consistent tests for directional MANOVA which, in simulations, outperform all their existing competitors. A real-data example involving the analysis of sunspots concludes the paper.

math.ST

Center-Outward Multiple-Output Lorenz Curves and Gini Indices a measure transportation approach

Based on measure transportation ideas and the related concepts of center-outward quantile functions, we propose multiple-output center-outward generalizations of the traditional univariate concepts of Lorenz and concentration functions, and the related Gini and Kakwani coefficients. These new concepts have a natural interpretation, either in terms of contributions of central ("middle-class") regions to the expectation of some variable of interest, or in terms of the physical notions of work and energy, which sheds new light on the nature of economic and social inequalities. Importantly, the proposed concepts pave the way to statistically sound definitions, based on multiple variables, of quantiles and quantile regions, and the concept of "middle class," of high relevance in various socio-economic contexts.

math.ST

Center-outward Rank- and Sign-based VARMA Portmanteau Tests: Chitturi, Hosking, and Li--McLeod revisited

The pseudo-Gaussian portmanteau tests of Chitturi, Hosking, and Li and McLeod for VARMA models are revisited from a Le Cam perspective, providing a precise and more rigorous description of the asymptotic behavior of the multivariate portmanteau test statistic, which depends on the dimension $d$ of the observations, the number $m$ of lags involved, and the length $n$ of the observation period. Then, based on the concepts of center-outward ranks and signs recently developed (Hallin, del Barrio, Cuesta-Albertos, and Matr\' an, {\it Annals of Statistics} 49, 1139--1165, 2021), a class of multivariate rank- and sign-based portmanteau test statistics is proposed which, under the null hypothesis and under a broad family of innovation densities, can be approximated by an asymptotically chi-square variable. The asymptotic properties of these tests are derived; simulations demonstrate their advantages over their classical pseudo-Gaussian counterpart.

math.ST

Nonparametric Multiple-Output Center-Outward Quantile Regression

Based on the novel concept of multivariate center-outward quantiles introduced recently in Chernozhukov et al. (2017) and Hallin et al. (2021), we are considering the problem of nonparametric multiple-output quantile regression. Our approach defines nested conditional center-outward quantile regression contours and regions with given conditional probability content irrespective of the underlying distribution; their graphs constitute nested center-outward quantile regression tubes. Empirical counterparts of these concepts are constructed, yielding interpretable empirical regions and contours which are shown to consistently reconstruct their population versions in the Pompeiu-Hausdorff topology. Our method is entirely non-parametric and performs well in simulations including heteroskedasticity and nonlinear trends; its power as a data-analytic tool is illustrated on some real datasets.

stat.ME

The integrated copula spectrum

Frequency domain methods form a ubiquitous part of the statistical toolbox for time series analysis. In recent years, considerable interest has been given to the development of new spectral methodology and tools capturing dynamics in the entire joint distributions and thus avoiding the limitations of classical, $L^2$-based spectral methods. Most of the spectral concepts proposed in that literature suffer from one major drawback, though: their estimation requires the choice of a smoothing parameter, which has a considerable impact on estimation quality and poses challenges for statistical inference. In this paper, associated with the concept of copula-based spectrum, we introduce the notion of copula spectral distribution function or integrated copula spectrum. This integrated copula spectrum retains the advantages of copula-based spectra but can be estimated without the need for smoothing parameters. We provide such estimators, along with a thorough theoretical analysis, based on a functional central limit theorem, of their asymptotic properties. We leverage these results to test various hypotheses that cannot be addressed by classical spectral methods, such as the lack of time-reversibility or asymmetry in tail dynamics.

math.ST

Distribution-free tests of multivariate independence based on center-outward quadrant, Spearman, Kendall, and van der Waerden statistics

Due to the lack of a canonical ordering in ${\mathbb R}^d$ for $d>1$, defining multivariate generalizations of the classical univariate ranks has been a long-standing open problem in statistics. Optimal transport has been shown to offer a solution in which multivariate ranks are obtained by transporting data points to a grid that approximates a uniform reference measure (Chernozhukov et al., 2017; Hallin, 2017; Hallin et al., 2021), thereby inducing ranks, signs, and a data-driven ordering of ${\mathbb R}^d$. We take up this new perspective to define and study multivariate analogues of the sign covariance/quadrant statistic, Spearman's rho, Kendall's tau, and van der Waerden covariances. The resulting tests of multivariate independence are fully distribution-free, hence uniformly valid irrespective of the actual (absolutely continuous) distribution of the observations. Our results provide the asymptotic distribution theory for these new test statistics, with asymptotic approximations to critical values to be used for testing independence between random vectors, as well as a power analysis of the resulting tests in an extension of the so-called (bivariate) Konijn model. This power analysis includes a multivariate Chernoff--Savage property guaranteeing that, under elliptical generalized Konijn models, the asymptotic relative efficiency of our van der Waerden tests with respect to Wilks' classical (pseudo-)Gaussian procedure is strictly larger than or equal to one, where equality is achieved under Gaussian distributions only. We similarly provide a lower bound for the asymptotic relative efficiency of our Spearman procedure with respect to Wilks' test, thus extending the classical result by Hodges and Lehmann on the asymptotic relative efficiency, in univariate location models, of Wilcoxon tests with respect to the Student ones.

math.ST

On the Finite-Sample Performance of Measure Transportation-Based Multivariate Rank Tests

Extending to dimension 2 and higher the dual univariate concepts of ranks and quantiles has remained an open problem for more than half a century. Based on measure transportation results, a solution has been proposed recently under the name center-outward ranks and quantiles which, contrary to previous proposals, enjoys all the properties that make univariate ranks a successful tool for statistical inference. Just as their univariate counterparts (to which they reduce in dimension one), center-outward ranks allow for the construction of distribution-free and asymptotically efficient tests for a variety of problems where the density of some noise or innovation remains unspecified. The actual implementation of these tests involves the somewhat arbitrary choice of a grid. While the asymptotic impact of that choice is nil, its finite-sample consequences are not. In this note, we investigate the finite-sample impact of that choice in the typical context of the multivariate two-sample location problem.

math.ST