Searcharxiv⌕ Search

arXiv subjects

Marc Hallin

Publications and source records attributed to Marc Hallin.

At least 37 records · Page 2Linked to original sources

Fully distribution-free center-outward rank tests for multiple-output regression and MANOVA

Extending rank-based inference to a multivariate setting such as multiple-output regression or MANOVA with unspecified d-dimensional error density has remained an open problem for more than half a century. None of the many solutions proposed so far is enjoying the combination of distribution-freeness and efficiency that makes rank-based inference a successful tool in the univariate setting. A concept of center-outward multivariate ranks and signs based on measure transportation ideas has been introduced recently. Center-outward ranks and signs are not only distribution-free but achieve in dimension d > 1 the (essential) maximal ancillarity property of traditional univariate ranks, hence carry all the "distribution-free information" available in the sample. We derive here the Hájek representation and asymptotic normality results required in the construction of center-outward rank tests for multiple-output regression and MANOVA. When based on appropriate spherical scores, these fully distribution-free tests achieve parametric efficiency in the corresponding models.

math.ST↗

The integrated copula spectrum

Frequency domain methods form a ubiquitous part of the statistical toolbox for time series analysis. In recent years, considerable interest has been given to the development of new spectral methodology and tools capturing dynamics in the entire joint distributions and thus avoiding the limitations of classical, $L^2$-based spectral methods. Most of the spectral concepts proposed in that literature suffer from one major drawback, though: their estimation requires the choice of a smoothing parameter, which has a considerable impact on estimation quality and poses challenges for statistical inference. In this paper, associated with the concept of copula-based spectrum, we introduce the notion of copula spectral distribution function or integrated copula spectrum. This integrated copula spectrum retains the advantages of copula-based spectra but can be estimated without the need for smoothing parameters. We provide such estimators, along with a thorough theoretical analysis, based on a functional central limit theorem, of their asymptotic properties. We leverage these results to test various hypotheses that cannot be addressed by classical spectral methods, such as the lack of time-reversibility or asymmetry in tail dynamics.

math.ST↗

On the Finite-Sample Performance of Measure Transportation-Based Multivariate Rank Tests

Extending to dimension 2 and higher the dual univariate concepts of ranks and quantiles has remained an open problem for more than half a century. Based on measure transportation results, a solution has been proposed recently under the name center-outward ranks and quantiles which, contrary to previous proposals, enjoys all the properties that make univariate ranks a successful tool for statistical inference. Just as their univariate counterparts (to which they reduce in dimension one), center-outward ranks allow for the construction of distribution-free and asymptotically efficient tests for a variety of problems where the density of some noise or innovation remains unspecified. The actual implementation of these tests involves the somewhat arbitrary choice of a grid. While the asymptotic impact of that choice is nil, its finite-sample consequences are not. In this note, we investigate the finite-sample impact of that choice in the typical context of the multivariate two-sample location problem.

math.ST↗

On universally consistent and fully distribution-free rank tests of vector independence

Rank correlations have found many innovative applications in the last decade. In particular, suitable rank correlations have been used for consistent tests of independence between pairs of random variables. Using ranks is especially appealing for continuous data as tests become distribution-free. However, the traditional concept of ranks relies on ordering data and is, thus, tied to univariate observations. As a result, it has long remained unclear how one may construct distribution-free yet consistent tests of independence between random vectors. This is the problem addressed in this paper, in which we lay out a general framework for designing dependence measures that give tests of multivariate independence that are not only consistent and distribution-free but which we also prove to be statistically efficient. Our framework leverages the recently introduced concept of center-outward ranks and signs, a multivariate generalization of traditional ranks, and adopts a common standard form for dependence measures that encompasses many popular examples. In a unified study, we derive a general asymptotic representation of center-outward rank-based test statistics under independence, extending to the multivariate setting the classical Hájek asymptotic representation results. This representation permits direct calculation of limiting null distributions and facilitates a local power analysis that provides strong support for the center-outward approach by establishing, for the first time, the nontrivial power of center-outward rank-based tests over root-$n$ neighborhoods within the class of quadratic mean differentiable alternatives.

math.ST↗

Factor Models for High-Dimensional Functional Time Series

In this paper, we set up the theoretical foundations for a high-dimensional functional factor model approach in the analysis of large cross-sections (panels) of functional time series (FTS). We first establish a representation result stating that, under mild assumptions on the covariance operator of the cross-section, we can represent each FTS as the sum of a common component driven by scalar factors loaded via functional loadings, and a mildly cross-correlated idiosyncratic component. Our model and theory are developed in a general Hilbert space setting that allows for mixed panels of functional and scalar time series. We then turn to the identification of the number of factors, and the estimation of the factors, their loadings, and the common components. We provide a family of information criteria for identifying the number of factors, and prove their consistency. We provide average error bounds for the estimators of the factors, loadings, and common component; our results encompass the scalar case, for which they reproduce and extend, under weaker conditions, well-established similar results. Under slightly stronger assumptions, we also provide uniform bounds for the estimators of factors, loadings, and common component, thus extending existing scalar results. Our consistency results in the asymptotic regime where the number $N$ of series and the number $T$ of time observations diverge thus extend to the functional context the "blessing of dimensionality" that explains the success of factor models in the analysis of high-dimensional (scalar) time series. We provide numerical illustrations that corroborate the convergence rates predicted by the theory, and provide finer understanding of the interplay between $N$ and $T$ for estimation purposes. We conclude with an application to forecasting mortality curves, where we demonstrate that our approach outperforms existing methods.

math.ST↗

Multivariate goodness-of-Fit tests based on Wasserstein distance

Goodness-of-fit tests based on the empirical Wasserstein distance are proposed for simple and composite null hypotheses involving general multivariate distributions. For group families, the procedure is to be implemented after preliminary reduction of the data via invariance.This property allows for calculation of exact critical values and p-values at finite sample sizes. Applications include testing for location--scale families and testing for families arising from affine transformations, such as elliptical distributions with given standard radial density and unspecified location vector and scatter matrix. A novel test for multivariate normality with unspecified mean vector and covariance matrix arises as a special case. For more general parametric families, we propose a parametric bootstrap procedure to calculate critical values. The lack of asymptotic distribution theory for the empirical Wasserstein distance means that the validity of the parametric bootstrap under the null hypothesis remains a conjecture. Nevertheless, we show that the test is consistent against fixed alternatives. To this end, we prove a uniform law of large numbers for the empirical distribution in Wasserstein distance, where the uniformity is over any class of underlying distributions satisfying a uniform integrability condition but no additional moment assumptions. The calculation of test statistics boils down to solving the well-studied semi-discrete optimal transport problem. Extensive numerical experiments demonstrate the practical feasibility and the excellent performance of the proposed tests for the Wasserstein distance of order p = 1 and p = 2 and for dimensions at least up to d = 5. The simulations also lend support to the conjecture of the asymptotic validity of the parametric bootstrap.

stat.ME↗

Rank-Based Testing for Semiparametric VAR Models: a measure transportation approach

We develop a class of tests for semiparametric vector autoregressive (VAR) models with unspecified innovation densities, based on the recent measure-transportation-based concepts of multivariate {\it center-outward ranks} and {\it signs}. We show that these concepts, combined with Le Cam's asymptotic theory of statistical experiments, yield novel testing procedures, which (a)~are valid under a broad class of innovation densities (possibly non-elliptical, skewed, and/or with infinite moments), (b)~are optimal (locally asymptotically maximin or most stringent) at selected ones, and (c) are robust against additive outliers. In order to do so, we establish a H\' ajek asymptotic representation result, of independent interest, for a general class of center-outward rank-based serial statistics. As an illustration, we consider the problems of testing the absence of serial correlation in multiple-output and possibly non-linear regression (an extension of the classical Durbin-Watson problem) and the sequential identification of the order $p$ of a vector autoregressive (VAR($p$)) model. A Monte Carlo comparative study of our tests and their routinely-applied Gaussian competitors demonstrates the benefits (in terms of size, power, and robustness) of our methodology; these benefits are particularly significant in the presence of asymmetric and leptokurtic innovation densities. A real data application concludes the paper.

math.ST↗

Center-Outward R-Estimation for Semiparametric VARMA Models

We propose a new class of R-estimators for semiparametric VARMA models in which the innovation density plays the role of the nuisance parameter. Our estimators are based on the novel concepts of multivariate center-outward ranks and signs. We show that these concepts, combined with Le Cam's asymptotic theory of statistical experiments, yield a class of semiparametric estimation procedures, which are efficient (at a given reference density), root-$n$ consistent, and asymptotically normal under a broad class of (possibly non elliptical) actual innovation densities. No kernel density estimation is required to implement our procedures. A Monte Carlo comparative study of our R-estimators and other routinely-applied competitors demonstrates the benefits of the novel methodology, in large and small sample. Proofs, computational aspects, and further numerical results are available in the supplementary material.

math.ST↗

Center-Outward Distribution Functions, Quantiles, Ranks, and Signs in $\mathbb{R}^d$

Univariate concepts as quantile and distribution functions involving ranks and signs, do not canonically extend to $\mathbb{R}^d, d\geq 2$. Palliating that has generated an abundant literature. Chapter 1 shows that, unlike the many definitions that have been proposed so far, the measure transportation-based ones introduced in Chernozhukov et al. (2017) enjoy all the properties that make univariate quantiles and ranks successful tools for semiparametric statistical inference. We therefore propose a new center-outward definition of multivariate distribution and quantile functions, along with their empirical counterparts, for which we obtain a Glivenko-Cantelli result. Our approach is geometric and, contrary to the Monge-Kantorovich one in Chernozhukov et al. (2017), does not require any moment assumptions. The resulting ranks and signs are strictly distribution-free, and maximal invariant under the action of a data-driven class of (order-preserving) transformations generating the family of absolutely continuous distributions; that property is the theoretical foundation of the semiparametric efficiency preservation property of ranks. The corresponding quantiles are equivariant under the same transformations. The empirical proposed distribution functions are defined at observed values only. A continuous extension to the entire $\mathbb{R}^d$, yielding continuous empirical quantile contours while preserving the monotonicity and Glivenko-Cantelli features is desirable. Such extension requires solving a nontrivial problem of smooth interpolation under cyclical monotonicity constraints. A complete solution of that problem is given in Chapter 2; we show that the resulting distribution and quantile functions are Lipschitz, and provide a sharp lower bound for the Lipschitz constants. A numerical study of empirical center-outward quantile contours and their consistency is conducted.

stat.ME↗

A note on the Regularity of Center-Outward Distribution and Quantile Functions

We provide sufficient conditions under which the center-outward distribution and quantile functions introduced in Chernozhukov et al.~(2017) and Hallin~(2017) are homeomorphisms, thereby extending a recent result by Figalli \cite{Fi2}. Our approach relies on Cafarelli's classical regularity theory for the solutions of the Monge-Ampère equation, but has to deal with difficulties related with the unboundedness at the origin of the density of the spherical uniform reference measure. Our conditions are satisfied by probabillities on Euclidean space with a general (bounded or unbounded) convex support which are not covered in~\cite{Fi2}. We provide some additional results about center-outward distribution and quantile functions, including the fact that quantile sets exhibit some weak form of convexity.

math.ST↗

Center-outward quantiles and the measurement of multivariate risk

All multivariate extensions of the univariate theory of risk measurement run into the same fundamental problem of the absence, in dimension d > 1, of a canonical ordering of Rd. Based on measure transportation ideas, several attempts have been made recently in the statistical literature to overcome that conceptual difficulty. In Hallin (2017), the concepts of center-outward distribution and quantile functions are developed as generalisations of the classical univariate concepts of distribution and quantile functions, along with their empirical versions. We propose a class of smooth approximations as an alternative to the interpolation developed in del Barrio et al. (2018). This approximation allows for the computation of some new empirical risk measures, based either on the convex potential associated with the proposed transports, or on the volumes of the resulting empirical quantile regions. We also discuss the role of such transports in the evaluation of the risk associated with multivariate regularly varying distributions. Some simulations and applications to case studies illustrate the value of the approach.

stat.ME↗

Optimal tests for elliptical symmetry: specified and unspecified location

Although the assumption of elliptical symmetry is quite common in multivariate analysis and widespread in a number of applications, the problem of testing the null hypothesis of ellipticity so far has not been addressed in a fully satisfactory way. Most of the literature in the area indeed addresses the null hypothesis of elliptical symmetry with specified location and actually addresses location rather than non-elliptical alternatives. In this paper, we are proposing new classes of testing procedures, both for specified and unspecified location. The backbone of our construction is Le Cam's asymptotic theory of statistical experiments, and optimality is to be understood locally and asymptotically within the family of generalized skew-elliptical distributions. The tests we are proposing are meeting all the desired properties of a``good'' test of elliptical symmetry: they have a simple asymptotic distribution under the entire null hypothesis of elliptical symmetry with unspecified radial density and shape parameter; they are affine-invariant, computationally fast, intuitively understandable, and not too demanding in terms of moments. While achieving optimality against generalized skew-elliptical alternatives, they remain quite powerful under a much broader class of non-elliptical distributions and significantly outperform the available competitors.

stat.ME↗

Generalized Dynamic Factor Models and Volatilities: Consistency, rates, and prediction intervals

Volatilities, in high-dimensional panels of economic time series with a dynamic factor structure on the levels or returns, typically also admit a dynamic factor decomposition. We consider a two-stage dynamic factor model method recovering the common and idiosyncratic components of both levels and log-volatilities. Specifically, in a first estimation step, we extract the common and idiosyncratic shocks for the levels, from which a log-volatility proxy is computed. In a second step, we estimate a dynamic factor model, which is equivalent to a multiplicative factor structure for volatilities, for the log-volatility panel. By exploiting this two-stage factor approach, we build one-step-ahead conditional prediction intervals for large $n \times T$ panels of returns. Those intervals are based on empirical quantiles, not on conditional variances; they can be either equal- or unequal- tailed. We provide uniform consistency and consistency rates results for the proposed estimators as both $n$ and $T$ tend to infinity. We study the finite-sample properties of our estimators by means of Monte Carlo simulations. Finally, we apply our methodology to a panel of asset returns belonging to the S&P100 index in order to compute one-step-ahead conditional prediction intervals for the period 2006-2013. A comparison with the componentwise GARCH benchmark (which does not take advantage of cross-sectional information) demonstrates the superiority of our approach, which is genuinely multivariate (and high-dimensional), nonparametric, and model-free.

econ.EM↗

Quantile Spectral Analysis for Locally Stationary Time Series

Classical spectral methods are subject to two fundamental limitations: they only can account for covariance-related serial dependencies, and they require second-order stationarity. Much attention has been devoted lately to quantile-based spectral methods that go beyond covariance-based serial dependence features. At the same time, covariance-based methods relaxing stationarity into much weaker {\it local stationarity} conditions have been developed for a variety of time-series models. Here, we are combining those two approaches by proposing quantile-based spectral methods for locally stationary processes. We therefore introduce a time-varying version of the copula spectra that have been recently proposed in the literature, along with a suitable local lag-window estimator. We propose a new definition of local {\it strict} stationarity that allows us to handle completely general non-linear processes without any moment assumptions, thus accommodating our quantile-based concepts and methods. We establish a central limit theorem for the new estimators, and illustrate the power of the proposed methodology by means of a simulation study. Moreover, in two empirical studies (namely of the Standard \& Poor's 500 series and a temperature dataset recorded in Hohenpeissenberg) we demonstrate that the new approach detects important variations in serial dependence structures both across time and across quantiles. Such variations remain completely undetected, and are actually undetectable, via classical covariance-based spectral methods.

math.ST↗

Networks, Dynamic Factors, and the Volatility Analysis of High-Dimensional Financial Series

We consider weighted directed networks for analysing, over the period 2000-2013, the interdependencies between volatilities of a large panel of stocks belonging to the S\&P100 index. In particular, we focus on the so-called {\it Long-Run Variance Decomposition Network} (LVDN), where the nodes are stocks, and the weight associated with edge $(i,j)$ represents the proportion of $h$-step-ahead forecast error variance of variable $i$ accounted for by variable $j$'s innovations. To overcome the curse of dimensionality, we decompose the panel into a component driven by few global, market-wide, factors, and an idiosyncratic one modelled by means of a sparse vector autoregression (VAR) model. Inversion of the VAR together with suitable identification restrictions, produces the estimated network, by means of which we can assess how {\it systemic} each firm is.~Our analysis demonstrates the prominent role of financial firms as sources of contagion, especially during the~2007-2008 crisis.

q-fin.ST↗

Quantile spectral processes: Asymptotic analysis and inference

Quantile- and copula-related spectral concepts recently have been considered by various authors. Those spectra, in their most general form, provide a full characterization of the copulas associated with the pairs $(X_t,X_{t-k})$ in a process $(X_t)_{t\in\mathbb{Z}}$, and account for important dynamic features, such as changes in the conditional shape (skewness, kurtosis), time-irreversibility, or dependence in the extremes that their traditional counterparts cannot capture. Despite various proposals for estimation strategies, only quite incomplete asymptotic distributional results are available so far for the proposed estimators, which constitutes an important obstacle for their practical application. In this paper, we provide a detailed asymptotic analysis of a class of smoothed rank-based cross-periodograms associated with the copula spectral density kernels introduced in Dette et al. [Bernoulli 21 (2015) 781-831]. We show that, for a very general class of (possibly nonlinear) processes, properly scaled and centered smoothed versions of those cross-periodograms, indexed by couples of quantile levels, converge weakly, as stochastic processes, to Gaussian processes. A first application of those results is the construction of asymptotic confidence intervals for copula spectral density kernels. The same convergence results also provide asymptotic distributions (under serially dependent observations) for a new class of rank-based spectral methods involving the Fourier transforms of rank-based serial statistics such as the Spearman, Blomqvist or Gini autocovariance coefficients.

math.ST↗

Monge-Kantorovich Depth, Quantiles, Ranks, and Signs

We propose new concepts of statistical depth, multivariate quantiles, ranks and signs, based on canonical transportation maps between a distribution of interest on $R^d$ and a reference distribution on the $d$-dimensional unit ball. The new depth concept, called Monge-Kantorovich depth, specializes to halfspace depth in the case of spherical distributions, but, for more general distributions, differs from the latter in the ability for its contours to account for non convex features of the distribution of interest. We propose empirical counterparts to the population versions of those Monge-Kantorovich depth contours, quantiles, ranks and signs, and show their consistency by establishing a uniform convergence property for empirical transport maps, which is of independent interest.

math.ST↗