SearcharxivSearch

arXiv subjects

Ivan Kojadinovic

Publications and source records attributed to Ivan Kojadinovic.

At least 19 recordsLinked to original sources

An iterated $I$-projection procedure for solving the generalized minimum information checkerboard copula problem

The minimum information copula principle initially suggested in \cite{MeeBed97} is a maximum entropy-like approach for finding the least informative copula, if it exists, that satisfies a certain number of expectation constraints specified either from domain knowledge or the available data. We first propose a generalization of this principle allowing the inclusion of additional constraints fixing certain higher-order margins of the copula. We next show that the associated optimization problem has a unique solution under a natural condition. As the latter problem is intractable in general we consider its version with all the probability measures involved in its formulation replaced by checkerboard approximations. This amounts to attempting to solve a so-called discrete $I$-projection linear problem. We then exploit the seminal results of \cite{Csi75} to derive an iterated procedure for solving the latter and provide theoretical guarantees for its convergence. The usefulness of the procedure is finally illustrated via numerical experiments in dimensions up to four with substantially finer discretizations than those encountered in the literature.

math.PR

The empirical discrete copula process

This paper develops a general inferential framework for discrete copulas on finite supports in any dimension. The copula of a multivariate discrete distribution is defined as Csiszar's I-projection (i.e., the minimum-Kullback-Leibler divergence projection) of its joint probability array onto the polytope of uniform-margins probability arrays of the same size, and its empirical estimator is obtained by applying that same projection to the array of empirical frequencies observed on the sample. Under the assumption of random sampling, strong consistency and root-n-asymptotic normality of the empirical copula array is established, with an explicit "sandwich" form for its covariance. The theory is illustrated by deriving the large-sample distribution of Yule's concordance coefficient (the natural analogue of Spearman's rho for bivariate discrete distributions) and by constructing a test for quasi-independence in multivariate contingency tables. Our results not only complete the foundations of discrete-copula inference but also connect directly to entropically regularised optimal transport and other minimum-divergence problems.

math.ST

On the differentiability of $ϕ$-projections in the discrete finite case

In the case of finite measures on finite spaces, we state conditions under which ϕ- projections are continuously differentiable. When the set on which one wishes to ϕ- project is convex, we show that the required assumptions are implied by easily verifiable conditions. In particular, for input probability vectors and a rather large class of ϕ-divergences, we obtain that ϕ-projections are continuously differentiable when projecting on a set defined by linear equalities. The obtained results are applied to ϕ- projection estimators (that is, minimum ϕ-divergence estimators). A first application, rooted in robust statistics, concerns the computation of the influence functions of such estimators. In a second set of applications, we derive their asymptotics when projecting on parametric sets of probability vectors, on sets of probability vectors generated from distributions with certain moments fixed and on Fréchet classes of bivariate probability arrays. The resulting asymptotics hold whether the element to be ϕ-projected belongs to the set on which one wishes to ϕ-project or not.

math.ST

Copula-like inference for discrete bivariate distributions with rectangular supports

After reviewing a large body of literature on the modeling of bivariate discrete distributions with finite support, \cite{Gee20} made a compelling case for the use of $I$-projections in the sense of \cite{Csi75} as a sound way to attempt to decompose a bivariate probability mass function (p.m.f.) into its two univariate margins and a bivariate p.m.f.\ with uniform margins playing the role of a discrete copula. From a practical perspective, the necessary $I$-projections on Fréchet classes can be carried out using the iterative proportional fitting procedure (IPFP), also known as Sinkhorn's algorithm or matrix scaling in the literature. After providing conditions under which a bivariate p.m.f.\ can be decomposed in the aforementioned sense, we investigate, for starting bivariate p.m.f.s with rectangular supports, nonparametric and parametric estimation procedures as well as goodness-of-fit tests for the underlying discrete copula. Related asymptotic results are provided and build upon a differentiability result for $I$-projections on Fréchet classes which can be of independent interest. Theoretical results are complemented by finite-sample experiments and a data example.

stat.ME

Subsampling (weighted smooth) empirical copula processes

A key tool to carry out inference on the unknown copula when modeling a continuous multivariate distribution is a nonparametric estimator known as the empirical copula. One popular way of approximating its sampling distribution consists of using the multiplier bootstrap. The latter is however characterized by a high implementation cost. Given the rank-based nature of the empirical copula, the classical empirical bootstrap of Efron does not appear to be a natural alternative, as it relies on resamples which contain ties. The aim of this work is to investigate the use of subsampling in the aforementioned framework. The latter consists of basing the inference on statistic values computed from subsamples of the initial data. One of its advantages in the rank-based context under consideration is that the formed subsamples do not contain ties. Another advantage is its asymptotic validity under minimalistic conditions. In this work, we show the asymptotic validity of subsampling for several (weighted, smooth) empirical copula processes both in the case of serially independent observations and time series. In the former case, subsampling is observed to be substantially better than the empirical bootstrap and equivalent, overall, to the multiplier bootstrap in terms of finite-sample performance.

math.ST

Resampling techniques for a class of smooth, possibly data-adaptive empirical copulas

We investigate the validity of two resampling techniques when carrying out inference on the underlying unknown copula using a recently proposed class of smooth, possibly data-adaptive nonparametric estimators that contains empirical Bernstein copulas (and thus the empirical beta copula). Following \cite{KirSegTsu21}, the first resampling technique is based on drawing samples from the smooth estimator and can only can be used in the case of independent observations. The second technique is a smooth extension of the so-called sequential dependent multiplier bootstrap and can thus be used in a time series setting and, possibly, for change-point analysis. The two studied resampling schemes are applied to confidence interval construction and the offline detection of changes in the cross-sectional dependence of multivariate time series, respectively. Monte Carlo experiments confirm the possible advantages of such smooth inference procedures over their non-smooth counterparts. A by-product of this work is the study of the weak consistency and finite-sample performance of two classes of smooth estimators of the first-order partial derivatives of a copula which can have applications in mean and quantile regression.

math.ST

A class of smooth, possibly data-adaptive nonparametric copula estimators containing the empirical beta copula

A broad class of smooth, possibly data-adaptive nonparametric copula estimators that contains empirical Bernstein copulas introduced by Sancetta and Satchell (and thus the empirical beta copula proposed by Segers, Sibuya and Tsukahara) is studied. Within this class, a subclass of estimators that depend on a scalar parameter determining the amount of marginal smoothing and a functional parameter controlling the shape of the smoothing region is specifically considered. Empirical investigations of the influence of these parameters suggest to focus on two particular data-adaptive smooth copula estimators that were found to be uniformly better than the empirical beta copula in all of the considered Monte Carlo experiments. Finally, with future applications to change-point detection in mind, conditions under which related sequential empirical copula processes converge weakly are provided.

math.ST

Multi-purpose open-end monitoring procedures for multivariate observations based on the empirical distribution function

We propose nonparametric open-end sequential testing procedures that can detect all types of changes in the contemporary distribution function of possibly multivariate observations. Their asymptotic properties are theoretically investigated under stationarity and under alternatives to stationarity. Monte Carlo experiments reveal their good finite-sample behavior in the case of continuous univariate, bivariate and trivariate observations. A short data example concludes the work.

stat.ME

On Stute's representation for a class of smooth, possibly data-adaptive empirical copula processes

Given a random sample from a continuous multivariate distribution, Stute's representation is obtained for empirical copula processes constructed from a broad class of smooth, possibly data-adaptive nonparametric copula estimators. The latter class contains for instance empirical Bernstein copulas introduced by Sancetta and Satchell and thus the empirical beta copula proposed by Segers, Sibuya and Tsukahara. The almost sure rate in Stute's representation is expressed in terms of a parameter controlling the speed at which the spread of the smoothing region decreases as the sample size increases.

math.ST

Nonparametric sequential change-point detection for multivariate time series based on empirical distribution functions

The aim of sequential change-point detection is to issue an alarm when it is thought that certain probabilistic properties of the monitored observations have changed. This work is concerned with nonparametric, closed-end testing procedures based on differences of empirical distribution functions that are designed to be particularly sensitive to changes in the comtemporary distribution of multivariate time series. The proposed detectors are adaptations of statistics used in a posteriori (offline) change-point testing and involve a weighting allowing to give more importance to recent observations. The resulting sequential change-point detection procedures are carried out by comparing the detectors to threshold functions estimated through resampling such that the probability of false alarm remains approximately constant over the monitoring period. A generic result on the asymptotic validity of such a way of estimating a threshold function is stated. As a corollary, the asymptotic validity of the studied sequential tests based on empirical distribution functions is proven when these are carried out using a dependent multiplier bootstrap for multivariate time series. Large-scale Monte Carlo experiments demonstrate the good finite-sample properties of the resulting procedures. The application of the derived sequential tests is illustrated on financial data.

stat.ME

Open-end nonparametric sequential change-point detection based on the retrospective CUSUM statistic

The aim of online monitoring is to issue an alarm as soon as there is significant evidence in the collected observations to suggest that the underlying data generating mechanism has changed. This work is concerned with open-end, nonparametric procedures that can be interpreted as statistical tests. The proposed monitoring schemes consist of computing the so-called retrospective CUSUM statistic (or minor variations thereof) after the arrival of each new observation. After proposing suitable threshold functions for the chosen detectors, the asymptotic validity of the procedures is investigated in the special case of monitoring for changes in the mean, both under the null hypothesis of stationarity and relevant alternatives. To carry out the sequential tests in practice, an approach based on an asymptotic regression model is used to estimate high quantiles of relevant limiting distributions. Monte Carlo experiments demonstrate the good finite-sample behavior of the proposed monitoring schemes and suggest that they are superior to existing competitors as long as changes do not occur at the very beginning of the monitoring. Extensions to statistics exhibiting an asymptotic mean-like behavior are briefly discussed. Finally, the application of the derived sequential change-point detection tests is succinctly illustrated on temperature anomaly data.

math.ST

Combining cumulative sum change-point detection tests for assessing the stationarity of univariate time series

We derive tests of stationarity for univariate time series by combining change-point tests sensitive to changes in the contemporary distribution with tests sensitive to changes in the serial dependence. The proposed approach relies on a general procedure for combining dependent tests based on resampling. After proving the asymptotic validity of the combining procedure under the conjunction of null hypotheses and investigating its consistency, we study rank-based tests of stationarity by combining cumulative sum change-point tests based on the contemporary empirical distribution function and on the empirical autocopula at a given lag. Extensions based on tests solely focusing on second-order characteristics are proposed next. The finite-sample behaviors of all the derived statistical procedures for assessing stationarity are investigated in large-scale Monte Carlo experiments and illustrations on two real data sets are provided. Extensions to multivariate time series are briefly discussed as well.

stat.ME

A note on conditional versus joint unconditional weak convergence in bootstrap consistency results

The consistency of a bootstrap or resampling scheme is classically validated by weak convergence of conditional laws. However, when working with stochastic processes in the space of bounded functions and their weak convergence in the Hoffmann-Jørgensen sense, an obstacle occurs: due to possible non-measurability, neither laws nor conditional laws are well-defined. Starting from an equivalent formulation of weak convergence based on the bounded Lipschitz metric, a classical circumvent is to formulate bootstrap consistency in terms of the latter distance between what might be called a \emph{conditional law} of the (non-measurable) bootstrap process and the law of the limiting process. The main contribution of this note is to provide an equivalent formulation of bootstrap consistency in the space of bounded functions which is more intuitive and easy to work with. Essentially, the equivalent formulation consists of (unconditional) weak convergence of the original process jointly with two bootstrap replicates. As a by-product, we provide two equivalent formulations of bootstrap consistency for statistics taking values in separable metric spaces: the first in terms of (unconditional) weak convergence of the statistic jointly with its bootstrap replicates, the second in terms of convergence in probability of the empirical distribution function of the bootstrap replicates. Finally, the asymptotic validity of bootstrap-based confidence intervals and tests is briefly revisited, with particular emphasis on the, in practice unavoidable, Monte Carlo approximation of conditional quantiles.

math.ST

Some copula inference procedures adapted to the presence of ties

When modeling the distribution of a multivariate continuous random vector using the so-called \emph{copula approach}, it is not uncommon to have ties in the coordinate samples of the available data because of rounding or lack of measurement precision. Yet, the vast majority of existing inference procedures on the underlying copula were both theoretically derived and practically implemented under the assumption of no ties. Applying them nonetheless can lead to strongly biased results. Some of the existing statistical tests can however be adapted to provide meaningful results in the presence of ties. It is the case of some tests of exchangeability, radial symmetry, extreme-value dependence and goodness of fit. Detailed algorithms for computing approximate p-values for the modified tests are provided and their finite-sample behaviors are empirically investigated through extensive Monte Carlo experiments. An illustration on a real-world insurance data set concludes the work.

stat.ME

Detecting distributional changes in samples of independent block maxima using probability weighted moments

The analysis of seasonal or annual block maxima is of interest in fields such as hydrology, climatology or meteorology. In connection with the celebrated method of block maxima, we study several tests that can be used to assess whether the available series of maxima is identically distributed. It is assumed that block maxima are independent but not necessarily generalized extreme value distributed. The asymptotic null distributions of the test statistics are investigated and the practical computation of approximate p-values is addressed. Extensive Monte-Carlo simulations show the adequate finite-sample behavior of the studied tests for a large number of realistic data generating scenarios. Illustrations on several environmental datasets conclude the work.

stat.ME

A dependent multiplier bootstrap for the sequential empirical copula process under strong mixing

Two key ingredients to carry out inference on the copula of multivariate observations are the empirical copula process and an appropriate resampling scheme for the latter. Among the existing techniques used for i.i.d. observations, the multiplier bootstrap of Rémillard and Scaillet (J. Multivariate Anal. 100 (2009) 377-386) frequently appears to lead to inference procedures with the best finite-sample properties. Bücher and Ruppert (J. Multivariate Anal. 116 (2013) 208-229) recently proposed an extension of this technique to strictly stationary strongly mixing observations by adapting the dependent multiplier bootstrap of Bühlmann (The blockwise bootstrap in time series and empirical processes (1993) ETH Zürich, Section 3.3) to the empirical copula process. The main contribution of this work is a generalization of the multiplier resampling scheme proposed by Bücher and Ruppert along two directions. First, the resampling scheme is now genuinely sequential, thereby allowing to transpose to the strongly mixing setting many of the existing multiplier tests on the unknown copula, including nonparametric tests for change-point detection. Second, the resampling scheme is now fully automatic as a data-adaptive procedure is proposed which can be used to estimate the bandwidth parameter. A simulation study is used to investigate the finite-sample performance of the resampling scheme and provides suggestions on how to choose several additional parameters. As by-products of this work, the validity of a sequential version of the dependent multiplier bootstrap for empirical processes of Bühlmann is obtained under weaker conditions on the strong mixing coefficients and the multipliers, and the weak convergence of the sequential empirical copula process is established under many serial dependence conditions.

math.ST

Dependent multiplier bootstraps for non-degenerate $U$-statistics under mixing conditions with applications

The asymptotic validity of a resampling method for two sequential processes constructed from non-degenerate $U$-statistics is established under mixing conditions. The resampling schemes, referred to as {\em dependent multiplier bootstraps}, result from an adaptation of the seminal approach of \cite{GomHor02} to mixing sequences. The proofs exploit recent results of \cite{DehWen10b} on degenerate $U$-statistics. A data-driven procedure for estimating a key bandwidth parameter involved in the resampling schemes is also suggested, making the use of the studied dependent multiplier bootstraps fully automatic. The derived results are applied to the construction of confidence intervals and to test for change-point detection. For such applications, Monte Carlo experiments suggest that the use of the proposed resampling approaches can have advantages over that of estimated asymptotic distributions.

math.ST

Detecting breaks in the dependence of multivariate extreme-value distributions

In environmental sciences, it is often of interest to assess whether the dependence between extreme measurements has changed during the observation period. The aim of this work is to propose a statistical test that is particularly sensitive to such changes. The resulting procedure is also extended to allow the detection of changes in the extreme-value dependence under the presence of known breaks in the marginal distributions. Simulations are carried out to study the finite-sample behavior of both versions of the proposed test. Illustrations on hydrological data sets conclude the work.

stat.ME