Searcharxiv⌕ Search

arXiv subjects

Salima El Kolei

Publications and source records attributed to Salima El Kolei.

9 recordsLinked to original sources

Goodness-of-fit testing of the distribution of posterior classification probabilities for validating model-based clustering

We present the first method for assessing the relevance of a model-based clustering result in a general framework. Standard validation criteria, like the adjusted Rand index, rely on external labels to assess partition accuracy; consequently, they are inapplicable to real-world clustering problems where labels are missing. In contrast, our method offers an internal goodness-of-fit diagnostic, since it evaluates the validity of the clustering mechanism by testing the specification of the posterior probabilities of classification defined on the unit simplex. Because this simplex dimension is fixed by the number of clusters, the procedure naturally circumvents the curse of dimensionality, making it applicable to high-dimensional data where traditional density-based tests fail. The testing procedure requires only a consistent estimator of the parameters and the associated posterior classification probabilities for each observation, and its implementation is straightforward, as no additional model fitting is needed. Under the null hypothesis, the method exploits the fact that any functional transformation of the posterior probabilities has the same expectation under both the model being tested and the true data-generating process. The resulting goodness-of-fit test is constructed via an empirical likelihood approach with a growing number of moment conditions, allowing asymptotic detection of any alternative. A block-splitting strategy, employed to account for parameter estimation, provides a vector of test statistics that behave like a vector of independent chi-square random variables. Therefore, the goodness-of-fit of the posterior classification probabilities is assessed via the goodness-of-fit of the vector of empirical likelihood ratio test statistics. Hence, based on the distribution of this vector of statistics, different goodness-of-fit tests (e.g., Kolmogorov-Smirnov) can be used to investigate the distribution of the vector of test statistics with an exact asymptotic significance level.

math.ST↗

Estimation of the Order of Non-Parametric Hidden Markov Models using the Singular Values of an Integral Operator

We are interested in assessing the order of a finite-state Hidden Markov Model (HMM) with the only two assumptions that the transition matrix of the latent Markov chain has full rank and that the density functions of the emission distributions are linearly independent. We introduce a new procedure for estimating this order by investigating the rank of some well-chosen integral operator which relies on the distribution of a pair of consecutive observations. This method circumvents the usual limits of the spectral method when it is used for estimating the order of an HMM: it avoids the choice of the basis functions; it does not require any knowledge of an upper-bound on the order of the HMM (for the spectral method, such an upper-bound is defined by the number of basis functions); it permits to easily handle different types of data (including continuous data, circular data or multivariate continuous data) with a suitable choice of kernel. The method relies on the fact that the order of the HMM can be identified from the distribution of a pair of consecutive observations and that this order is equal to the rank of some integral operator (\emph{i.e.} the number of its singular values that are non-zero). Since only the empirical counter-part of the singular values of the operator can be obtained, we propose a data-driven thresholding procedure. An upper-bound on the probability of overestimating the order of the HMM is established. Moreover, sufficient conditions on the bandwidth used for kernel density estimation and on the threshold are stated to obtain the consistency of the estimator of the order of the HMM. The procedure is easily implemented since the values of all the tuning parameters are determined by the sample size.

math.ST↗

Investigating swimming technical skills by a double partition clustering of multivariate functional data allowing for dimension selection

Investigating technical skills of swimmers is a challenge for performance improvement, that can be achieved by analyzing multivariate functional data recorded by Inertial Measurement Units (IMU). To investigate technical levels of front-crawl swimmers, a new model-based approach is introduced to obtain two complementary partitions reflecting, for each swimmer, its swimming pattern and its ability to reproduce it. Contrary to the usual approaches for functional data clustering, the proposed approach also considers the information of the residuals resulting from the functional basis decomposition. Indeed, after decomposing into functional basis both the original signal (measuring the swimming pattern) and the signal of squared residuals (measuring the ability to reproduce the swimming pattern), the method fits the joint distribution of the coefficients related to both decompositions by considering dependency between both partitions. Modeling this dependency is mandatory since the difficulty of reproducing a swimming pattern depends on its shape. Moreover, a sparse decomposition of the distribution within components that permits a selection of the relevant dimensions during clustering is proposed. The partitions obtained on the IMU data aggregate the kinematical stroke variability linked to swimming technical skills and allow relevant biomechanical strategy for front-crawl sprint performance to be identified.

stat.AP↗

Nonparametric estimation in a regression model with additive and multiplicative noise

In this paper, we consider an unknown functional estimation problem in a general nonparametric regression model with the feature of having both multiplicative and additive noise.We propose two new wavelet estimators in this general context. We prove that they achieve fast convergence rates under the mean integrated square error over Besov spaces. The obtained rates have the particularity of being established under weak conditions on the model. A numerical study in a context comparable to stochastic frontier estimation (with the difference that the boundary is not necessarily a production function) supports the theory.

math.ST↗

Adaptive Density Estimation on Bounded Domains

We study the estimation, in Lp-norm, of density functions defined on [0,1]^d. We construct a new family of kernel density estimators that do not suffer from the so-called boundary bias problem and we propose a data-driven procedure based on the Goldenshluger and Lepski approach that jointly selects a kernel and a bandwidth. We derive two estimators that satisfy oracle-type inequalities. They are also proved to be adaptive over a scale of anisotropic or isotropic Sobolev-Slobodetskii classes (which are particular cases of Besov or Sobolev classical classes). The main interest of the isotropic procedure is to obtain adaptive results without any restriction on the smoothness parameter.

math.ST↗

Analysis, detection and correction of misspecified discrete time state space models

Misspecifications (i.e. errors on the parameters) of state space models lead to incorrect inference of the hidden states. This paper studies weakly nonlin-ear state space models with additive Gaussian noises and proposes a method for detecting and correcting misspecifications. The latter induce a biased estimator of the hidden state but also happen to induce correlation on innovations and other residues. This property is used to find a well-defined objective function for which an optimisation routine is applied to recover the true parameters of the model. It is argued that this method can consistently estimate the bias on the parameter. We demonstrate the algorithm on various models of increasing complexity.

stat.AP↗

Parametric inference of hidden discrete-time diffusion processes by deconvolution

We study a new parametric approach for hidden discrete-time diffusion models. This method is based on contrast minimization and deconvolution and leads to estimate a large class of stochastic models with nonlinear drift and nonlinear diffusion. It can be applied, for example, for ecological and financial state space models. After proving consistency and asymptotic normality of the estimation, leading to asymptotic confidence intervals, we provide a thorough numerical study, which compares many classical methods used in practice (Non Linear Least Square estimator, Monte Carlo Expectation Maxi-mization Likelihood estimator and Bayesian estimators) to estimate stochastic volatility model. We prove that our estimator clearly outperforms the Maximum Likelihood Estimator in term of computing time, but also most of the other methods. We also show that this contrast method is the most stable and also does not need any tuning parameter.

math.ST↗

Propagation of initial errors on the parameters for linear and Gaussian state space models

For linear and Gaussian state space models parametrized by $θ_0 \in Θ\subset \mathbb{R}^r, r \geq 1$ corresponding to the vector of parameters of the model, the Kalman filter gives exactly the solution for the optimal filtering under weak assumptions. This result supposes that $θ_0$ is perfectly known. In most real applications, this assumption is not realistic since $θ_0$ is unknown and has to be estimated. In this paper, we analysis the Kalman filter for a biased estimator of $θ_0$. We show the propagation of this bias on the estimation of the hidden state. We give an expression of this propagation for linear and Gaussian state space models and we extend this result for almost linear models estimated by the Extended Kalman filter. An illustration is given for the autoregressive process with measurement noises widely studied in econometrics to model economic and financial data.

stat.OT↗

Parametric estimation of hidden stochastic model by contrast minimization and deconvolution: application to the Stochastic Volatility Model

We study a new parametric approach for particular hidden stochastic models such as the Stochastic Volatility model. This method is based on contrast minimization and deconvolution. After proving consistency and asymptotic normality of the estimation leading to asymptotic confidence intervals, we provide a thorough numerical study, which compares most of the classical methods that are used in practice (Quasi Maximum Likelihood estimator, Simulated Expectation Maximization Likelihood estimator and Bayesian estimators). We prove that our estimator clearly outperforms the Maximum Likelihood Estimator in term of computing time, but also most of the other methods. We also show that this contrast method is the most robust with respect to non Gaussianity of the error and also does not need any tuning parameter.

stat.AP↗