SearcharxivSearch

arXiv subjects

Salah Khardani

Publications and source records attributed to Salah Khardani.

5 recordsLinked to original sources

On the Dirichlet-kernel Gasser--Müller estimator and its competitors for fixed design regression on the simplex

A Dirichlet-kernel Gasser-Müller (D-GM) estimator is introduced for fixed design regression on the simplex, extending the univariate analog due to Chen [Statist. Sinica, vol. 10(1) (2000), pp. 73-91]. Its pointwise bias and variance, asymptotic normality, and mean integrated squared error are investigated. Some simulation experiments are conducted to compare its small-sample performance with that of two recently proposed alternatives: the Dirichlet-kernel Nadaraya-Watson (D-NW) and local linear (D-LL) estimators. The simulation results reveal that the D-LL estimator is best among the D-LL, D-NW, and D-GM estimators and that the proposed D-GM estimator is worst. A real data analysis is also reported for the GEMAS dataset to analyze the relationship between soil composition and pH levels across various agricultural and grazing lands in Europe.

math.ST

Dirichlet kernel density estimation on the simplex with missing data

Nonparametric density estimation for compositional data supported on the simplex is examined under a missing at random mechanism. Rather than imputing missing values and estimating the density from a completed data set, we adopt a strategy based on inverse probability weighting. The proposed estimator uses an adaptive Dirichlet kernel, which ensures nonnegativity on the simplex and favorable behavior near the boundary. When the observation probabilities are unknown, they are estimated through a Nadaraya-Watson regression step. The large-sample properties of the estimator are derived, including pointwise bias and variance expansions, optimal smoothing rates, and asymptotic normality. A simulation study investigates its finite-sample performance under varying sample sizes and missing rates. Simulations show our method outperforms inverse-probability-weighted kernel density estimators based on additive and isometric log-ratio transformations of the data for certain target densities. The methodology is further illustrated through an application to leukocyte composition data from the National Health and Nutrition Examination Survey (NHANES), which allows for the identification of the modal immune profile in the sampled population.

stat.ME

A Bernstein polynomial approach for the estimation of cumulative distribution functions in the presence of missing data

We study nonparametric estimation of univariate cumulative distribution functions (CDFs) pertaining to data missing at random. The proposed estimators smooth the inverse probability weighted (IPW) empirical CDF with the Bernstein operator, yielding monotone, $[0,1]$-valued curves that automatically adapt to bounded supports. We analyze two versions: a pseudo estimator that uses known propensities and a feasible estimator that uses propensities estimated nonparametrically from discrete auxiliary variables, the latter scenario being much more common in practice. For both, we derive pointwise bias and variance expansions, establish the optimal polynomial degree $m$ with respect to the mean integrated squared error, and prove the asymptotic normality. A key finding is that the feasible estimator has a smaller variance than the pseudo estimator by an explicit nonnegative correction term. We also develop an efficient degree selection procedure via least-squares cross-validation. Monte Carlo experiments show that, for small to moderate sample sizes, the Bernstein-smoothed pseudo and feasible estimators outperform their unsmoothed counterparts and the integrated version of the IPW kernel density estimator proposed by Dubnicka (2009), under certain models. A real-data application to fasting plasma glucose from the 2017-2018 NHANES survey illustrates the method in a practical setting. All code needed to reproduce our analyses is readily accessible on GitHub.

math.ST

Dirichlet kernel density estimation for strongly mixing sequences on the simplex

This paper investigates the theoretical properties of Dirichlet kernel density estimators for compositional data supported on simplices, for the first time addressing scenarios involving time-dependent observations characterized by strong mixing conditions. We establish rigorous results for the asymptotic normality and mean squared error of these estimators, extending previous findings from the independent and identically distributed (iid) context to the more general setting of strongly mixing processes. To demonstrate its practical utility, the estimator is applied to monthly market-share compositions of several Renault vehicle classes over a twelve-year period, with bandwidth selection performed via leave-one-out least squares cross-validation. Our findings underscore the reliability and strength of Dirichlet kernel techniques when applied to temporally dependent compositional data.

math.ST

Kernel-smoothed conditional quantiles of randomly censored functional stationary ergodic data

This paper, investigates the conditional quantile estimation of a scalar random response and a functional random covariate (i.e. valued in some infinite-dimensional space) whenever {\it functional stationary ergodic data with random censorship} are considered. We introduce a kernel type estimator of the conditional quantile function. We establish the strong consistency with rate of this estimator as well as the asymptotic normality which induces a confidence interval that is usable in practice since it does not depend on any unknown quantity. An application to electricity peak demand interval prediction with censored smart meter data is carried out to show the performance of the proposed estimator.

math.ST