SearcharxivSearch

arXiv subjects

William Kengne

Publications and source records attributed to William Kengne.

At least 19 recordsLinked to original sources

Adaptive deep nonparametric regression from dependent data under covariate shift

Covariate shift often occurs because, in many real applications, the source and the target observations may be generated from different distributions. In this case, the standard metric under the source distribution is not appropriate. This paper considers deep neural network estimators for nonparametric quantile and Huber regression under covariate shift and from dependent observations. We deal with a generalized Bernstein-type inequality that is satisfied by many classical models, including i.i.d. observations, $\phi$-mixing, strong mixing, and $\mathcal{C}$-mixing processes. To perform the covariate shift phenomenon, we propose a sparse-penalized deep neural network (SPDNN) estimator that takes into account the discrepancy between the source and target distributions of the data. When the density ratio (between the source and target distributions of the covariate) is unknown, a two steps pre-training procedure is carried out: the first step is devoted to the construction of a least squares SPDNN estimator of the density ratio; which is used in the second step to perform a pre-training reweighted SPDNN estimator of the regression function. For both the quantile and the Huber regression, non-asymptotic error bounds of the proposed SPDNN estimators are established in the class of H\"older smooth functions. These estimators can adaptively attain (up to a logarithmic factor) the minimax optimal convergence rate from i.i.d. data as well as from several classical time series models.

stat.ML

Deep regression learning from dependent observations with minimum error entropy principle

This paper considers nonparametric regression from strongly mixing observations. The proposed approach is based on deep neural networks with minimum error entropy (MEE) principle. We study two estimators: the non-penalized deep neural network (NPDNN) and the sparse-penalized deep neural network (SPDNN) predictors. Upper bounds of the expected excess risk are established for both estimators over the classes of H\"older and composition H\"older functions. For the models with Gaussian error, the rates of the upper bound obtained match (up to a logarithmic factor) with the lower bounds established in \cite{schmidt2020nonparametric}, showing that both the MEE-based NPDNN and SPDNN estimators from strongly mixing data can achieve the minimax optimal convergence rate.

stat.ML

A general framework for deep learning

This paper develops a general approach for deep learning for a setting that includes nonparametric regression and classification. We perform a framework from data that fulfills a generalized Bernstein-type inequality, including independent, $\phi$-mixing, strongly mixing and $\mathcal{C}$-mixing observations. Two estimators are proposed: a non-penalized deep neural network estimator (NPDNN) and a sparse-penalized deep neural network estimator (SPDNN). For each of these estimators, bounds of the expected excess risk on the class of H\"older smooth functions and composition H\"older functions are established. Applications to independent data, as well as to $\phi$-mixing, strongly mixing, $\mathcal{C}$-mixing processes are considered. For each of these examples, the upper bounds of the expected excess risk of the proposed NPDNN and SPDNN predictors are derived. It is shown that both the NPDNN and SPDNN estimators are minimax optimal (up to a logarithmic factor) in many classical settings.

math.ST

Deep learning from strongly mixing observations: Sparse-penalized regularization and minimax optimality

The explicit regularization and optimality of deep neural networks estimators from independent data have made considerable progress recently. The study of such properties on dependent data is still a challenge. In this paper, we carry out deep learning from strongly mixing observations, and deal with the squared and a broad class of loss functions. We consider sparse-penalized regularization for deep neural network predictor. For a general framework that includes, regression estimation, classification, time series prediction,$\cdots$, oracle inequality for the expected excess risk is established and a bound on the class of Hölder smooth functions is provided. For nonparametric regression from strong mixing data and sub-exponentially error, we provide an oracle inequality for the $L_2$ error and investigate an upper bound of this error on a class of Hölder composition functions. For the specific case of nonparametric autoregression with Gaussian and Laplace errors, a lower bound of the $L_2$ error on this Hölder composition class is established. Up to logarithmic factor, this bound matches its upper bound; so, the deep neural network estimator attains the minimax optimal rate.

stat.ML

Minimax optimality of deep neural networks on dependent data via PAC-Bayes bounds

In a groundbreaking work, Schmidt-Hieber (2020) proved the minimax optimality of deep neural networks with ReLu activation for least-square regression estimation over a large class of functions defined by composition. In this paper, we extend these results in many directions. First, we remove the i.i.d. assumption on the observations, to allow some time dependence. The observations are assumed to be a Markov chain with a non-null pseudo-spectral gap. Then, we study a more general class of machine learning problems, which includes least-square and logistic regression as special cases. Leveraging on PAC-Bayes oracle inequalities and a version of Bernstein inequality due to Paulin (2015), we derive upper bounds on the estimation risk for a generalized Bayesian estimator. In the case of least-square regression, this bound matches (up to a logarithmic factor) the lower bound of Schmidt-Hieber (2020). We establish a similar lower bound for classification with the logistic loss, and prove that the proposed DNN estimator is optimal in the minimax sense.

stat.ML

Robust deep learning from weakly dependent data

Recent developments on deep learning established some theoretical properties of deep neural networks estimators. However, most of the existing works on this topic are restricted to bounded loss functions or (sub)-Gaussian or bounded input. This paper considers robust deep learning from weakly dependent observations, with unbounded loss function and unbounded input/output. It is only assumed that the output variable has a finite $r$ order moment, with $r >1$. Non asymptotic bounds for the expected excess risk of the deep neural network estimator are established under strong mixing, and $ψ$-weak dependence assumptions on the observations. We derive a relationship between these bounds and $r$, and when the data have moments of any order (that is $r=\infty$), the convergence rate is close to some well-known results. When the target predictor belongs to the class of Hölder smooth functions with sufficiently large smoothness index, the rate of the expected excess risk for exponentially strongly mixing data is close to or as same as those for obtained with i.i.d. samples. Application to robust nonparametric regression and robust nonparametric autoregression are considered. The simulation study for models with heavy-tailed errors shows that, robust estimators with absolute loss and Huber loss function outperform the least squares method.

stat.ML

Penalized deep neural networks estimator with general loss functions under weak dependence

This paper carries out sparse-penalized deep neural networks predictors for learning weakly dependent processes, with a broad class of loss functions. We deal with a general framework that includes, regression estimation, classification, times series prediction, $\cdots$ The $ψ$-weak dependence structure is considered, and for the specific case of bounded observations, $θ_\infty$-coefficients are also used. In this case of $θ_\infty$-weakly dependent, a non asymptotic generalization bound within the class of deep neural networks predictors is provided. For learning both $ψ$ and $θ_\infty$-weakly dependent processes, oracle inequalities for the excess risk of the sparse-penalized deep neural networks estimators are established. When the target function is sufficiently smooth, the convergence rate of these excess risk is close to $\mathcal{O}(n^{-1/3})$. Some simulation results are provided, and application to the forecast of the particulate matter in the Vitória metropolitan area is also considered.

stat.ML

Sparse-penalized deep neural networks estimator under weak dependence

We consider the nonparametric regression and the classification problems for $ψ$-weakly dependent processes. This weak dependence structure is more general than conditions such as, mixing, association, $\ldots$. A penalized estimation method for sparse deep neural networks is performed. In both nonparametric regression and binary classification problems, we establish oracle inequalities for the excess risk of the sparse-penalized deep neural networks estimators. Convergence rates of the excess risk of these estimators are also derived. The simulation results displayed show that, the proposed estimators overall work well than the non penalized estimators.

stat.ML

Excess risk bound for deep learning under weak dependence

This paper considers deep neural networks for learning weakly dependent processes in a general framework that includes, for instance, regression estimation, time series prediction, time series classification. The $ψ$-weak dependence structure considered is quite large and covers other conditions such as mixing, association,$\ldots$ Firstly, the approximation of smooth functions by deep neural networks with a broad class of activation functions is considered. We derive the required depth, width and sparsity of a deep neural network to approximate any Hölder smooth function, defined on any compact set $\mx$. Secondly, we establish a bound of the excess risk for the learning of weakly dependent observations by deep neural networks. When the target function is sufficiently smooth, this bound is close to the usual $\mathcal{O}(n^{-1/2})$.

stat.ML

Deep learning for $ψ$-weakly dependent processes

In this paper, we perform deep neural networks for learning $ψ$-weakly dependent processes. Such weak-dependence property includes a class of weak dependence conditions such as mixing, association,$\cdots$ and the setting considered here covers many commonly used situations such as: regression estimation, time series prediction, time series classification,$\cdots$ The consistency of the empirical risk minimization algorithm in the class of deep neural networks predictors is established. We achieve the generalization bound and obtain a learning rate, which is less than $\mathcal{O}(n^{-1/α})$, for all $α> 2 $. Applications to binary time series classification and prediction in affine causal models with exogenous covariates are carried out. Some simulation results are provided, as well as an application to the US recession data.

stat.ML

Density power divergence for general integer-valued time series with multivariate exogenous covariate

In this article, we study a robust estimation method for a general class of integer-valued time series models. The conditional distribution of the process belongs to a broad class of distribution and unlike classical autoregressive framework, the conditional mean of the process also depends on some multivariate exogenous covariate. We derive a robust inference procedure based on the minimum density power divergence. Under certain regularity conditions, we establish that the proposed estimator is consistent and asymptotically normal. Simulation experiments are conducted to illustrate the empirical performances of the estimator. An application to the number of transactions per minute for the stock Ericsson B is also provided.

math.ST

Statistical learning for $ψ$-weakly dependent processes

We consider statistical learning question for $ψ$-weakly dependent processes, that unifies a large class of weak dependence conditions such as mixing, association,$\cdots$ The consistency of the empirical risk minimization algorithm is established. We derive the generalization bounds and provide the learning rate, which, on some H{ö}lder class of hypothesis, is close to the usual $O(n^{-1/2})$ obtained in the {\it i.i.d.} case. Application to time series prediction is carried out with an example of causal models with exogenous covariates.

math.ST

Some asymptotic results for time series model selection

We consider the model selection problem for a large class of time series models, including, multivariate count processes, causal processes with exogenous covariates. A procedure based on a general penalized contrast is proposed. Some asymptotic results for weak and strong consistency are established. The non consistency issue is addressed, and a class of penalty term, that does not ensure consistency is provided. Examples of continuous valued and multivariate count autoregressive time series are considered.

math.ST

Efficient and Consistent Data-Driven Model Selection for Time Series

This paper studies the model selection problem in a large class of causal time series models, which includes both the ARMA or AR($\infty$) processes, as well as the GARCH or ARCH($\infty$), APARCH, ARMA-GARCH and many others processes. We first study the asymptotic behavior of the ideal penalty that minimizes the risk induced by a quasi-likelihood estimation among a finite family of models containing the true model. Then, we provide general conditions on the penalty term for obtaining the consistency and efficiency properties. We notably prove that consistent model selection criteria outperform classical AIC criterion in terms of efficiency. Finally, we derive from a Bayesian approach the usual BIC criterion, and by keeping all the second order terms of the Laplace approximation, a data-driven criterion denoted KC'. Monte-Carlo experiments exhibit the obtained asymptotic results and show that KC' criterion does better than the AIC and BIC ones in terms of consistency and efficiency.

math.ST

Inference and model selection in general causal time series with exogenous covariates

In this paper, we study a general class of causal processes with exogenous covariates, including many classical processes such as the ARMA-GARCH, APARCH, ARMAX, GARCH-X and APARCH-X processes. Under some Lipschitz-type conditions, the existence of a $τ$-weakly dependent strictly stationary and ergodic solution is established. We provide conditions for the strong consistency and derive the asymptotic distribution of the quasi-maximum likelihood estimator (QMLE), both when the true parameter is an interior point of the parameter's space and when it belongs to the boundary. A significance Wald-type test of parameter is developed. This test is quite extensive and includes the test of nullity of the parameter's components, which in particular, allows us to assess the relevance of the exogenous covariates. Relying on the QMLE of the model, we also propose a penalized criterion to address the problem of the model selection for this class. The weak and the strong consistency of the procedure are established. Finally, Monte Carlo simulations are conducted to numerically illustrate the main results.

math.ST

Epidemic change-point detection in general causal time series

We consider an epidemic change-point detection in a large class of causal time series models, including among other processes, AR($\infty$), ARCH($\infty$), TARCH($\infty$), ARMA-GARCH. A test statistic based on the Gaussian quasi-maximum likelihood estimator of the parameter is proposed. It is shown that, under the null hypothesis of no change, the test statistic converges to a distribution obtained from a difference of two Brownian bridge and diverges to infinity under the epidemic alternative. Numerical results for simulation and real data example are provided.

math.ST

A general procedure for change-point detection in multivariate time series

We consider the change-point detection in multivariate continuous and integer valued time series. We propose a Wald-type statistic based on the estimator performed by a general contrast function; which can be constructed from the likelihood, a quasi-likelihood, a least squares method, etc. Sufficient conditions are provided to ensure that the statistic convergences to a well-known distribution under the null hypothesis (of no change) and diverges to infinity under the alternative; which establishes the consistency of the procedure. Some examples are detailed to illustrate the scope of application of the proposed procedure. Simulation experiments are conducted to illustrate the asymptotic results.

math.ST

Epidemic change-point detection in general integer-valued time series

In this paper, we consider the structural change in a class of discrete valued time series, which the true conditional distribution of the observations is assumed to be unknown. The conditional mean of the process depends on a parameter $θ^*$ which may change over time. We provide sufficient conditions for the consistency and the asymptotic normality of the Poisson quasi-maximum likelihood estimator (QMLE) of the model. We consider an epidemic change-point detection and propose a test statistic based on the QMLE of the parameter. Under the null hypothesis of a constant parameter (no change), the test statistic converges to a distribution obtained from a difference of two Brownian bridge. The test statistic diverges to infinity under the epidemic alternative, which establishes that the proposed procedure is consistent in power. The effectiveness of the proposed procedure is illustrated by simulated and real data examples.

math.ST