SearcharxivSearch

arXiv subjects

Gabriela Ciuperca

Publications and source records attributed to Gabriela Ciuperca.

At least 19 recordsLinked to original sources

Transfert learning and adaptive LASSO quantile

We propose for a quantile regression an estimation method for transferring knowledge using two $L_1$ penalties based on an estimator obtained from a source database. The proposed transfer learning estimator satisfies the properties of consistency and sparsity. Its convergence rate and asymptotic behavior are studied in several scenarios. This knowledge transfer results in a shorter computation time than that of the standard adaptive LASSO estimator. Another advantage of our method is that it can be applied to models with non-Gaussian errors. In addition, in order to implement the computing of the adaptive transfer LASSO quantile estimator, we propose an algorithm. The simulations confirm the theoretical results and demonstrate that the adaptive learning estimator, calculated using the proposed algorithm, is more competitive than the LASSO estimators. Finally, we illustrate the practical utility of the proposed transfer learning estimator and algorithm using a real-data application involving the physicochemical properties of protein tertiary structures.

stat.ME

Model selection by cross-validation in an expectile linear regression

For linear models that may have asymmetric errors, we study variable selection by cross-validation. The data are split into training and validation sets, with the number of observations in the validation set much larger than in the training set. For the model coefficients, the expectile or adaptive LASSO expectile estimators are calculated on the training set. These estimators will be used to calculate the cross-validation mean score (CVS) on the validation set. We show that the model that minimizes CVS is consistent in two cases: when the number of explanatory variables is fixed or when it depends on the number of observations. Monte Carlo simulations confirm the theoretical results and demonstrate the superiority of our estimation method compared to two others in the literature. The usefulness of the CV expectile model selection technique is illustrated by applying it to real data sets.

stat.ME

Right-censored models on massive data

This article considers the automatic selection problem of the relevant explanatory variables in a right-censored model on a massive database. We propose and study four aggregated censored adaptive LASSO estimators constructed by dividing the observations in such a way as to keep the consistency of the estimator of the survival curve. We show that these estimators have the same theoretical oracle properties as the one built on the full database. Moreover, by Monte Carlo simulations we obtain that their calculation time is smaller than that of the full database. The simulations confirm also the theoretical properties. For optimal tuning parameter selection, we propose a BIC-type criterion.

math.ST

Right-censored models by the expectile method

Based on the expectile loss function and the adaptive LASSO penalty, the paper proposes and studies the estimation methods for the accelerated failure time (AFT) model. In this approach, we need to estimate the survival function of the censoring variable by the Kaplan-Meier estimator. The AFT model parameters are first estimated by the expectile method and afterwards, when the number of explanatory variables can be large, by the adaptive LASSO expectile method which directly carries out the automatic selection of variables. We also obtain the convergence rate and asymptotic normality for the two estimators, while showing the sparsity property for the censored adaptive LASSO expectile estimator. A numerical study using Monte Carlo simulations confirms the theoretical results and demonstrates the competitive performance of the two proposed estimators. The usefulness of these estimators is illustrated by applying them to three survival data sets.

math.ST

Smoothed empirical likelihood estimation and automatic variable selection for an expectile high-dimensional model with possibly missing response variable

We consider a linear model which can have a large number of explanatory variables, the errors with an asymmetric distribution or some values of the explained variable are missing at random. In order to take in account these several situations, we consider the non parametric empirical likelihood (EL) estimation method. Because a constraint in EL contains an indicator function then a smoothed function instead of the indicator will be considered. Two smoothed expectile maximum EL methods are proposed, one of which will automatically select the explanatory variables. For each of the methods we obtain the convergence rate of the estimators and their asymptotic normality. The smoothed expectile empirical log-likelihood ratio process follow asymptotically a chi-square distribution and moreover the adaptive LASSO smoothed expectile maximum EL estimator satisfies the sparsity property which guarantees the automatic selection of zero model coefficients. In order to implement these methods, we propose four algorithms.

stat.ME

Automatic selection by penalized asymmetric Lq-norm in an high-dimensional model with grouped variables

The paper focuses on the automatic selection of the grouped explanatory variables in an high-dimensional model, when the model errors are asymmetric. After introducing the model and notations, we define the adaptive group LASSO expectile estimator for which we prove the oracle properties: the sparsity and the asymptotic normality. Afterwards, the results are generalized by considering the asymmetric $L_q$-norm loss function. The theoretical results are obtained in several cases with respect to the number of variable groups. This number can be fixed or dependent on the sample size $n$, with the possibility that it is of the same order as $n$. Note that these new estimators allow us to consider weaker assumptions on the data and on the model errors than the usual ones. Simulation study demonstrates the competitive performance of the proposed penalized expectile regression, especially when the samples size is close to the number of explanatory variables and model errors are asymmetrical. An application on air pollution data is considered.

math.ST

Real-time detection of a change-point in a linear expectile model

In the present paper we address the real-time detection problem of a change-point in the coefficients of a linear model with the possibility that the model errors are asymmetrical and that the explanatory variables number is large. We build test statistics based on the cumulative sum (CUSUM) of the expectile function derivatives calculated on the residuals obtained by the expectile and adaptive LASSO expectile estimation methods. The asymptotic distribution of these statistics are obtained under the hypothesis that the model does not change. Moreover, we prove that they diverge when the model changes at an unknown observation. The asymptotic study of the test statistics under these two hypotheses allows us to find the asymptotic critical region and the stopping time, that is the observation where the model will change. The empirical performance is investigated by a comparative simulation study with other statistics of CUSUM type. Two examples on real data are also presented to demonstrate its interest in practice.

stat.ME

Detection of similar successive groups in a model with diverging number of variable groups

In this paper, a linear model with grouped explanatory variables is considered. The idea is to perform an automatic detection of different successive groups of the unknown coefficients under the assumption that the number of groups is of the same order as the sample size. The standard least squares loss function and the quantile loss function are both used together with the fused and adaptive fused penalty to simultaneously estimate and group the unknown parameters. The proper convergence rate is given for the obtained estimators and the upper bound for the number of different successive group is derived. A simulation study is used to compare the empirical performance of the proposed fused and adaptive fused estimators and a real application on the air quality data demonstrates the practical applicability of the proposed methods.

stat.ME

Change-point detection in a linear model by adaptive fused quantile method

A novel approach to quantile estimation in multivariate linear regression models with change-points is proposed: the change-point detection and the model estimation are both performed automatically, by adopting either the quantile fused penalty or the adaptive version of the quantile fused penalty. These two methods combine the idea of the check function used for the quantile estimation and the $L_1$ penalization principle known from the signal processing and, unlike some standard approaches, the presented methods go beyond typical assumptions usually required for the model errors, such as sub-Gaussian or Normal distribution. They can effectively handle heavy-tailed random error distributions, and, in general, they offer a more complex view on the data as one can obtain any conditional quantile of the target distribution, not just the conditional mean. The consistency of detection is proved and proper convergence rates for the parameter estimates are derived. The empirical performance is investigated via an extensive comparative simulation study and practical utilization is demonstrated using a real data example.

math.ST

Adaptive elastic-net selection in a quantile model with diverging number of variable groups

In real applications of the linear model, the explanatory variables are very often naturally grouped, the most common example being the multivariate variance analysis. In the present paper, a quantile model with structure group is considered, the number of groups can diverge with sample size. We introduce and study the adaptive elastic-net group estimator, for improving the parameter estimation accuracy. This method allows automatic selection, with a probability converging to one, of significant groups and further the non zero parameter estimators are asymptotically normal. The convergence rate of the adaptive elastic-net group quantile estimator is also obtained, rate which depends on the number of groups. In order to put the estimation method into practice, an algorithm based on the subgradient method is proposed and implemented. The Monte Carlo simulations show that the adaptive elastic-net group quantile estimations are more accurate that other existing group estimations in the literature. Moreover, the numerical study confirms the theoretical results and the usefulness of the proposed estimation method.

math.ST

Change-point Detection by the Quantile LASSO Method

A simultaneous change-point detection and estimation in a piece-wise constant model is a common task in modern statistics. If, in addition, the whole estimation can be performed automatically, in just one single step without going through any hypothesis tests for non-identifiable models, or unwieldy classical a-posterior methods, it becomes an interesting, but also challenging idea. In this paper we introduce the estimation method based on the quantile LASSO approach. Unlike standard LASSO approaches, our method does not rely on typical assumptions usually required for the model errors, such as sub-Gaussian or Normal distribution. The proposed quantile LASSO method can effectively handle heavy-tailed random error distributions, and, in general, it offers a more complex view of the data as one can obtain any conditional quantile of the target distribution, not just the conditional mean. It is proved that under some reasonable assumptions the number of change-points is not underestimated with probability tenting to one, and, in addition, when the number of change-points is estimated correctly, the change-point estimates provided by the quantile LASSO are consistent. Numerical simulations are used to demonstrate these results and to illustrate the empirical performance robust favor of the proposed quantile LASSO method.

math.ST

Variable selection in high-dimensional linear model with possibly asymmetric or heavy-tailed errors

We consider the problem of automatic variable selection in a linear model with asymmetric or heavy-tailed errors when the number of explanatory variables diverges with the sample size. For this high-dimensional model, the penalized least square method is not appropriate and the quantile framework makes the inference more difficult because to the non differentiability of the loss function. We propose and study an estimation method by penalizing the expectile process with an adaptive LASSO penalty. Two cases are considered: the number of model parameters is smaller and afterwards larger than the sample size, the two cases being distinct by the adaptive penalties considered. For each case we give the rate convergence and establish the oracle properties of the adaptive LASSO expectile estimator. The proposed estimators are evaluated through Monte Carlo simulations and compared with the adaptive LASSO quantile estimator. We applied also our estimation method to real data in genetics when the number of parameters is greater than the sample size.

math.ST

Adaptive Fused LASSO in Grouped Quantile Regression

This paper considers quantile model with grouped explanatory variables. In order to have the sparsity of the parameter groups but also the sparsity between two successive groups of variables, we propose and study an adaptive fused group LASSO quantile estimator. The number of variable groups can be fixed or divergent. We find the convergence rate under classical assumptions and we show that the proposed estimator satisfies the oracle properties.

math.ST

Adaptive group LASSO selection in quantile models

The paper considers a linear model with grouped explanatory variables. If the model errors are not with zero mean and bounded variance or if model contains outliers, then the least squares framework is not appropriate. Thus, the quantile regression is an interesting alternative. In order to automatically select the relevant variable groups, we propose and study here the adaptive group LASSO quantile estimator. We establish the sparsity and asymptotic normality of the proposed estimator in two cases: fixed number and divergent number of variable groups. Numerical study by Monte Carlo simulations confirms the theoretical results and illustrates the performance of the proposed estimator.

math.ST

Real time change-point detection in a nonlinear quantile model

Most studies in real time change-point detection either focus on the linear model or use the CUSUM method under classical assumptions on model errors. This paper considers the sequential change-point detection in a nonlinear quantile model. A test statistic based on the CUSUM of the quantile process subgradient is proposed and studied. Under null hypothesis that the model does not change, the asymptotic distribution of the test statistic is determined. Under alternative hypothesis that at some unknown observation there is a change in model, the proposed test statistic converges in probability to $\infty$. These results allow to build the critical regions on open-end and on closed-end procedures. Simulation results, using Monte Carlo technique, investigate the performance of the test statistic, specially for heavy-tailed error distributions. We also compare it with the classical CUSUM test statistic.

math.ST

Empirical likelihood test for high-dimensional two-sample model

A non parametric method based on the empirical likelihood is proposed for detecting the change in the coefficients of high-dimensional linear model where the number of model variables may increase as the sample size increases. This amounts to testing the null hypothesis of no change against the alternative of one change in the regression coefficients. Based on the theoretical asymptotic behaviour of the empirical likelihood ratio statistic, we propose, for a fixed design, a simpler test statistic, easier to use in practice. The asymptotic normality of the proposed test statistic under the null hypothesis is proved, a result which is different from the $χ^2$ law for a model with a fixed variable number. Under alternative hypothesis, the test statistic diverges. We can then find the asymptotic confidence region for the difference of parameters of the two phases. Some Monte-Carlo simulations study the behaviour of the proposed test statistic.

math.ST

Model selection in high-dimensional quantile regression with seamless $L_0$ penalty

In this paper we are interested in parameters estimation of linear model when number of parameters increases with sample size. Without any assumption about moments of the model error, we propose and study the seamless $L_0$ quantile estimator. For this estimator we first give the convergence rate. Afterwards, we prove that it correctly distinguishes between zero and nonzero parameters and that the estimators of the nonzero parameters are asymptotically normal. A consistent BIC criterion to select the tuning parameters is given.

math.ST

Estimation in a change-point nonlinear quantile model

This paper considers a nonlinear quantile model with change-points. The quantile estimation method, which as a particular case includes median model, is more robust with respect to other traditional methods when model errors contain outliers. Under relatively weak assumptions, the convergence rate and asymptotic distribution of change-point and of regression parameter estimators are obtained. Numerical study by Monte Carlo simulations shows the performance of the proposed method for nonlinear model with change-points.

math.ST