Searcharxiv⌕ Search

arXiv subjects

Valentin Patilea

Publications and source records attributed to Valentin Patilea.

35 records · Page 2Linked to original sources

A likelihood-based approach for cure regression models

We propose a new likelihood-based approach for estimation, inference and variable selection for parametric cure regression models in time-to-event analysis under random right-censoring. In this context, it often happens that some subjects are "cured", i.e., they will never experience the event of interest. Then, the sample of censored observations is an unlabeled mixture of cured and "susceptible" subjects. Using inverse probability censoring weighting (IPCW), we propose a likelihood-based estimation procedure for the cure regression model without making assumptions about the distribution of survival times for the susceptible subjects. The IPCW approach does require a preliminary estimate of the censoring distribution, for which general parametric, semi- or non-parametric approaches can be used. The incorporation of a penalty term in our estimation procedure is straightforward; in particular, we propose L1-type penalties for variable selection. Our theoretical results are derived under mild assumptions. Simulation experiments and real data analysis illustrate the effectiveness of the new approach.

stat.ME↗

Modified Cox regression with current status data

In survival analysis, the lifetime under study is not always observed. In certain applications, for some individuals, the value of the lifetime is only known to be smaller or larger than some random duration. This framework represent an extension of standard situations where the lifetime is only left or only right randomly censored. We consider the case where the independent observation units include also some covariates, and we propose two semiparametric regression models. The new models extend the standard Cox proportional hazard model to the situation of a more complex censoring mechanism. However, like in Cox's model, in both models the nonparametric baseline hazard function still could be expressed as an explicit functional of the distribution of the observations. This allows to define the estimator of the finite-dimensional parameters as the maximum of a likelihood-type criterion which is an explicit function of the data. Given an estimate of the finite-dimensional parameter, the estimation of the baseline cumulative hazard function is straightforward.

math.ST↗

A semiparametric single-index estimator for a class of estimating equation models

We propose a two-step pseudo-maximum likelihood procedure for semiparametric single-index regression models where the conditional variance is a known function of the regression and an additional parameter. The Poisson single-index regression with multiplicative unobserved heterogeneity is an example of such models. Our procedure is based on linear exponential densities with nuisance parameter. The pseudo-likelihood criterion we use contains a nonparametric estimate of the index regression and therefore a rule for choosing the smoothing parameter is needed. We propose an automatic and natural rule based on the joint maximization of the pseudo-likelihood with respect to the index parameter and the smoothing parameter. We derive the asymptotic properties of the semiparametric estimator of the index parameter and the asymptotic behavior of our `optimal' smoothing parameter. The finite sample performances of our methodology are analyzed using simulated and real data.

math.ST↗

A General Approach for Cure Models in Survival Analysis

In survival analysis it often happens that some subjects under study do not experience the event of interest; they are considered to be `cured'. The population is thus a mixture of two subpopulations: the one of cured subjects, and the one of `susceptible' subjects. When covariates are present, a so-called mixture cure model can be used to model the conditional survival function of the population. It depends on two components: the probability of being cured and the conditional survival function of the susceptible subjects. In this paper we propose a novel approach to estimate a mixture cure model when the data are subject to random right censoring. We work with a parametric model for the cure proportion (like e.g. a logistic model), while the conditional survival function of the uncured subjects is unspecified. The approach is based on an inversion which allows to write the survival function as a function of the distribution of the observable random variables. This leads to a very general class of models, which allows a flexible and rich modeling of the conditional survival function. We show the identifiability of the proposed model, as well as the weak consistency and the asymptotic normality of the model parameters. We also consider in more detail the case where kernel estimators are used for the nonparametric part of the model. The new estimators are compared with the estimators from a Cox mixture cure model via finite sample simulations. Finally, we apply the new model and estimation procedure on two medical data sets.

math.ST↗

Nonparametric testing for no-effect with functional responses and functional covariates

This paper examines the problem of nonparametric testing for the no-effect of a random covariate (or predictor) on a functional response. This means testing whether the conditional expectation of the response given the covariate is almost surely zero or not, without imposing any model relating response and covariate. The covariate could be univariate, multivariate or functional. Our test statistic is a quadratic form involving univariate nearest neighbor smoothing and the asymptotic critical values are given by the standard normal law. When the covariate is multidimensional or functional, a preliminary dimension reduction device is used which allows the effect of the covariate to be summarized into a univariate random quantity. The test is able to detect not only linear but nonparametric alternatives. The responses could have conditional variance of unknown form and the law of the covariate does not need to be known. An empirical study with simulated and real data shows that the test performs well in applications.

math.ST↗

Nonparametric model checks of single-index assumptions

Semiparametric single-index assumptions are convenient and widely used dimen\-sion reduction approaches that represent a compromise between the parametric and fully nonparametric models for regressions or conditional laws. In a mean regression setup, the SIM assumption means that the conditional expectation of the response given the vector of covariates is the same as the conditional expectation of the response given a scalar projection of the covariate vector. In a conditional distribution modeling, under the SIM assumption the conditional law of a response given the covariate vector coincides with the conditional law given a linear combination of the covariates. Several estimation techniques for single-index models are available and commonly used in applications. However, the problem of testing the goodness-of-fit seems less explored and the existing proposals still have some major drawbacks. In this paper, a novel kernel-based approach for testing SIM assumptions is introduced. The covariate vector needs not have a density and only the index estimated under the SIM assumption is used in kernel smoothing. Hence the effect of high-dimensional covariates is mitigated while asymptotic normality of the test statistic is obtained. Irrespective of the fixed dimension of the covariate vector, the new test detects local alternatives approaching the null hypothesis slower than $n^{-1/2}h^{-1/4},$ where $h$ is the bandwidth used to build the test statistic and $n$ is the sample size. A wild bootstrap procedure is proposed for finite sample corrections of the asymptotic critical values. The small sample performances of our test compared to existing procedures are illustrated through simulations.

math.ST↗

Testing for the significance of functional covariates in regression models

Regression models with a response variable taking values in a Hilbert space and hybrid covariates are considered. This means two sets of regressors are allowed, one of finite dimension and a second one functional with values in a Hilbert space. The problem we address is the test of the effect of the functional covariates. This problem occurs for instance when checking the goodness-of-fit of some regression models for functional data. The significance test for functional regressors in nonparametric regression with hybrid covariates and scalar or functional responses is another example where the core problem is the test on the effect of functional covariates. We propose a new test based on kernel smoothing. The test statistic is asymptotically standard normal under the null hypothesis provided the smoothing parameter tends to zero at a suitable rate. The one-sided test is consistent against any fixed alternative and detects local alternatives à la Pitman approaching the null hypothesis. In particular we show that neither the dimension of the outcome nor the dimension of the functional covariates influences the theoretical power of the test against such local alternatives. Simulation experiments and a real data application illustrate the performance of the new test with finite samples.

math.ST↗

Powerful nonparametric checks for quantile regression

We address the issue of lack-of-fit testing for a parametric quantile regression. We propose a simple test that involves one-dimensional kernel smoothing, so that the rate at which it detects local alternatives is independent of the number of covariates. The test has asymptotically gaussian critical values, and wild bootstrap can be applied to obtain more accurate ones in small samples. Our procedure appears to be competitive with existing ones in simulations. We illustrate the usefulness of our test on birthweight data.

math.ST↗

A Significance Test for Covariates in Nonparametric Regression

We consider testing the significance of a subset of covariates in a nonparametric regression. These covariates can be continuous and/or discrete. We propose a new kernel-based test that smoothes only over the covariates appearing under the null hypothesis, so that the curse of dimensionality is mitigated. The test statistic is asymptotically pivotal and the rate of which the test detects local alternatives depends only on the dimension of the covariates under the null hypothesis. We show the validity of wild bootstrap for the test. In small samples, our test is competitive compared to existing procedures.

math.ST↗

Single index regression models in the presence of censoring depending on the covariates

Consider a random vector (X',Y)', where X is d-dimensional and Y is one-dimensional. We assume that Y is subject to random right censoring. The aim of this paper is twofold. First, we propose a new estimator of the joint distribution of (X',Y)'. This estimator overcomes the common curse-of-dimensionality problem, by using a new dimension reduction technique. Second, we assume that the relation between X and Y is given by a mean regression single index model, and propose a new estimator of the parameters in this model. The asymptotic properties of all proposed estimators are obtained.

math.ST↗

Testing second order dynamics for autoregressive processes in presence of time-varying variance

The volatility modeling for autoregressive univariate time series is considered. A benchmark approach is the stationary ARCH model of Engle (1982). Motivated by real data evidence, processes with non constant unconditional variance and ARCH effects have been recently introduced. We take into account such possible non stationarity and propose simple testing procedures for ARCH effects. Adaptive McLeod and Li's portmanteau and ARCH-LM tests for checking for second order dynamics are provided. The standard versions of these tests, commonly used by practitioners, suppose constant unconditional variance. We prove the failure of these standard tests with time-varying unconditional variance. The theoretical results are illustrated by mean of simulated and real data.

stat.ME↗

Projection-based nonparametric goodness-of-fit testing with functional covariates

This paper studies the problem of nonparametric testing for the effect of a random functional covariate on a real-valued error term. The covariate takes values in $L^2[0,1]$, the Hilbert space of the square-integrable real-valued functions on the unit interval. The error term could be directly observed as a response or \emph{estimated} from a functional parametric model, like for instance the functional linear regression. Our test is based on the remark that checking the no-effect of the functional covariate is equivalent to checking the nullity of the conditional expectation of the error term given a sufficiently rich set of projections of the covariate. Such projections could be on elements of norm 1 from finite-dimension subspaces of $L^2[0,1]$. Next, the idea is to search a finite-dimension element of norm 1 that is, in some sense, the least favorable for the null hypothesis. Finally, it remains to perform a nonparametric check of the nullity of the conditional expectation of the error term given the scalar product between the covariate and the selected least favorable direction. For such finite-dimension search and nonparametric check we use a kernel-based approach. As a result, our test statistic is a quadratic form based on univariate kernel smoothing and the asymptotic critical values are given by the standard normal law. The test is able to detect nonparametric alternatives, including the polynomial ones. The error term could present heteroscedasticity of unknown form. We do no require the law of the covariate $X$ to be known. The test could be implemented quite easily and performs well in simulations and real data applications. We illustrate the performance of our test for checking the functional linear regression model.

math.ST↗

A uniform Berry--Esseen theorem on $M$-estimators for geometrically ergodic Markov chains

Let $\{X_n\}_{n\ge0}$ be a $V$-geometrically ergodic Markov chain. Given some real-valued functional $F$, define $M_n(α):=n^{-1}\sum_{k=1}^nF(α,X_{k-1},X_k)$, $α\in\mathcal{A}\subset \mathbb {R}$. Consider an $M$ estimator $\hatα_n$, that is, a measurable function of the observations satisfying $M_n(\hatα_n)\leq \min_{α\in\mathcal{A}}M_n(α)+c_n$ with $\{c_n\}_{n\geq1}$ some sequence of real numbers going to zero. Under some standard regularity and moment assumptions, close to those of the i.i.d. case, the estimator $\hatα_n$ satisfies a Berry--Esseen theorem uniformly with respect to the underlying probability distribution of the Markov chain.

math.ST↗

Semiparametric efficiency bounds for seemingly unrelated conditional moment restrictions

This paper addresses the problem of semiparametric efficiency bounds for conditional moment restriction models with different conditioning variables. We characterize such an efficiency bound, that in general is not explicit, as a limit of explicit efficiency bounds for a decreasing sequence of unconditional (marginal) moment restriction models. An iterative procedure for approximating the efficient score when this is not explicit is provided. Our theoretical results complete and extend existing results in the literature, provide new insight for the theory of semiparametric efficiency bounds literature and open the door to new applications. In particular, we investigate a class of regression-like (mean regression, quantile regression,...) models with missing data.

math.ST↗

Corrected portmanteau tests for VAR models with time-varying variance

The problem of test of fit for Vector AutoRegressive (VAR) processes with unconditionally heteroscedastic errors is studied. The volatility structure is deterministic but time-varying and allows for changes that are commonly observed in economic or financial multivariate series. Our analysis is based on the residual autocovariances and autocorrelations obtained from Ordinary Least Squares (OLS), Generalized Least Squares (GLS)and Adaptive Least Squares (ALS) estimation of the autoregressive parameters. The ALS approach is the GLS approach adapted to the unknown time-varying volatility that is then estimated by kernel smoothing. The properties of the three types of residual autocovariances and autocorrelations are derived. In particular it is shown that the ALS and GLS residual autocorrelations are asymptotically equivalent. It is also found that the asymptotic distribution of the OLS residual autocorrelations can be quite different from the standard chi-square asymptotic distribution obtained in a correctly specified VAR model with iid innovations. As a consequence the standard portmanteau tests are unreliable in our framework. The correct critical values of the standard portmanteau tests based on the OLS residuals are derived. Moreover, modified portmanteau statistics based on ALS residual autocorrelations are introduced. Portmanteau tests with modified statistics based on OLS and ALS residuals and standard chi-square asymptotic distributions under the null hypothesis are also proposed. An extension of our portmanteau approaches to testing the lag length in a vector error correction type model for co-integrating relations is briefly investigated. The finite sample properties of the goodness-of-fit tests we consider are investigated by Monte Carlo experiments. The theoretical results are also illustrated using two U.S. economic data sets.

stat.ME↗

Adaptive estimation of vector autoregressive models with time-varying variance: application to testing linear causality in mean

Linear Vector AutoRegressive (VAR) models where the innovations could be unconditionally heteroscedastic and serially dependent are considered. The volatility structure is deterministic and quite general, including breaks or trending variances as special cases. In this framework we propose Ordinary Least Squares (OLS), Generalized Least Squares (GLS) and Adaptive Least Squares (ALS) procedures. The GLS estimator requires the knowledge of the time-varying variance structure while in the ALS approach the unknown variance is estimated by kernel smoothing with the outer product of the OLS residuals vectors. Different bandwidths for the different cells of the time-varying variance matrix are also allowed. We derive the asymptotic distribution of the proposed estimators for the VAR model coefficients and compare their properties. In particular we show that the ALS estimator is asymptotically equivalent to the infeasible GLS estimator. This asymptotic equivalence is obtained uniformly with respect to the bandwidth(s) in a given range and hence justifies data-driven bandwidth rules. Using these results we build Wald tests for the linear Granger causality in mean which are adapted to VAR processes driven by errors with a non stationary volatility. It is also shown that the commonly used standard Wald test for the linear Granger causality in mean is potentially unreliable in our framework. Monte Carlo experiments illustrate the use of the different estimation approaches for the analysis of VAR models with stable innovations.

stat.ME↗

Product-limit estimators of the survival function with twice censored data

A model for competing (resp. complementary) risks survival data where the failure time can be left (resp. right) censored is proposed. Product-limit estimators for the survival functions of the individual risks are derived. We deduce the strong convergence of our estimators on the whole real half-line without any additional assumptions and their asymptotic normality under conditions concerning only the observed distribution. When the observations are generated according to the double censoring model introduced by Turnbull, the product-limit estimators represent upper and lower bounds for Turnbull's estimator.

math.ST↗