SearcharxivSearch

arXiv subjects

Ingrid Van Keilegom

Publications and source records attributed to Ingrid Van Keilegom.

At least 19 recordsLinked to original sources

High-dimensional censored MIDAS logistic regression for corporate survival forecasting

This paper addresses the challenge of forecasting corporate distress, a problem marked by three key statistical hurdles: (i) right censoring, (ii) high-dimensional predictors, and (iii) mixed-frequency data. To overcome these complexities, we introduce a novel high-dimensional censored MIDAS (Mixed Data Sampling) logistic regression. Our approach handles censoring through inverse probability weighting and achieves accurate estimation with numerous mixed-frequency predictors by employing a sparse-group penalty. We establish finite-sample bounds for the estimation error, accounting for censoring, MIDAS approximation error, and heavy tails. For statistical inference, we develop a de-sparsified version of the proposed penalized estimator and establish its asymptotic theory, which enables valid statistical inference in high-dimensional settings with censoring. We show that censoring induces a nonstandard variance structure for the de-sparsified estimator, a feature that, to the best of our knowledge, has not been studied in the existing literature. The superior performance of the method is demonstrated through Monte Carlo simulations. Finally, we present an extensive application of our methodology to predict the financial distress of Chinese-listed firms and to identify covariates that are statistically significant for predicting distress. Our novel procedure is implemented in the R package \texttt{Survivalml}.

econ.EM

Copula based dependent censoring in cure models with covariates

In survival analysis, the time-to-event variable T is frequently subject to right censoring. Individuals may withdraw from the study for various reasons, or may not experience the event of interest before the end of follow-up. In this paper, we distinguish between two types of censoring: a potentially dependent censoring time C, which may be stochastically related to T, and an independent administrative censoring time A. In addition, the data may exhibit a cure fraction, meaning that some individuals will never experience the event. We build upon a recent work about a fully parametric mixture cure model, which accounts for dependent censoring through copulas. The proposed extension incorporates administrative censoring and allows covariates to affect all model parameters. This framework enables a more accurate modelling of the dependence between survival and censoring times while providing greater flexibility through covariate effects, leading to more individualised estimation of the cure fraction, the dependence structure, and other clinically relevant quantities. Moreover, the presence of covariates allows for weaker identification conditions.

stat.ME

Hybrid combinations of parametric and empirical likelihoods

This paper develops a hybrid likelihood (HL) method based on a compromise between parametric and nonparametric likelihoods. Consider the setting of a parametric model for the distribution of an observation $Y$ with parameter $θ$. Suppose there is also an estimating function $m(\cdot,μ)$ identifying another parameter $μ$ via $E\,m(Y,μ)=0$, at the outset defined independently of the parametric model. To borrow strength from the parametric model while obtaining a degree of robustness from the empirical likelihood method, we formulate inference about $θ$ in terms of the hybrid likelihood function $H_n(θ)=L_n(θ)^{1-a}R_n(μ(θ))^a$. Here $a\in[0,1)$ represents the extent of the compromise, $L_n$ is the ordinary parametric likelihood for $θ$, $R_n$ is the empirical likelihood function, and $μ$ is considered through the lens of the parametric model. We establish asymptotic normality of the corresponding HL estimator and a version of the Wilks theorem. We also examine extensions of these results under misspecification of the parametric model, and propose methods for selecting the balance parameter $a$.

stat.ME

Testing for sufficient follow-up in cure models with categorical covariates

In survival analysis, estimating the fraction of 'immune' or 'cured' subjects who will never experience the event of interest, requires a sufficiently long follow-up period. A few statistical tests have been proposed to test the assumption of sufficient follow-up, i.e. whether the right extreme of the censoring distribution exceeds that of the survival time of the uncured subjects. However, in practice the problem remains challenging. To address this, a relaxed notion of 'practically' sufficient follow-up has been introduced recently, suggesting that the follow-up would be considered sufficiently long if the probability for the event occurring after the end of the study is very small. All these existing tests do not incorporate covariate information, which might affect the cure rate and the survival times. We extend the test for 'practically' sufficient follow-up to settings with categorical covariates. While a straightforward intersection-union type test could reject the null hypothesis of insufficient follow-up only if such hypothesis is rejected for all covariate values, in practice this approach is overly conservative and lacks power. To improve upon this, we propose a novel test procedure that relies on the test decision for one properly chosen covariate value. Our approach relies on the assumption that the conditional density of the uncured survival time is a non-increasing function of time in the tail region. We show that both methods yield tests of asymptotically level $α$ and investigate their finite sample performance through simulations. The practical application of the methods is illustrated using a skin melanoma dataset.

stat.ME

Tests of exogeneity in duration models with censored data

Consider the setting in which a researcher is interested in the causal effect of a treatment $Z$ on a duration time $T$, which is subject to right censoring. We assume that $T=φ(X,Z,U)$, where $X$ is a vector of baseline covariates, $φ(X,Z,U)$ is strictly increasing in the error term $U$ for each $(X,Z)$ and $U\sim \mathcal{U}[0,1]$. Therefore, the model is nonparametric and nonseparable. We propose nonparametric tests for the hypothesis that $Z$ is exogenous, meaning that $Z$ is independent of $U$ given $X$. The test statistics rely on an instrumental variable $W$ that is independent of $U$ given $X$. We assume that $X,W$ and $Z$ are all categorical. Test statistics are constructed for the hypothesis that the conditional rank $V_T= F_{T \mid X,Z}(T \mid X,Z)$ is independent of $(X,W)$ jointly. Under an identifiability condition on $φ$, this hypothesis is equivalent to $Z$ being exogenous. However, note that $V_T$ is censored by $V_C =F_{T \mid X,Z}(C \mid X,Z)$, which complicates the construction of the test statistics significantly. We derive the limiting distributions of the proposed tests and prove that our estimator of the distribution of $V_T$ converges to the uniform distribution at a rate faster than the usual parametric $n^{-1/2}$-rate. We demonstrate that the test statistics and bootstrap approximations for the critical values have a good finite sample performance in various Monte Carlo settings. Finally, we illustrate the tests with an empirical application to the National Job Training Partnership Act (JTPA) Study.

econ.EM

Survival analysis under label shift

Let P represent the source population with complete data, containing covariate $\mathbf{Z}$ and response $T$, and Q the target population, where only the covariate $\mathbf{Z}$ is available. We consider a setting with both label shift and label censoring. Label shift assumes that the marginal distribution of $T$ differs between $P$ and $Q$, while the conditional distribution of $\mathbf{Z}$ given $T$ remains the same. Label censoring refers to the case where the response $T$ in $P$ is subject to random censoring. Our goal is to leverage information from the label-shifted and label-censored source population $P$ to conduct statistical inference in the target population $Q$. We propose a parametric model for $T$ given $\mathbf{Z}$ in $Q$ and estimate the model parameters by maximizing an approximate likelihood. This allows for statistical inference in $Q$ and accommodates a range of classical survival models. Under the label shift assumption, the likelihood depends not only on the unknown parameters but also on the unknown distribution of $T$ in $P$ and $\mathbf{Z}$ in $Q$, which we estimate nonparametrically. The asymptotic properties of the estimator are rigorously established and the effectiveness of the method is demonstrated through simulations and a real data application. This work is the first to combine survival analysis with label shift, offering a new research direction in this emerging topic.

stat.ME

Estimation of the complier causal hazard ratio under dependent censoring

In this work, we are interested in studying the causal effect of an endogenous binary treatment on a dependently censored duration outcome. By dependent censoring, it is meant that the duration time ($T$) and right censoring time ($C$) are not statistically independent of each other, even after conditioning on the measured covariates. The endogeneity issue is handled by making use of a binary instrumental variable for the treatment. To deal with the dependent censoring problem, it is assumed that on the stratum of compliers: (i) $T$ follows a semiparametric proportional hazards model; (ii) $C$ follows a fully parametric model; and (iii) the relation between $T$ and $C$ is modeled by a parametric copula, such that the association parameter can be left unspecified. In this framework, the treatment effect of interest is the complier causal hazard ratio (CCHR). We devise an estimation procedure that is based on a weighted maximum likelihood approach, where the weights are the probabilities of an observation coming from a complier. The weights are estimated non-parametrically in a first stage, followed by the estimation of the CCHR. Novel conditions under which the model is identifiable are given, a two-step estimation procedure is proposed and some important asymptotic properties are established. Simulations are used to assess the validity and finite-sample performance of the estimation procedure. Finally, we apply the approach to estimate the CCHR of both job training programs on unemployment duration and periodic screening examinations on time until death from breast cancer. The data come from the National Job Training Partnership Act study and the Health Insurance Plan of Greater New York experiment respectively.

econ.EM

Non-parametric cure models through extreme-value tail estimation

In survival analysis, the estimation of the proportion of subjects who will never experience the event of interest, termed the cure rate, has received considerable attention recently. Its estimation can be a particularly difficult task when follow-up is not sufficient, that is when the censoring mechanism has a smaller support than the distribution of the target data. In the latter case, non-parametric estimators were recently proposed using extreme value methodology, assuming that the distribution of the susceptible population is in the Fréchet or Gumbel max-domains of attraction. In this paper, we take the extreme value techniques one step further, to jointly estimate the cure rate and the extreme value index, using probability plotting methodology, and in particular using the full information contained in the top order statistics. In other words, under sufficient or insufficient follow-up, we reconstruct the immune proportion. To this end, a Peaks-over-Threshold approach is proposed under the Gumbel max-domain assumption. Next, the approach is also transferred to more specific models such as Pareto, log-normal and Weibull tail models, allowing to recognize the most important tail characteristics of the susceptible population. We establish the asymptotic behavior of our estimators under regularization. Though simulation studies, our estimators are show to rival and often outperform established models, even when purely considering cure rate estimation. Finally, we provide an application of our method to Norwegian birth registry data.

math.ST

A flexible control function approach for survival data subject to different types of censoring

This paper addresses the problem of identifying and estimating the causal effect of a treatment in the presence of unmeasured confounding and various types of right-censoring. Examples of these censoring mechanisms are administrative censoring, competing risks and dependent censoring (e.g. loss to follow-up). Different parametric transformations are applied to each event time, resulting in a regression model with a more additive structure and error terms that are approximately normal and homoscedastic. The transformed event times are modeled using a joint regression framework, assuming multivariate Gaussian error terms with an unspecified covariance matrix. A control function approach is used to deal with unmeasured confounding. The model is shown to be identifiable and a two-step estimation procedure is proposed. This estimator is proven to yield consistent and asymptotically normal estimates. Furthermore, a goodness-of-fit test for the model's validity is developed. Simulations are conducted to examine the finite-sample performance of the proposed estimator under various scenarios. Finally, the methodology is applied to investigate the causal effect of job training programs on unemployment duration using data from the National Job Training Partnership Act (JTPA) study.

math.ST

Bounds for the regression parameters in dependently censored survival models

We propose a semiparametric model to study the effect of covariates on the distribution of a censored event time while making minimal assumptions about the censoring mechanism. The result is a partially identified model, in the sense that we obtain bounds on the covariate effects, which are allowed to be time-dependent. Moreover, these bounds can be interpreted as classical confidence intervals and are obtained by aggregating information in the conditional Peterson bounds over the entire covariate space. As a special case, our approach can be used to study the popular Cox proportional hazards model while leaving the censoring distribution as well as its dependence with the time of interest completely unspecified. A simulation study illustrates good finite sample performance of the method, and several data applications in both economics and medicine demonstrate its practicability on real data. All developed methodology is implemented in R and made available in the package depCensoring.

stat.ME

Quantile regression under dependent censoring with unknown association

The study of survival data often requires taking proper care of the censoring mechanism that prohibits complete observation of the data. Under right censoring, only the first occurring event is observed: either the event of interest, or a competing event like withdrawal of a subject from the study. The corresponding identifiability difficulties led many authors to imposing (conditional) independence or a fully known dependence between survival and censoring times, both of which are not always realistic. However, recent results in survival literature showed that parametric copula models allow identification of all model parameters, including the association parameter, under appropriately chosen marginal distributions. The present paper is the first one to apply such models in a quantile regression context, hence benefiting from its well-known advantages in terms of e.g. robustness and richer inference results. The parametric copula is supplemented with a likewise parametric, yet flexible, enriched asymmetric Laplace distribution for the survival times conditional on the covariates. Its asymmetric Laplace basis provides its close connection to quantiles, while the extension with Laguerre orthogonal polynomials ensures sufficient flexibility for increasing polynomial degrees. The distributional flavour of the quantile regression presented, comes with advantages of both theoretical and computational nature. All model parameters are proven to be identifiable, consistent, and asymptotically normal. Finally, performance of the model and of the proposed estimation procedure is assessed through extensive simulation studies as well as an application on liver transplant data.

math.ST

Copula based dependent censoring in cure models

In this paper we consider a time-to-event variable $T$ that is subject to random right censoring, and we assume that the censoring time $C$ is stochastically dependent on $T$ and that there is a positive probability of not observing the event. There are various situations in practice where this happens, and appropriate models and methods need to be considered to avoid biased estimators of the survival function or incorrect conclusions in clinical trials. We consider a fully parametric model for the bivariate distribution of $(T,C)$, that takes these features into account. The model depends on a parametric copula (with unknown association parameter) and on parametric marginal distributions for $T$ and $C$. Sufficient conditions are developed under which the model is identified, and an estimation procedure is proposed. In particular, our model allows to identify and estimate the association between $T$ and $C$, even though only the smallest of these variables is observable. The finite sample performance of the estimated parameters is illustrated by means of a thorough simulation study and the analysis of breast cancer data.

stat.ME

Nonparametric covariate hypothesis tests for the cure rate in mixture cure models

In lifetime data, like cancer studies, theremay be long term survivors, which lead to heavy censoring at the end of the follow-up period. Since a standard survival model is not appropriate to handle these data, a cure model is needed. In the literature, covariate hypothesis tests for cure models are limited to parametric and semiparametric methods.We fill this important gap by proposing a nonparametric covariate hypothesis test for the probability of cure in mixture cure models. A bootstrap method is proposed to approximate the null distribution of the test statistic. The procedure can be applied to any type of covariate, and could be extended to the multivariate setting. Its efficiency is evaluated in a Monte Carlo simulation study. Finally, the method is applied to a colorectal cancer dataset.

stat.ME

Nonparametric incidence estimation and bootstrap bandwidth selection in mixture cure models

A completely nonparametric method for the estimation of mixture cure models is proposed. A nonparametric estimator of the incidence is extensively studied and a nonparametric estimator of the latency is presented. These estimators, which are based on the Beran estimator of the conditional survival function, are proved to be the local maximum likelihood estimators. An i.i.d. representation is obtained for the nonparametric incidence estimator. As a consequence, an asymptotically optimal bandwidth is found. Moreover, a bootstrap bandwidth selection method for the nonparametric incidence estimator is proposed. The introduced nonparametric estimators are compared with existing semiparametric approaches in a simulation study, in which the performance of the bootstrap bandwidth selector is also assessed. Finally, the method is applied to a database of colorectal cancer from the University Hospital of A Coruña (CHUAC).

stat.ME

Instrumental variable estimation of the proportional hazards model by presmoothing

We consider instrumental variable estimation of the proportional hazards model of Cox (1972). The instrument and the endogenous variable are discrete but there can be (possibly continuous) exogenous covariables. By making a rank invariance assumption, we can reformulate the proportional hazards model into a semiparametric version of the instrumental variable quantile regression model of Chernozhukov and Hansen (2005). A naïve estimation approach based on conditional moment conditions generated by the model would lead to a highly nonconvex and nonsmooth objective function. To overcome this problem, we propose a new presmoothing methodology. First, we estimate the model nonparametrically - and show that this nonparametric estimator has a closed-form solution in the leading case of interest of randomized experiments with one-sided noncompliance. Second, we use the nonparametric estimator to generate ``proxy'' observations for which exogeneity holds. Third, we apply the usual partial likelihood estimator to the ``proxy'' data. While the paper focuses on the proportional hazards model, our presmoothing approach could be applied to estimate other semiparametric formulations of the instrumental variable quantile regression model. Our estimation procedure allows for random right-censoring. We show asymptotic normality of the resulting estimator. The approach is illustrated via simulation studies and an empirical application to the Illinois

econ.EM

Testing for sufficient follow-up in censored survival data by using extremes

In survival analysis, it often happens that some individuals, referred to as cured individuals, never experience the event of interest. When analyzing time-to-event data with a cure fraction, it is crucial to check the assumption of `sufficient follow-up', which means that the right extreme of the censoring time distribution is larger than that of the survival time distribution for the non-cured individuals. However, the available methods to test this assumption are limited in the literature. In this article, we study the problem of testing whether follow-up is sufficient for light-tailed distributions and develop a simple novel test. The proposed test statistic compares an estimator of the non-cure proportion under sufficient follow-up to one without the assumption of sufficient follow-up. A bootstrap procedure is employed to approximate the critical values of the test. We also carry out extensive simulations to evaluate the finite sample performance of the test and illustrate the practical use with applications to leukemia and breast cancer datasets.

stat.ME

Testing for homogeneous treatment effects in linear and nonparametric instrumental variable models

The hypothesis of homogeneous treatment effects is central to the instrumental variables literature. This assumption signifies that treatment effects are constant across all subjects. It allows to interpret instrumental variable estimates as average treatment effects over the whole population of the study. When this assumption does not hold, the bias of instrumental variable estimators can be larger than that of naive estimators ignoring endogeneity. This paper develops two tests for the assumption of homogeneous treatment effects when the treatment is endogenous and an instrumental variable is available. The tests leverage a covariable that is (jointly with the error terms) independent of a coordinate of the instrument. This covariate does not need to be exogenous. The first test assumes that the potential outcomes are linear in the regressors and is computationally simple. The second test is nonparametric and relies on Tikhonov regularization. The treatment can be either discrete or continuous. We show that the tests have asymptotically correct level and asymptotic power equal to one against a range of alternatives. Simulations demonstrate that the proposed tests attain excellent finite sample performances. The methodology is also applied to the evaluation of returns to schooling and the effect of price on demand in a fish market.

econ.EM

Comparison of Quantile Regression Curves with Censored Data

This paper proposes a new test for the comparison of conditional quantile curves when the outcome of interest, typically a duration, is subject to right censoring. The test can be applied both in the case of two independent samples and for paired data, and can be used for the comparison of quantiles at a fixed quantile level, a finite set of levels or a range of quantile levels. The asymptotic distribution of the proposed test statistics is obtained both under the null hypothesis and under local alternatives. We describe a bootstrap procedure in order to approximate the critical values, and present the results of a simulation study, in which the performance of the tests for small and moderate sample sizes is studied and compared with the behavior of alternative tests. Finally, we apply the proposed tests on a data set concerning diabetic retinopathy.

stat.ME