SearcharxivSearch

arXiv subjects

Suvra Pal

Publications and source records attributed to Suvra Pal.

15 recordsLinked to original sources

A Sequential Quadratic Hamiltonian-Based Estimation Method for Box-Cox Transformation Cure Model

We propose an enhanced estimation method for the Box-Cox transformation (BCT) cure rate model parameters by introducing a generic maximum likelihood estimation algorithm, the sequential quadratic Hamiltonian (SQH) scheme, which is based on a gradient-free approach. We apply the SQH algorithm to the BCT cure model and, through an extensive simulation study, compare its model fitting results with those obtained using the recently developed non-linear conjugate gradient (NCG) algorithm. Since the NCG method has already been shown to outperform the well-known expectation maximization algorithm, our focus is on demonstrating the superiority of the SQH algorithm over NCG. First, we show that the SQH algorithm produces estimates with smaller bias and root mean square error for all BCT cure model parameters, resulting in more accurate and precise cure rate estimates. We then demonstrate that, being gradient-free, the SQH algorithm requires less CPU time to generate estimates compared to the NCG algorithm, which only computes the gradient and not the Hessian. These advantages make the SQH algorithm the preferred estimation method over the NCG method for the BCT cure model. Finally, we apply the SQH algorithm to analyze a well-known melanoma dataset and present the results.

stat.CO

A New Cure Rate Model with Discrete and Multiple Exposures

Cure rate models are mostly used to study data arising from cancer clinical trials. Its use in the context of infectious diseases has not been explored well. In 2008, Tournoud and Ecochard first proposed a mechanistic formulation of cure rate model in the context of infectious diseases with multiple exposures to infection. However, they assumed a simple Poisson distribution to capture the unobserved pathogens at each exposure time. In this paper, we propose a new cure rate model to study infectious diseases with discrete multiple exposures to infection. Our formulation captures both over-dispersion and under-dispersion with respect to the count on pathogens at each time of exposure. We also propose a new estimation method based on the expectation maximization algorithm to calculate the maximum likelihood estimates of the model parameters. We carry out a detailed Monte Carlo simulation study to demonstrate the performance of the proposed model and estimation algorithm. The flexibility of our proposed model also allows us to carry out a model discrimination. For this purpose, we use both likelihood ratio test and information-based criteria. Finally, we illustrate our proposed model using a recently collected data on COVID-19.

stat.ME

Likelihood-Based Inference for Semi-Parametric Transformation Cure Models with Interval Censored Data

A simple yet effective way of modeling survival data with cure fraction is by considering Box-Cox transformation cure model (BCTM) that unifies mixture and promotion time cure models. In this article, we numerically study the statistical properties of the BCTM when applied to interval censored data. Time-to-events associated with susceptible subjects are modeled through proportional hazards structure that allows for non-homogeneity across subjects, where the baseline hazard function is estimated by distribution-free piecewise linear function with varied degrees of non-parametricity. Due to missing cured statuses for right censored subjects, maximum likelihood estimates of model parameters are obtained by developing an expectation-maximization (EM) algorithm. Under the EM framework, the conditional expectation of the complete data log-likelihood function is maximized by considering all parameters (including the Box-Cox transformation parameter $α$) simultaneously, in contrast to conventional profile-likelihood technique of estimating $α$. The robustness and accuracy of the model and estimation method are established through a detailed simulation study under various parameter settings, and an analysis of real-life data obtained from a smoking cessation study.

stat.ME

Parametric quantile autoregressive conditional duration models with application to intraday value-at-risk

The modeling of high-frequency data that qualify financial asset transactions has been an area of relevant interest among statisticians and econometricians -- above all, the analysis of time series of financial durations. Autoregressive conditional duration (ACD) models have been the main tool for modeling financial transaction data, where duration is usually defined as the time interval between two successive events. These models are usually specified in terms of a time-varying mean (or median) conditional duration. In this paper, a new extension of ACD models is proposed which is built on the basis of log-symmetric distributions reparametrized by their quantile. The proposed quantile log-symmetric conditional duration autoregressive model allows us to model different percentiles instead of the traditionally used conditional mean (or median) duration. We carry out an in-depth study of theoretical properties and practical issues, such as parameter estimation using maximum likelihood method and diagnostic analysis based on residuals. A detailed Monte Carlo simulation study is also carried out to evaluate the performance of the proposed models and estimation method in retrieving the true parameter values as well as to evaluate a form of residuals. Finally, the proposed class of models is applied to a price duration data set and then used to derive a semi-parametric intraday value-at-risk (IVaR) model.

stat.ME

Bivariate autoregressive conditional models: A new method for jointly modeling duration and number of transactions of irregularly spaced financial data

In this paper, a new approach to bivariate modeling of autoregressive conditional duration (ACD) models is proposed. Specifically, we consider the joint modeling of durations and the number of transactions made during the spell. The proposed bivariate ACD model is based on log-symmetric distributions, which are useful for modeling strictly positive, asymmetric and light- and heavy-tailed data, such as transaction-level high-frequency financial data. A Monte Carlo simulation is performed for the assessment of the estimation method and the evaluation of a form of residuals. A real financial transactions data set is analyzed in order to illustrate the proposed method.

stat.AP

A Semi-parametric Promotion Time Cure Model with Support Vector Machine

The promotion time cure rate model (PCM) is an extensively studied model for the analysis of time-to-event data in the presence of a cured subgroup. There are several strategies proposed in the literature to model the latency part of PCM. However, there aren't many strategies proposed to investigate the effects of covariates on the incidence part of PCM. In this regard, most existing studies assume the boundary separating the cured and non-cured subjects with respect to the covariates to be linear. As such, they can only capture simple effects of the covariates on the cured/non-cured probability. In this manuscript, we propose a new promotion time cure model that uses the support vector machine (SVM) to model the incidence part. The proposed model inherits the features of the SVM and provides flexibility in capturing non-linearity in the data. To the best of our knowledge, this is the first work that integrates the SVM with PCM model. For the estimation of model parameters, we develop an expectation maximization algorithm where we make use of the sequential minimal optimization technique together with the Platt scaling method to obtain the posterior probabilities of cured/uncured. A detailed simulation study shows that the proposed model outperforms the existing logistic regression-based PCM model as well as the spline regression-based PCM model, which is also known to capture non linearity in the data. This is true in terms of bias and mean square error of different quantities of interest, and also in terms of predictive and classification accuracies of cure. Finally, we illustrate the applicability and superiority of our model using the data from a study on leukemia patients who went through bone marrow transplantation.

stat.ME

Scale-mixture Birnbaum-Saunders quantile regression models applied to personal accident insurance data

The modeling of personal accident insurance data has been a topic of extreme relevance in the insurance literature. This kind of data often exhibits positive skewness and heavy tails. In this work, we propose a new quantile regression model based on the scale-mixture Birnbaum-Saunders distribution for modeling personal accident insurance data. The maximum likelihood estimates of the model parameters are obtained via the EM algorithm. Two Monte Carlo simulation studies are performed using the R software. The first study aims to analyze the performances of the EM algorithm to obtain the maximum likelihood estimates, and the randomized quantile and generalized Cox-Snell residuals. In the second simulation study, the size and power of the the Wald, likelihood ratio, score and gradient tests are evaluated. The two simulation studies are conducted considering different quantiles of interest and sample sizes. Finally, a real insurance data set is analyzed to illustrate the proposed approach.

stat.ME

Optimal personalized therapies in colon-cancer induced immune response using a Fokker-Planck framework

In this paper, a new stochastic framework to determine optimal combination therapies in colon cancer-induced immune response is presented. The dynamics of colon cancer is described through an Itö stochastic process, whose probability density function evolution is governed by the Fokker-Planck equation. An open-loop control optimization problem is proposed to determine the optimal combination therapies. Numerical results with combination therapies comprising of the chemotherapy drug \ind{Doxorubicin} and immunotherapy drug IL-2 validate the proposed framework.

math.OC

A Support Vector Machine Based Cure Rate Model For Interval Censored Data

The mixture cure rate model is the most commonly used cure rate model in the literature. In the context of mixture cure rate model, the standard approach to model the effect of covariates on the cured or uncured probability is to use a logistic function. This readily implies that the boundary classifying the cured and uncured subjects is linear. In this paper, we propose a new mixture cure rate model based on interval censored data that uses the support vector machine (SVM) to model the effect of covariates on the uncured or the cured probability (i.e., on the incidence part of the model). Our proposed model inherits the features of the SVM and provides flexibility to capture classification boundaries that are non-linear and more complex. Furthermore, the new model can be used to model the effect of covariates on the incidence part when the dimension of covariates is high. The latency part is modeled by a proportional hazards structure. We develop an estimation procedure based on the expectation maximization (EM) algorithm to estimate the cured/uncured probability and the latency model parameters. Our simulation study results show that the proposed model performs better in capturing complex classification boundaries when compared to the existing logistic regression based mixture cure rate model. We also show that our model's ability to capture complex classification boundaries improve the estimation results corresponding to the latency parameters. For illustrative purpose, we present our analysis by applying the proposed methodology to an interval censored data on smoking cessation.

stat.ME

Parametric quantile autoregressive moving average models with exogenous terms applied to Walmart sales data

Parametric autoregressive moving average models with exogenous terms (ARMAX) have been widely used in the literature. Usually, these models consider a conditional mean or median dynamics, which limits the analysis. In this paper, we introduce a class of quantile ARMAX models based on log-symmetric distributions. This class is indexed by quantile and dispersion parameters. It not only accommodates the possibility to model bimodal and/or light/heavy-tailed distributed data but also accommodates heteroscedasticity. We estimate the model parameters by using the conditional maximum likelihood method. Furthermore, we carry out an extensive Monte Carlo simulation study to evaluate the performance of the proposed models and the estimation method in retrieving the true parameter values. Finally, the proposed class of models and the estimation method are applied to a dataset on the competition "M5 Forecasting - Accuracy" that corresponds to the daily sales history of several Walmart products. The results indicate that the proposed log-symmetric quantile ARMAX models have good performance in terms of model fitting and forecasting.

stat.ME

A Fokker-Planck feedback control framework for optimal personalized therapies in colon cancer-induced angiogenesis

In this paper, a new framework for obtaining personalized optimal treatment strategies in colon cancer-induced angiogenesis is presented. The dynamics of colon cancer is given by a Itó stochastic process, which helps in modeling the randomness present in the system. The stochastic dynamics is then represented by the Fokker-Planck (FP) partial differential equation (PDE) that governs the evolution of the associated probability density function. The optimal therapies are obtained using a three step procedure. First, a finite dimensional FP-constrained optimization problem is formulated that takes input individual noisy patient data, and is solved to obtain the unknown parameters corresponding to the individual tumor characteristics. Next, a sensitivity analysis of the optimal parameter set is used to determine the parameters to be controlled, thus, helping in assessing the types of treatment therapies. Finally, a feedback FP control problem is solved to determine the optimal combination therapies. Numerical results with the combination drug, comprising of Bevacizumab and Capecitabine, demonstrate the efficiency of the proposed framework.

q-bio.QM

A Simplified Stochastic EM Algorithm for Cure Rate Model with Negative Binomial Competing Risks: An Application to Breast Cancer Data

In this paper, a long-term survival model under competing risks is considered. The unobserved number of competing risks is assumed to follow a negative binomial distribution that can capture both over- and under-dispersion. Considering the latent competing risks as missing data, a variation of the well-known expectation maximization (EM) algorithm, called the stochastic EM algorithm (SEM), is developed. It is shown that the SEM algorithm avoids calculation of complicated expectations, which is a major advantage of the SEM algorithm over the EM algorithm. The proposed procedure also allows the objective function to be split into two simpler functions, one corresponding to the parameters associated with the cure rate and the other corresponding to the parameters associated with the progression times. The advantage of this approach is that each simple function, with lower parameter dimension, can be maximized independently. An extensive Monte Carlo simulation study is carried out to compare the performances of the SEM and EM algorithms. Finally, a breast cancer survival data is analyzed and it is shown that the SEM algorithm performs better than the EM algorithm.

stat.ME

A Stochastic Version of the EM Algorithm for Mixture Cure Rate Model with Exponentiated Weibull Family of Lifetimes

Handling missing values plays an important role in the analysis of survival data, especially, the ones marked by cure fraction. In this paper, we discuss the properties and implementation of stochastic approximations to the expectation-maximization (EM) algorithm to obtain maximum likelihood (ML) type estimates in situations where missing data arise naturally due to right censoring and a proportion of individuals are immune to the event of interest. A flexible family of three parameter exponentiated-Weibull (EW) distributions is assumed to characterize lifetimes of the non-immune individuals as it accommodates both monotone (increasing and decreasing) and non-monotone (unimodal and bathtub) hazard functions. To evaluate the performance of the SEM algorithm, an extensive simulation study is carried out under various parameter settings. Using likelihood ratio test we also carry out model discrimination within the EW family of distributions. Furthermore, we study the robustness of the SEM algorithm with respect to outliers and algorithm starting values. Few scenarios where stochastic EM (SEM) algorithm outperforms the well-studied EM algorithm are also examined in the given context. For further demonstration, a real survival data on cutaneous melanoma is analyzed using the proposed cure rate model with EW lifetime distribution and the proposed estimation technique. Through this data, we illustrate the applicability of the likelihood ratio test towards rejecting several well-known lifetime distributions that are nested within the wider class of EW distributions.

stat.ME

A New Non-Linear Conjugate Gradient Algorithm for Destructive Cure Rate Model and a Simulation Study: Illustration with Negative Binomial Competing Risks

In this paper, we propose a new estimation methodology based on a projected non-linear conjugate gradient (PNCG) algorithm with an efficient line search technique. We develop a general PNCG algorithm for a survival model incorporating a proportion cure under a competing risks setup, where the initial number of competing risks are exposed to elimination after an initial treatment (known as destruction). In the literature, expectation maximization (EM) algorithm has been widely used for such a model to estimate the model parameters. Through an extensive Monte Carlo simulation study, we compare the performance of our proposed PNCG with that of the EM algorithm and show the advantages of our proposed method. Through simulation, we also show the advantages of our proposed methodology over other optimization algorithms (including other conjugate gradient type methods) readily available as R software packages. To show these we assume the initial number of competing risks to follow a negative binomial distribution although our general algorithm allows one to work with any competing risks distribution. Finally, we apply our proposed algorithm to analyze a well-known melanoma data.

math.ST

A New Estimation Algorithm for Box-Cox Transformation Cure Rate Model and Comparison With EM Algorithm

In this paper, we develop a new estimation procedure based on the non-linear conjugate gradient (NCG) algorithm for the Box-Cox transformation cure rate model. We compare the performance of the NCG algorithm with the well-known expectation maximization (EM) algorithm through a simulation study and show the advantages of the NCG algorithm over the EM algorithm. In particular, we show that the NCG algorithm allows simultaneous maximization of all model parameters when the likelihood surface is flat with respect to a Box-Cox model parameter. This is a big advantage over the EM algorithm, where a profile likelihood approach has been proposed in the literature that may not provide satisfactory results. We finally use the NCG algorithm to analyze a well-known melanoma data and show that it results in a better fit.

stat.CO