Searcharxiv⌕ Search

arXiv subjects

Francisco J. Rubio

Publications and source records attributed to Francisco J. Rubio.

6 recordsLinked to original sources

On models for the estimation of the excess mortality hazard in case of insufficiently stratified life tables

In cancer epidemiology using population-based data, regression models for the excess mortality hazard is a useful method to estimate cancer survival and to describe the association between prognosis factors and excess mortality. This method requires expected mortality rates from general population life tables: each cancer patient is assigned an expected (background) mortality rate obtained from the life tables, typically at least according to their age and sex, from the population they belong to. However, those life tables may be insufficiently stratified, as some characteristics such as deprivation, ethnicity, and comorbidities, are not available in the life tables for a number of countries. This may affect the background mortality rate allocated to each patient, and it has been shown that not including relevant information for assigning an expected mortality rate to each patient induces a bias in the estimation of the regression parameters of the excess hazard model. We propose two parametric corrections in excess hazard regression models, including a single-parameter or a random effect (frailty), to account for possible mismatches in the life table and thus misspecification of the background mortality rate. In an extensive simulation study, the good statistical performance of the proposed approach is demonstrated, and we illustrate their use on real population-based data of lung cancer patients. We present conditions and limitations of these methods, and provide some recommendations for their use in practice.

stat.ME↗

On a general structure for hazard-based regression models: an application to population-based cancer research

The proportional hazards model represents the most commonly assumed hazard structure when analysing time to event data using regression models. We study a general hazard structure which contains, as particular cases, proportional hazards, accelerated hazards, and accelerated failure time structures, as well as combinations of these. We propose an approach to apply these different hazard structures, based on a flexible parametric distribution (Exponentiated Weibull) for the baseline hazard. This distribution allows us to cover the basic hazard shapes of interest in practice: constant, bathtub, increasing, decreasing, and unimodal. In an extensive simulation study, we evaluate our approach in the context of excess hazard modelling, which is the main quantity of interest in descriptive cancer epidemiology. This study exhibits good inferential properties of the proposed model, as well as good performance when using the Akaike Information Criterion for selecting the hazard structure. An application on lung cancer data illustrates the usefulness of the proposed model.

stat.ME↗

Objective priors for the number of degrees of freedom of a multivariate t distribution and the t-copula

An objective Bayesian approach to estimate the number of degrees of freedom $(ν)$ for the multivariate $t$ distribution and for the $t$-copula, when the parameter is considered discrete, is proposed. Inference on this parameter has been problematic for the multivariate $t$ and, for the absence of any method, for the $t$-copula. An objective criterion based on loss functions which allows to overcome the issue of defining objective probabilities directly is employed. The support of the prior for $ν$ is truncated, which derives from the property of both the multivariate $t$ and the $t$-copula of convergence to normality for a sufficiently large number of degrees of freedom. The performance of the priors is tested on simulated scenarios. The R codes and the replication material are available as a supplementary material of the electronic version of the paper and on real data: daily logarithmic returns of IBM and of the Center for Research in Security Prices Database.

stat.ME↗

Tractable Bayesian variable selection: beyond normality

Bayesian variable selection often assumes normality, but the effects of model misspecification are not sufficiently understood. There are sound reasons behind this assumption, particularly for large $p$: ease of interpretation, analytical and computational convenience. More flexible frameworks exist, including semi- or non-parametric models, often at the cost of some tractability. We propose a simple extension of the Normal model that allows for skewness and thicker-than-normal tails but preserves tractability. It leads to easy interpretation and a log-concave likelihood that facilitates optimization and integration. We characterize asymptotically parameter estimation and Bayes factor rates, in particular studying the effects of model misspecification. Under suitable conditions misspecified Bayes factors are consistent and induce sparsity at the same asymptotic rates than under the correct model. However, the rates to detect signal are altered by an exponential factor, often resulting in a loss of sensitivity. These deficiencies can be ameliorated by inferring the error distribution from the data, a simple strategy that can improve inference substantially. Our work focuses on the likelihood and can thus be combined with any likelihood penalty or prior, but here we focus on non-local priors to induce extra sparsity and ameliorate finite-sample effects caused by misspecification. Our results highlight the practical importance of focusing on the likelihood rather than solely on the prior, when it comes to Bayesian variable selection. The methodology is available in R package `mombf'.

stat.ME↗

Bayesian linear regression with skew-symmetric error distributions with applications to survival analysis

We study Bayesian linear regression models with skew-symmetric scale mixtures of normal error distributions. These kinds of models can be used to capture departures from the usual assumption of normality of the errors in terms of heavy tails and asymmetry. We propose a general non-informative prior structure for these regression models and show that the corresponding posterior distribution is proper under mild conditions. We extend these propriety results to cases where the response variables are censored. The latter scenario is of interest in the context of accelerated failure time models, which are relevant in survival analysis. We present a simulation study that demonstrates good frequentist properties of the posterior credible intervals associated to the proposed priors. This study also sheds some light on the trade-off between increased model flexibility and the risk of over-fitting. We illustrate the performance of the proposed models with real data. Although we focus on models with univariate response variables, we also present some extensions to the multivariate case in the Supporting Web Material.

stat.AP↗

Survival and lifetime data analysis with a flexible class of distributions

We introduce a general class of continuous univariate distributions with positive support obtained by transforming the class of two-piece distributions. We show that this class of distributions is very flexible, easy to implement, and contains members that can capture different tail behaviours and shapes, producing also a variety of hazard functions. The proposed distributions represent a flexible alternative to the classical choices such as the log-normal, Gamma, and Weibull distributions. We investigate empirically the inferential properties of the proposed models through an extensive simulation study. We present some applications using real data in the contexts of time-to-event and accelerated failure time models. In the second kind of applications, we explore the use of these models in the estimation of the distribution of the individual remaining life.

stat.AP↗