SearcharxivSearch

arXiv subjects

F. J. Rubio

Publications and source records attributed to F. J. Rubio.

17 recordsLinked to original sources

An objective non-local prior for skew-symmetric models

We propose an objective non-local prior for testing symmetry against skew-symmetric alternatives. The prior is derived through a formal construction rule by assigning a uniform distribution to a discrepancy-based measure of the shape parameter's effect. This approach avoids the need for user-specified hyperparameters and produces a weakly informative prior tailored to the skew-symmetric family. We illustrate the use of the proposed prior in the context of testing normality against skew-normal alternatives through both a simulation study and a real-data application.

stat.ME

Hazard-based distributional regression via ordinary differential equations

The hazard function is central to the formulation of commonly used survival regression models such as the proportional hazards and accelerated failure time models. However, these models rely on a shared baseline hazard, which, when specified parametrically, can only capture limited shapes. To overcome this limitation, we propose a general class of parametric survival regression models obtained by modelling the hazard function using autonomous systems of ordinary differential equations (ODEs). Covariate information is incorporated via transformed linear predictors on the parameters of the ODE system. Our framework capitalises on the interpretability of parameters in common ODE systems, enabling the identification of covariate values that produce qualitatively distinct hazard shapes associated with different attractors of the system of ODEs. This provides deeper insights into how covariates influence survival dynamics. We develop efficient Bayesian computational tools, including parallelised evaluation of the log-posterior, which facilitates integration with general-purpose Markov Chain Monte Carlo samplers. We also derive conditions for posterior asymptotic normality, enabling fast approximations of the posterior. A central contribution of our work lies in the case studies. We demonstrate the methodology using clinical trial data with crossing survival curves, and a study of cancer recurrence times where our approach reveals how the efficacy of interventions (treatments) on hazard and survival are influenced by patient characteristics.

stat.ME

On harmonic oscillator hazard functions

We propose a parametric hazard model obtained by enforcing positivity in the damped harmonic oscillator. The resulting model has closed-form hazard and cumulative hazard functions, facilitating likelihood and Bayesian inference on the parameters. We show that this model can capture a range of hazard shapes, such as increasing, decreasing, unimodal, bathtub, and oscillatory patterns, and characterize the tails of the corresponding survival function. We illustrate the use of this model in survival analysis using real data.

stat.ME

Dynamic survival analysis: modelling the hazard function via ordinary differential equations

The hazard function represents one of the main quantities of interest in the analysis of survival data. We propose a general approach for parametrically modelling the dynamics of the hazard function using systems of autonomous ordinary differential equations (ODEs). This modelling approach can be used to provide qualitative and quantitative analyses of the evolution of the hazard function over time. Our proposal capitalises on the extensive literature of ODEs which, in particular, allow for establishing basic rules or laws on the dynamics of the hazard function via the use of autonomous ODEs. We show how to implement the proposed modelling framework in cases where there is an analytic solution to the system of ODEs or where an ODE solver is required to obtain a numerical solution. We focus on the use of a Bayesian modelling approach, but the proposed methodology can also be coupled with maximum likelihood estimation. A simulation study is presented to illustrate the performance of these models and the interplay of sample size and censoring. Two case studies using real data are presented to illustrate the use of the proposed approach and to highlight the interpretability of the corresponding models. We conclude with a discussion on potential extensions of our work and strategies to include covariates into our framework. Although we focus on examples on Medical Statistics, the proposed framework is applicable in any context where the interest lies on estimating and interpreting the dynamics hazard function.

stat.ME

On near-redundancy and identifiability of parametric hazard regression models under censoring

We study parametric inference on a rich class of hazard regression models in the presence of right-censoring. Previous literature has reported some inferential challenges, such as multimodal or flat likelihood surfaces, in this class of models for some particular data sets. We formalize the study of these inferential problems by linking them to the concepts of near-redundancy and practical non-identifiability of parameters. We show that the maximum likelihood estimators of the parameters in this class of models are consistent and asymptotically normal. Thus, the inferential problems in this class of models are related to the finite-sample scenario, where it is difficult to distinguish between the fitted model and a nested non-identifiable (i.e., parameter-redundant) model. We propose a method for detecting near-redundancy, based on distances between probability distributions. We also employ methods used in other areas for detecting practical non-identifiability and near-redundancy, including the inspection of the profile likelihood function and the Hessian method. For cases where inferential problems are detected, we discuss alternatives such as using model selection tools to identify simpler models that do not exhibit these inferential problems, increasing the sample size, or extending the follow-up time. We illustrate the performance of the proposed methods through a simulation study. Our simulation study reveals a link between the presence of near-redundancy and practical non-identifiability. Two illustrative applications using real data, with and without inferential problems, are presented.

stat.ME

Individual frailty excess hazard models in cancer epidemiology

Unobserved individual heterogeneity is a common challenge in population cancer survival studies. This heterogeneity is usually associated with the combination of model misspecification and the failure to record truly relevant variables. We investigate the effects of unobserved individual heterogeneity in the context of excess hazard models, one of the main tools in cancer epidemiology. We propose an individual excess hazard frailty model to account for individual heterogeneity. This represents an extension of frailty modelling to the relative survival framework. In order to facilitate the inference on the parameters of the proposed model, we select frailty distributions which produce closed-form expressions of the marginal hazard and survival functions. The resulting model allows for an intuitive interpretation, in which the frailties induce a selection of the healthier individuals among survivors. We model the excess hazard using a flexible parametric model with a general hazard structure which facilitates the inclusion of time-dependent effects. We illustrate the performance of the proposed methodology through a simulation study. We present a real-data example using data from lung cancer patients diagnosed in England, and discuss the impact of not accounting for unobserved heterogeneity on the estimation of net survival. The methodology is implemented in the R package IFNS.

stat.ME

On a prior based on the Wasserstein information matrix

We introduce a prior for the parameters of univariate continuous distributions, based on the Wasserstein information matrix, which is invariant under reparameterisations. We discuss the links between the proposed prior with information geometry. We present sufficient conditions for the propriety of the posterior distribution for general classes of models. We present a simulation study that shows that the induced posteriors have good frequentist properties.

math.ST

A Unifying Framework for Flexible Excess Hazard Modeling with Applications in Cancer Epidemiology

Excess hazard modeling is one of the main tools in population-based cancer survival research. Indeed, this setting allows for direct modeling of the survival due to cancer even in the absence of reliable information on the cause of death, which is common in population-based cancer epidemiology studies. We propose a unifying link-based additive modeling framework for the excess hazard that allows for the inclusion of many types of covariate effects, including spatial and time-dependent effects, using any type of smoother, such as thin plate, cubic splines, tensor products and Markov random fields. In addition, this framework accounts for all types of censoring as well as left-truncation. Estimation is conducted by using an efficient and stable penalized likelihood-based algorithm whose empirical performance is evaluated through extensive simulation studies. Some theoretical and asymptotic results are discussed. Two case studies are presented using population-based cancer data from patients diagnosed with breast (female), colon and lung cancers in England. The results support the presence of non-linear and time-dependent effects as well as spatial variation. The proposed approach is available in the R package GJRM.

stat.ME

Flexible linear mixed models with improper priors for longitudinal and survival data

We propose a Bayesian approach using improper priors for hierarchical linear mixed models with flexible random effects and residual error distributions. The error distribution is modelled using scale mixtures of normals, which can capture tails heavier than those of the normal distribution. This generalisation is useful to produce models that are robust to the presence of outliers. The case of asymmetric residual errors is also studied. We present general results for the propriety of the posterior that also cover cases with censored observations, allowing for the use of these models in the contexts of popular longitudinal and survival analyses. We consider the use of copulas with flexible marginals for modelling the dependence between the random effects, but our results cover the use of any random effects distribution. Thus, our paper provides a formal justification for Bayesian inference in a very wide class of models (covering virtually all of the literature) under attractive prior structures that limit the amount of required user elicitation.

stat.ME

Flexible objective Bayesian linear regression with applications in survival analysis

We study objective Bayesian inference for linear regression models with residual errors distributed according to the class of two-piece scale mixtures of normal distributions. These models allow for capturing departures from the usual assumption of normality of the errors in terms of heavy tails, asymmetry, and certain types of heteroscedasticity. We propose a general noninformative, scale-invariant, prior structure and provide sufficient conditions for the propriety of the posterior distribution of the model parameters, which cover cases when the response variables are censored. These results allow us to apply the proposed models in the context of survival analysis. This paper represents an extension to the Bayesian framework of the models proposed in Rubio and Hong (2015). We present a simulation study that shows good frequentist properties of the posterior credible intervals as well as point estimators associated to the proposed priors. We illustrate the performance of these models with real data in the context of survival analysis of cancer patients.

stat.AP

Bayesian modelling of skewness and kurtosis with two-piece scale and shape distributions

We formalise and generalise the definition of the family of univariate double two--piece distributions, obtained by using a density--based transformation of unimodal symmetric continuous distributions with a shape parameter. The resulting distributions contain five interpretable parameters that control the mode, as well as the scale and shape in each direction. Four-parameter subfamilies of this class of distributions that capture different types of asymmetry are discussed. We propose interpretable scale and location-invariant benchmark priors and derive conditions for the propriety of the corresponding posterior distribution. The prior structures used allow for meaningful comparisons through Bayes factors within flexible families of distributions. These distributions are applied to data from finance, internet traffic and medicine, comparing them with appropriate competitors.

stat.ME

On modelling asymmetric data using two-piece sinh-arcsinh distributions

We introduce the univariate two--piece sinh-arcsinh distribution, which contains two shape parameters that separately control skewness and kurtosis. We show that this new model can capture higher levels of asymmetry than the original sinh-arcsinh distribution (Jones and Pewsey, 2009), in terms of some asymmetry measures, while keeping flexibility of the tails and tractability. We illustrate the performance of the proposed model with real data, and compare it to appropriate alternatives. Although we focus on the study of the univariate versions of the proposed distributions, we point out some multivariate extensions.

stat.AP

On the Independence Jeffreys prior for skew--symmetric models with applications

We study the Jeffreys prior of the skewness parameter of a general class of scalar skew--symmetric models. It is shown that this prior is symmetric about 0, proper, and with tails $O(λ^{-3/2})$ under mild regularity conditions. We also calculate the independence Jeffreys prior for the case with unknown location and scale parameters. Sufficient conditions for the existence of the corresponding posterior distribution are investigated for the case when the sampling model belongs to the family of skew--symmetric scale mixtures of normal distributions. The usefulness of these results is illustrated using the skew--logistic model and two applications with real data.

stat.ME

Nonparametric inference for $P(X<Y)$ with paired variables

We propose two classes of nonparametric point estimators of $θ=P(X<Y)$ in the case where $(X,Y)$ are paired, possibly dependent, absolutely continuous random variables. The proposed estimators are based on nonparametric estimators of the joint density of $(X,Y)$ and the distribution function of $Z=Y - X$. We explore the use of several density and distribution function estimators and characterise the convergence of the resulting estimators of $θ$. We consider the use of bootstrap methods to obtain confidence intervals. The performance of these estimators is illustrated using simulated and real data. These examples show that not accounting for pairing and dependence may lead to erroneous conclusions about the relationship between $X$ and $Y$.

stat.ME

A Simple Approach to Maximum Intractable Likelihood Estimation

Approximate Bayesian Computation (ABC) can be viewed as an analytic approximation of an intractable likelihood coupled with an elementary simulation step. Such a view, combined with a suitable instrumental prior distribution permits maximum-likelihood (or maximum-a-posteriori) inference to be conducted, approximately, using essentially the same techniques. An elementary approach to this problem which simply obtains a nonparametric approximation of the likelihood surface which is then used as a smooth proxy for the likelihood in a subsequent maximisation step is developed here and the convergence of this class of algorithms is characterised theoretically. The use of non-sufficient summary statistics in this context is considered. Applying the proposed method to four problems demonstrates good performance. The proposed approach provides an alternative for approximating the maximum likelihood estimator (MLE) in complex scenarios.

stat.ME