SearcharxivSearch

arXiv subjects

Giampiero Marra

Publications and source records attributed to Giampiero Marra.

11 recordsLinked to original sources

A Bivariate Transformation Model for Time-to-Event Data Affected by Unobserved Confounding: Revisiting the Illinois Reemployment Bonus Experiment

Motivated by empirical studies investigating treatment effects in survival analysis, we propose a bivariate transformation model to quantify the impact of a binary treatment on a time-to-event outcome. The model equations are connected through a bivariate Gaussian distribution, with the dependence parameter capturing unobserved confounding, and are specified as functions of additive predictors to flexibly account for the impacts of observed confounders. Moreover, the baseline survival function is estimated using monotonic P-splines, the effects of binary or factor instruments can be regularized through a ridge penalty approach, and interactions between treatment and observed confounders can be incorporated to accommodate potential variations in treatment effects across subgroups. The proposal naturally provides the survival average treatment effect. Parameter estimation is achieved via an efficient and stable penalized maximum likelihood estimation approach, and intervals constructed using related inferential results. We revisit a dataset from the Illinois Reemployment Bonus Experiment to estimate the effect of a cash bonus on the probability of remaining unemployed at several time points, unveiling interesting insights. The modeling framework is incorporated into the R package GJRM, enabling researchers and practitioners to employ the proposed model and ensuring the reproducibility of results.

stat.ME

Spline-Based Multi-State Models for Analyzing Disease Progression

Motivated by disease progression-related studies, we propose an estimation method for fitting general non-homogeneous multi-state Markov models. The proposal can handle many types of multi-state processes, with several states and various combinations of observation schemes (e.g., intermittent, exactly observed, censored), and allows for the transition intensities to be flexibly modelled through additive (spline-based) predictors. The algorithm is based on a computationally efficient and stable penalized maximum likelihood estimation approach which exploits the information provided by the analytical Hessian matrix of the model log-likelihood. The proposed modeling framework is employed in case studies that aim at modeling the onset of cardiac allograft vasculopathy, and cognitive decline due to aging, where novel patterns are uncovered. To support applicability and reproducibility, all developed tools are implemented in the R package flexmsm.

stat.ME

Testing for similarity of multivariate mixed outcomes using generalised joint regression models with application to efficacy-toxicity responses

A common problem in clinical trials is to test whether the effect of an explanatory variable on a response of interest is similar between two groups, e.g. patient or treatment groups. In this regard, similarity is defined as equivalence up to a pre-specified threshold that denotes an acceptable deviation between the two groups. This issue is typically tackled by assessing if the explanatory variable's effect on the response is similar. This assessment is based on, for example, confidence intervals of differences or a suitable distance between two parametric regression models. Typically, these approaches build on the assumption of a univariate continuous or binary outcome variable. However, multivariate outcomes, especially beyond the case of bivariate binary response, remain underexplored. This paper introduces an approach based on a generalised joint regression framework exploiting the Gaussian copula. Compared to existing methods, our approach accommodates various outcome variable scales, such as continuous, binary, categorical, and ordinal, including mixed outcomes in multi-dimensional spaces. We demonstrate the validity of this approach through a simulation study and an efficacy-toxicity case study, hence highlighting its practical relevance.

stat.ME

Beyond unidimensional poverty analysis using distributional copula models for mixed ordered-continuous outcomes

Poverty is a multidimensional concept often comprising a monetary outcome and other welfare dimensions such as education, subjective well-being or health, that are measured on an ordinal scale. In applied research, multidimensional poverty is ubiquitously assessed by studying each poverty dimension independently in univariate regression models or by combining several poverty dimensions into a scalar index. This inhibits a thorough analysis of the potentially varying interdependence between the poverty dimensions. We propose a multivariate copula generalized additive model for location, scale and shape (copula GAMLSS or distributional copula model) to tackle this challenge. By relating the copula parameter to covariates, we specifically examine if certain factors determine the dependence between poverty dimensions. Furthermore, specifying the full conditional bivariate distribution, allows us to derive several features such as poverty risks and dependence measures coherently from one model for different individuals. We demonstrate the approach by studying two important poverty dimensions: income and education. Since the level of education is measured on an ordinal scale while income is continuous, we extend the bivariate copula GAMLSS to the case of mixed ordered-continuous outcomes. The new model is integrated into the GJRM package in R and applied to data from Indonesia. Particular emphasis is given to the spatial variation of the income-education dependence and groups of individuals at risk of being simultaneously poor in both education and income dimensions.

stat.ME

Modelling the Extremes of Seasonal Viruses and Hospital Congestion: The Example of Flu in a Swiss Hospital

Viruses causing flu or milder coronavirus colds are often referred to as "seasonal viruses" as they tend to subside in warmer months. In other words, meteorological conditions tend to impact the activity of viruses, and this information can be exploited for the operational management of hospitals. In this study, we use three years of daily data from one of the biggest hospitals in Switzerland and focus on modelling the extremes of hospital visits from patients showing flu-like symptoms and the number of positive cases of flu. We propose employing a discrete Generalized Pareto distribution for the number of positive and negative cases, and a Generalized Pareto distribution for the odds of positive cases. Our modelling framework allows for the parameters of these distributions to be linked to covariate effects, and for outlying observations to be dealt with via a robust estimation approach. Because meteorological conditions may vary over time, we use meteorological and not calendar variations to explain hospital charge extremes, and our empirical findings highlight their significance. We propose a measure of hospital congestion and a related tool to estimate the resulting CaRe (Charge-at-Risk-estimation) under different meteorological conditions. The relevant numerical computations can be easily carried out using the freely available GJRM R package. The introduced approach could be applied to several types of seasonal disease data such as those derived from the new virus SARS-CoV-2 and its COVID-19 disease which is at the moment wreaking havoc worldwide. The empirical effectiveness of the proposed method is assessed through a simulation study.

stat.ME

Robust Fitting for Generalized Additive Models for Location, Scale and Shape

The validity of estimation and smoothing parameter selection for the wide class of generalized additive models for location, scale and shape (GAMLSS) relies on the correct specification of a likelihood function. Deviations from such assumption are known to mislead any likelihood-based inference and can hinder penalization schemes meant to ensure some degree of smoothness for non-linear effects. We propose a general approach to achieve robustness in fitting GAMLSSs by limiting the contribution of observations with low log-likelihood values. Robust selection of the smoothing parameters can be carried out either by minimizing information criteria that naturally arise from the robustified likelihood or via an extended Fellner-Schall method. The latter allows for automatic smoothing parameter selection and is particularly advantageous in applications with multiple smoothing parameters. We also address the challenge of tuning robust estimators for models with non-linear effects by proposing a novel median downweighting proportion criterion. This enables a fair comparison with existing robust estimators for the special case of generalized additive models, where our estimator competes favorably. The overall good performance of our proposal is illustrated by further simulations in the GAMLSS setting and by an application to functional magnetic resonance brain imaging using bivariate smoothing splines.

stat.ME

Generalised Joint Regression for Count Data with a Focus on Modelling Football Matches

We propose a versatile joint regression framework for count responses. The method is implemented in the R add-on package GJRM and allows for modelling linear and non-linear dependence through the use of several copulae. Moreover, the parameters of the marginal distributions of the count responses and of the copula can be specified as flexible functions of covariates. Motivated by a football application, we also discuss an extension which forces the regression coefficients of the marginal (linear) predictors to be equal via a suitable penalisation. Model fitting is based on a trust region algorithm which estimates simultaneously all the parameters of the joint models. We investigate the proposal's empirical performance in two simulation studies, the first one designed for arbitrary count data, the other one reflecting football-specific settings. Finally, the method is applied to FIFA World Cup data, showing its competitiveness to the standard approach with regard to predictive performance.

stat.AP

Penalised maximum likelihood estimation in multistate models for interval-censored data

Multistate models can be used to describe transitions over time across states. In the presence of interval-censored times for transitions, the likelihood is constructed using transition probabilities. Models are specified using proportional hazards model for the transitions. Time-dependency is usually defined by parametric models, which can be too restrictive. Nonparametric hazards specification with splines allow for flexible modelling of time-dependency without making strong model assumptions. Penalised maximum likelihood is used to estimate the models. Selecting the optimal amount of smoothing is challenging as the problem involves multiple penalties. We propose an automatic and efficient method to estimate multistate models with splines in the presence of interval-censoring. The method is illustrated with a data analysis and a simulation study.

stat.ME

A Bivariate Copula Additive Model for Location, Scale and Shape

Rigby & Stasinopoulos (2005) introduced generalized additive models for location, scale and shape (GAMLSS) where the response distribution is not restricted to belong to the exponential family and its parameters can be specified as functions of additive predictors that allows for several types of covariate effects (e.g., linear, non-linear, random and spatial effects). In many empirical situations, however, modeling simultaneously two or more responses conditional on some covariates can be of considerable relevance. In this article, we extend the scope of GAMLSS by introducing a bivariate copula additive model with continuous margins for location, scale and shape. The framework permits the copula dependence and marginal distribution parameters to be estimated simultaneously and, like in GAMLSS, each parameter to be modeled using an additive predictor. Parameter estimation is achieved within a penalized likelihood framework using a trust region algorithm with integrated automatic multiple smoothing parameter selection. The proposed approach allows for straightforward inclusion of potentially any parametric continuous marginal distribution and copula function. The models can be easily used via the copulaReg() function in the R package SemiParBIVProbit. The usefulness of the proposal is illustrated on two case studies (which use electricity price and demand data, and birth records) and on simulated data.

stat.ME

Discrete Responses in Bivariate Generalized Additive Models

A conceptual framework for the analysis of dichotomous and ordinal polychotomous responses within a penalized multivariate Generalized Linear Model is introduced. The proposed structure allows for a rather flexible predictor specification through the inclusion of non-parametric and spatial covariate effects, and the characterisation of the distribution of the stochastic model components with copulae of univariate marginals. Analytic derivations for the particular case of Gaussian marginals within a bivariate system of dichotomous outcomes are also provided, and the framework is subsequently illustrated through the estimation of the HIV prevalence in Zambia using the 2007 DHS dataset.

stat.ME

Bankruptcy Prediction of Small and Medium Enterprises Using a Flexible Binary Generalized Extreme Value Model

We introduce a binary regression accounting-based model for bankruptcy prediction of small and medium enterprises (SMEs). The main advantage of the model lies in its predictive performance in identifying defaulted SMEs. Another advantage, which is especially relevant for banks, is that the relationship between the accounting characteristics of SMEs and response is not assumed a priori (e.g., linear, quadratic or cubic) and can be determined from the data. The proposed approach uses the quantile function of the generalized extreme value distribution as link function as well as smooth functions of accounting characteristics to flexibly model covariate effects. Therefore, the usual assumptions in scoring models of symmetric link function and linear or pre-specied covariate-response relationships are relaxed. Out-of-sample and out-of-time validation on Italian data shows that our proposal outperforms the commonly used (logistic) scoring model for different default horizons.

stat.ME