SearcharxivSearch

arXiv subjects

Irantzu Barrio

Publications and source records attributed to Irantzu Barrio.

6 recordsLinked to original sources

Variable Domain Multivariate Functional Principal Component Analysis

Multivariate functional principal component analysis (MFPCA) is a powerful dimension reduction technique for analyzing multiple functional variables simultaneously. However, existing MFPCA methods assume that all functional observations are recorded over a common, fixed domain. This assumption is often violated in practical applications where the observation period varies across subjects, leading to what is known as variable domain functional data. We propose a novel approach for MFPCA that explicitly accommodates variable domains by extending existing multivariate functional principal component analysis to the variable domain setting. Our methodology involves performing univariate variable domain FPCA for each functional variable separately, stacking the resulting univariate scores, and then smoothing the empirical covariance matrix of these stacked scores over the domain length. This allows us to estimate multivariate eigenfunctions and scores that properly account for varying observation periods. We demonstrate through extensive simulation studies that our proposed method outperforms approaches that ignore the variable domain structure and rely on binning strategies. The practical utility of our method is illustrated through an application analyzing body temperature and capillary oxygen saturation (SpO$_2$) trajectories in COVID-19 hospital admitted patients, where patients experienced varying lengths of stay and monitoring periods.

stat.ME

Design-Based Inference for the AUC with Complex Survey Data

Complex survey data are usually collected following complex sampling designs. Accounting for the sampling design is essential to obtain unbiased estimates and valid inferences when analyzing complex survey data. The area under the receiver operating characteristic curve (AUC) is routinely used to assess the discriminative ability of predictive models for binary outcomes. However, valid inference for the AUC under complex sampling designs remains challenging. Although bootstrap techniques are widely applied under simple random sampling for variance estimation in this framework, traditional implementations do not account for complex designs. In this work, we propose a design-based framework for AUC inference. In particular, replicate weights methods are used to construct confidence intervals and hypothesis tests. The performance of replicate weights methods and the traditional non-design-based bootstrap for this purpose has been analyzed through an extensive simulation study. Design-based methods achieve coverage probabilities close to nominal levels and appropriate rejection rates under the null hypothesis. In contrast, the traditional non-design-based bootstrap method tends to underestimate the variance, leading to undercoverage and inflated rejection rates. Differences between methods decrease as the number of selected clusters per stratum increases. An application to data from the National Health and Nutrition Examination Survey (NHANES) illustrates the practical relevance of the proposed framework. The methods have been incorporated into the svyROC R package.

stat.ME

Dynamic prediction of death risk given a renewal hospitalization process

Predicting the risk of death for chronic patients is highly valuable for informed medical decision-making. This paper proposes a general framework for dynamic prediction of the risk of death of a patient given her hospitalization history, which is generally available to physicians. Predictions are based on a joint model for the death and hospitalization processes, thereby avoiding the potential bias arising from selection of survivors. The framework accommodates various submodels for the hospitalization process. In particular, we study prediction of the risk of death in a renewal model for hospitalizations, a common approach to recurrent event modelling. In the renewal model, the distribution of hospitalizations throughout the follow-up period impacts the risk of death. This result differs from prediction in the Poisson model, previously studied, where only the number of hospitalizations matters. We apply our methodology to a prospective, observational cohort study of 512 patients treated for COPD in one of six outpatient respiratory clinics run by the Respiratory Service of Galdakao University Hospital, with a median follow-up of 4.7 years. We find that more concentrated hospitalizations increase the risk of death.

stat.ME

Proposal of a general framework to categorize continuous predictor variables

The use of discretized variables in the development of prediction models is a common practice, in part because the decision-making process is more natural when it is based on rules created from segmented models. Although this practice is perhaps more common in medicine, it is extensible to any area of knowledge where a predictive model helps in decision-making. Therefore, providing researchers with a useful and valid categorization method could be a relevant issue when developing prediction models. In this paper, we propose a new general methodology that can be applied to categorize a predictor variable in any regression model where the response variable belongs to the exponential family distribution. Furthermore, it can be applied in any multivariate context, allowing to categorize more than one continuous covariate simultaneously. In addition, a computationally very efficient method is proposed to obtain the optimal number of categories, based on a pseudo-BIC proposal. Several simulation studies have been conducted in which the efficiency of the method with respect to both the location and the number of estimated cut-off points is shown. Finally, the categorization proposal has been applied to a real data set of 543 patients with chronic obstructive pulmonary disease from Galdakao Hospital's five outpatient respiratory clinics, who were followed up for 10 years. We applied the proposed methodology to jointly categorize the continuous variables six-minute walking test and forced expiratory volume in one second in a multiple Poisson generalized additive model for the response variable rate of the number of hospital admissions by years of follow-up. The location and number of cut-off points obtained were clinically validated as being in line with the categorizations used in the literature.

stat.ME

A study on group fairness in healthcare outcomes for nursing home residents during the COVID-19 pandemic in the Basque Country

We explore the effect of nursing home status on healthcare outcomes such as hospitalisation, mortality and in-hospital mortality during the COVID-19 pandemic. Some claim that in specific Autonomous Communities (geopolitical divisions) in Spain, elderly people in nursing homes had restrictions on access to hospitals and treatments, which raised a public outcry about the fairness of such measures. In this work, the case of the Basque Country is studied under a rigorous statistical approach and a physician's perspective. As fairness/unfairness is hard to model mathematically and has strong real-world implications, this work concentrates on the following simplification: establishing if the nursing home status had a direct effect on healthcare outcomes once accounted for other meaningful patients' information such as age, health status and period of the pandemic, among others. The methods followed here are a combination of established techniques as well as new proposals from the fields of causality and fair learning. The current analysis suggests that as a group, people in nursing homes were significantly less likely to be hospitalised, and considerably more likely to die, even in hospitals, compared to their non-residents counterparts during most of the pandemic. Further data collection and analysis are needed to guarantee that this is solely/mainly due to nursing home status.

stat.ME

Estimation of logistic regression parameters for complex survey data: a real data based simulation study

In complex survey data, each sampled observation has assigned a sampling weight, indicating the number of units that it represents in the population. Whether sampling weights should or not be considered in the estimation process of model parameters is a question that still continues to generate much discussion among researchers in different fields. We aim to contribute to this debate by means of a real data based simulation study in the framework of logistic regression models. In order to study their performance, three methods have been considered for estimating the coefficients of the logistic regression model: a) the unweighted model, b) the weighted model, and c) the unweighted mixed model. The results suggest the use of the weighted logistic regression model, showing the importance of using sampling weights in the estimation of the model parameters.

stat.ME