SearcharxivSearch

arXiv subjects

Sarah Friedrich

Publications and source records attributed to Sarah Friedrich.

17 recordsLinked to original sources

A Guide to Estimating Conditional Average Treatment Effects in Competing Risks Settings

Conditional average treatment effects (CATEs) are central to treatment decision-making in personalized medicine. In competing risks settings, estimating CATEs from survival data allows for patient-specific assessments of treatment effectiveness for a specific event of interest while properly accounting for alternative event types. This distinction is essential in the presence of comorbidities, where competing causes of death may otherwise confound the therapeutic benefit. Focusing on right-censored survival times with binary treatment, we examine CATEs defined as covariate-conditional differences in the absolute risk for the event of interest at a fixed time. To this end, we study meta-learners which adapt machine learning algorithms for CATE estimation in competing risks scenarios. We systematically compare six meta-learners, combining Cox regression or random survival forests for risk modeling with elastic net regression or random forests for direct CATE modeling. To provide practical guidance on model selection, we evaluate their performance in multiple simulation settings, that differ in hazard complexity, treatment heterogeneity, treatment assignment, event type distribution and censoring. To facilitate applied use, we provide the R package, crsurvlearners, which implements all considered approaches.

stat.AP

Missing Data Imputation in the Context of Propensity Score Analysis: A Systematic Review

Missing data is a common challenge in observational studies. Another challenge stems from the observational nature of the study itself. Here, propensity score analysis can be used as a technique to replicate conditions similar to those found in clinical trials. With regard to the missing data, a majority of studies only analyze the complete cases, but this has several pitfalls. In this review, we investigate which methods are used for the handling of missing data in the context of propensity score analyses. Therefore, we searched PubMed for the keywords propensity score and missing data, restricting our search to the time between January 2010 and February 2024. The PRISMA statement was followed in this review. A total of 147 articles were included in the analyses. A major finding of this study is that although the usage of multiple imputation (MI) has risen over time, only a limited number of studies describe the mechanism of missing data and the details of the MI algorithm. Keywords Missing data, Propensity Score, Observational Data, Multiple Imputation, Systematic Review

stat.OT

Descriptive Discriminant Analysis of Multivariate Repeated Measures Data: A Use Case

Psychological research often focuses on examining group differences in a set of numeric variables for which normality is doubtful. Longitudinal studies enable the investigation of developmental trends. For instance, a recent study (Voormolen et al (2020), https://doi.org/10.3390/jcm9051525) examined the relation of complicated and uncomplicated mild traumatic brain injury (mTBI) with multidimensional outcomes measured at three- and six-months after mTBI. The data were analyzed using robust repeated measures multivariate analysis of variance (MANOVA), resulting in significant differences between groups and across time points, then followed up by univariate ANOVAs per variable as is typically done. However, this approach ignores the multivariate aspect of the original analyses. We propose descriptive discriminant analysis (DDA) as an alternative, which is a robust multivariate technique recommended for examining significant MANOVA results and has not yet been applied to multivariate repeated measures data. We provide a tutorial with annotated R code demonstrating its application to these empirical data.

stat.AP

Linear classification methods for multivariate repeated measures data -- a simulation study

Researchers in the behavioral and social sciences use linear discriminant analysis (LDA) for predictions of group membership (classification) and for identifying the variables most relevant to group separation among a set of continuous correlated variables (description). \\ In these and other disciplines, longitudinal data are often collected which provide additional temporal information. Linear classification methods for repeated measures data are more sensitive to actual group differences by taking the complex correlations between time points and variables into account, but are rarely discussed in the literature. Moreover, psychometric data rarely fulfill the multivariate normality assumption.\\ In this paper, we compare existing linear classification algorithms for nonnormally distributed multivariate repeated measures data in a simulation study based on psychological questionnaire data comprising Likert scales. The results show that in data without any specific assumed structure and larger sample sizes, the robust alternatives to standard repeated measures LDA may not be needed. To our knowledge, this is one of the few studies discussing repeated measures classification techniques, and the first one comparing multiple alternatives among each other.

stat.ME

Resampling-based confidence intervals and bands for the average treatment effect in observational studies with competing risks

The g-formula can be used to estimate the treatment effect while accounting for confounding bias in observational studies. With regard to time-to-event endpoints, possibly subject to competing risks, the construction of valid pointwise confidence intervals and time-simultaneous confidence bands for the causal risk difference is complicated, however. A convenient solution is to approximate the asymptotic distribution of the corresponding stochastic process by means of resampling approaches. In this paper, we consider three different resampling methods, namely the classical nonparametric bootstrap, the influence function equipped with a resampling approach as well as a martingale-based bootstrap version. We set up a simulation study to compare the accuracy of the different techniques, which reveals that the wild bootstrap should in general be preferred if the sample size is moderate and sufficient data on the event of interest have been accrued. For illustration, the three resampling methods are applied to data on the long-term survival in patients with early-stage Hodgkin's disease.

stat.ME

Asymptotic properties of resampling-based processes for the average treatment effect in observational studies with competing risks

In observational studies with time-to-event outcomes, the g-formula can be used to estimate a treatment effect in the presence of confounding factors. However, the asymptotic distribution of the corresponding stochastic process is complicated and thus not suitable for deriving confidence intervals or time-simultaneous confidence bands for the average treatment effect. A common remedy are resampling-based approximations, with Efron's nonparametric bootstrap being the standard tool in practice. We investigate the large sample properties of three different resampling approaches and prove their asymptotic validity in a setting with time-to-event data subject to competing risks.

math.ST

Liquid-organic time projection chamber for detecting low energy antineutrinos

The MeV region of antineutrino energy is of special interest for physics research and for monitoring nuclear nonproliferation. Whereas liquid scintillation detectors are typically used to detect the Inverse Beta Decay (IBD), it has recently been proposed to detect it with a liquid-organic Time Projection Chamber, which could allow a full reconstruction of the particle tracks of the IBD final state. We present the first comprehensive simulation-based study of the expected signatures. Their unequivocal signature could enable a background-minimized detection of electron antineutrinos using information on energy, location and direction of all final state particles. We show that the positron track reflects the antineutrino's vertex. It can also be used to determine the initial neutrino energy. In addition, we investigate the possibility to reconstruct the antineutrino direction on an event-by-event basis by the energy deposition of the neutron-induced proton recoils. Our simulations indicate that this could be a promising approach which should be further studied through experiments with a detector prototype.

physics.ins-det

On the role of benchmarking data sets and simulations in method comparison studies

Method comparisons are essential to provide recommendations and guidance for applied researchers, who often have to choose from a plethora of available approaches. While many comparisons exist in the literature, these are often not neutral but favour a novel method. Apart from the choice of design and a proper reporting of the findings, there are different approaches concerning the underlying data for such method comparison studies. Most manuscripts on statistical methodology rely on simulation studies and provide a single real-world data set as an example to motivate and illustrate the methodology investigated. In the context of supervised learning, in contrast, methods are often evaluated using so-called benchmarking data sets, i.e. real-world data that serve as gold standard in the community. Simulation studies, on the other hand, are much less common in this context. The aim of this paper is to investigate differences and similarities between these approaches, to discuss their advantages and disadvantages and ultimately to develop new approaches to the evaluation of methods picking the best of both worlds. To this aim, we borrow ideas from different contexts such as mixed methods research and Clinical Scenario Evaluation.

stat.ME

On the role of data, statistics and decisions in a pandemic

A pandemic poses particular challenges to decision-making because of the need to continuously adapt decisions to rapidly changing evidence and available data. For example, which countermeasures are appropriate at a particular stage of the pandemic? How can the severity of the pandemic be measured? What is the effect of vaccination in the population and which groups should be vaccinated first? The process of decision-making starts with data collection and modeling and continues to the dissemination of results and the subsequent decisions taken. The goal of this paper is to give an overview of this process and to provide recommendations for the different steps from a statistical perspective. In particular, we discuss a range of modeling techniques including mathematical, statistical and decision-analytic models along with their applications in the COVID-19 context. With this overview, we aim to foster the understanding of the goals of these modeling approaches and the specific data requirements that are essential for the interpretation of results and for successful interdisciplinary collaborations. A special focus is on the role played by data in these different models, and we incorporate into the discussion the importance of statistical literacy, and of effective dissemination and communication of findings.

stat.OT

Causal inference methods for small non-randomized studies: Methods and an application in COVID-19

The usual development cycles are too slow for the development of vaccines, diagnostics and treatments in pandemics such as the ongoing SARS-CoV-2 pandemic. Given the pressure in such a situation, there is a risk that findings of early clinical trials are overinterpreted despite their limitations in terms of size and design. Motivated by a non-randomized open-label study investigating the efficacy of hydroxychloroquine in patients with COVID-19, we describe in a unified fashion various alternative approaches to the analysis of non-randomized studies. A widely used tool to reduce the impact of treatment-selection bias are so-called propensity score (PS) methods. Conditioning on the propensity score allows one to replicate the design of a randomized controlled trial, conditional on observed covariates. Extensions include the g-computation approach, which is less frequently applied, in particular in clinical studies. Moreover, doubly robust estimators provide additional advantages. Here, we investigate the properties of propensity score based methods including three variations of doubly robust estimators in small sample settings, typical for early trials, in a simulation study. R code for the simulations is provided.

stat.AP

Is there a role for statistics in artificial intelligence?

The research on and application of artificial intelligence (AI) has triggered a comprehensive scientific, economic, social and political discussion. Here we argue that statistics, as an interdisciplinary scientific field, plays a substantial role both for the theoretical and practical understanding of AI and for its future development. Statistics might even be considered a core element of AI. With its specialist knowledge of data evaluation, starting with the precise formulation of the research question and passing through a study design stage on to analysis and interpretation of the results, statistics is a natural partner for other disciplines in teaching, research and practice. This paper aims at contributing to the current discussion by highlighting the relevance of statistical methodology in the context of AI development. In particular, we discuss contributions of statistics to the field of artificial intelligence concerning methodological development, planning and design of studies, assessment of data quality and data collection, differentiation of causality and associations and assessment of uncertainty in results. Moreover, the paper also deals with the equally necessary and meaningful extension of curricula in schools and universities.

cs.CY

More powerful logrank permutation tests for two-sample survival data

Weighted logrank tests are a popular tool for analyzing right censored survival data from two independent samples. Each of these tests is optimal against a certain hazard alternative, for example the classical logrank test for proportional hazards. But which weight function should be used in practical applications? We address this question by a flexible combination idea leading to a testing procedure with broader power. Beside the test's asymptotic exactness and consistency its power behaviour under local alternatives is derived. All theoretical properties can be transferred to a permutation version of the test, which is even finitely exact under exchangeability and showed a better finite sample performance in our simulation study. The procedure is illustrated in a real data example.

math.ST

Nonparametric MANOVA in Mann-Whitney effects

Multivariate analysis of variance (MANOVA) is a powerful and versatile method to infer and quantify main and interaction effects in metric multivariate multi-factor data. It is, however, neither robust against change in units nor a meaningful tool for ordinal data. Thus, we propose a novel nonparametric MANOVA. Contrary to existing rank-based procedures we infer hypotheses formulated in terms of meaningful Mann-Whitney-type effects in lieu of distribution functions. The tests are based on a quadratic form in multivariate rank effect estimators and critical values are obtained by the bootstrap. This newly developed procedure consequently provides asymptotically exact and consistent inference for general models such as the nonparametric Behrens-Fisher problem as well as multivariate one-, two-, and higher-way crossed layouts. Computer simulations in small samples confirm the reliability of the developed method for ordinal as well as metric data with covariance heterogeneity. Finally, an analysis of a real data example illustrates the applicability and correct interpretation of the results.

math.ST

Analysis of Multivariate Data and Repeated Measures Designs with the R Package MANOVA.RM

The numerical availability of statistical inference methods for a modern and robust analysis of longitudinal- and multivariate data in factorial experiments is an essential element in research and education. While existing approaches that rely on specific distributional assumptions of the data (multivariate normality and/or characteristic covariance matrices) are implemented in statistical software packages, there is a need for user-friendly software that can be used for the analysis of data that do not fulfill the aforementioned assumptions and provide accurate p-value and confidence interval estimates. Therefore, newly developed statistical methods for the analysis of repeated measures designs and multivariate data that neither assume multivariate normality nor specific covariance matrices have been implemented in the freely available R-package MANOVA.RM. The package is equipped with a graphical user interface for plausible applications in academia and other educational purpose. Several motivating examples illustrate the application of the methods.

stat.CO

MATS: Inference for potentially Singular and Heteroscedastic MANOVA

In many experiments in the life sciences, several endpoints are recorded per subject. The analysis of such multivariate data is usually based on MANOVA models assuming multivariate normality and covariance homogeneity. These assumptions, however, are often not met in practice. Furthermore, test statistics should be invariant under scale transformations of the data, since the endpoints may be measured on different scales. In the context of high-dimensional data, Srivastava and Kubokawa (2013) proposed such a test statistic for a specific one-way model, which, however, relies on the assumption of a common non-singular covariance matrix. We modify and extend this test statistic to factorial MANOVA designs, incorporating general heteroscedastic models. In particular, our only distributional assumption is the existence of the group-wise covariance matrices, which may even be singular. We base inference on quantiles of resampling distributions, and derive confidence regions and ellipsoids based on these quantiles. In a simulation study, we extensively analyze the behavior of these procedures. Finally, the methods are applied to a data set containing information on the 2016 presidential elections in the USA with unequal and singular empirical covariance matrices.

stat.AP

Using EEG, SPECT, and Multivariate Resampling Methods to Differentiate Between Alzheimer's and other Cognitive Impairments

The incidence of Alzheimer's disease (AD) and other forms of dementia is increasing in most western countries. For a precise and early diagnosis, several examination modalities exist, among them single-photon emission computed tomography (SPECT) and the electroencephalogram (EEG). The latter is highly available, free of radiation hazards, and non-invasive. Thus, its diagnostic utility regarding different stages of dementia is of great interest in neurological research, along with the question of whether its utility depends on age or sex of the person being examined. However, SPECT or EEG measurements are intrinsically multivariate, and there has been a shortage of sufficiently general inferential techniques for the analysis of multivariate data in factorial designs when neither multivariate normality nor equality of covariance matrices across groups should be assumed. We adapt an asymptotic model based (parametric) bootstrap approach to this situation, demonstrate its ability, and use it for a truly multivariate analysis of the EEG and SPECT measurements, taking into account demographic factors such as age and sex. These multivariate results are supplemented by marginal effects bootstrap inference whose theoretical properties can be derived analogously to the multivariate methods. Both inference approaches can have advantages in particular situations, as illustrated in the data analysis.

stat.AP

Permuting longitudinal data despite all the dependencies

For general repeated measures designs the Wald-type statistic (WTS) is an asymptotically valid procedure allowing for unequal covariance matrices and possibly non-normal multivariate observations. The drawback of this procedure is the poor performance for small to moderate samples, i.e. decisions based on the WTS may become quite liberal. It is the aim of the present paper to improve its small sample behavior by means of a novel permutation procedure. In particular, it is shown that a permutation version of the WTS inherits its good large sample properties while yielding a very accurate finite sample control of the type-I error as shown in extensive simulations. Moreover, the new permutation method is motivated by a practical data set of a split plot design with a factorial structure on the repeated measures.

stat.ME