Searcharxiv⌕ Search

arXiv subjects

Frank E. Harrell Jr

Publications and source records attributed to Frank E. Harrell Jr.

3 recordsLinked to original sources

Methods for adjusting for covariate measurement error in flexible modelling of functional form: results of a blinded, controlled neutral comparison simulation study

Covariate measurement error is pervasive in epidemiological research and distorts estimated exposure-outcome associations, yet correction methods have been studied almost exclusively under linear modelling assumptions. Their behaviour when the underlying association is non-linear and is itself estimated with flexible regression, remains poorly characterised. We report a blinded, multi-stage neutral comparison simulation study, conducted within the STRATOS initiative, evaluating measurement error correction coupled with flexible modelling of functional form. Six families of correction methods (pointwise and coefficient-based Simulation Extrapolation [SIMEX], Bayesian inference on the logit and risk scales, Multiple Imputation [MI], and Regression Calibration [RC]) were each combined with B-splines (BS), penalised splines (PS), fractional polynomials (FP), and natural splines (NS), yielding 23 analytic methods. Methods were applied to case-control data generated under five functional forms (J-shape, linear, two threshold models, and saturation) across simulated datasets spanning varying sample sizes, replication substudy sizes, error magnitudes, and error distributions, with classical additive error and a replication substudy for error calibration. Performance was assessed by the log mean squared error of the estimated function over the central 95 % of the exposure distribution. Pointwise SIMEX was the most accurate and most robust approach overall, followed by Bayesian methods and RC when paired with PS, FP, or NS; MI performed less well, and Bayesian estimation with unpenalised BS performed worst. PS, FP, and NS were near-equivalent, whereas BS was consistently inferior. No single method dominated across all scenarios, underscoring the value of sensitivity analyses.

stat.ME↗

Navigating the Landscape of Hierarchical Multi-Component Strategies: GPC, DOOR, and MOST

There is a growing recognition of the importance to involve patients in every stage of drug development. This shift acknowledges that patients' perspectives, experiences, and preferences are essential for ensuring that treatments meet real-world needs. In this context, a new body of statistical literature has emerged, focusing not only on the simultaneous consideration of multiple outcomes that reflect patients' overall experiences, but also on their structured prioritization. We refer to this class of approaches as hierarchical multi-component statistical methods. Among these, two influential frameworks - generalized pairwise comparisons (GPC) and desirability of outcome ranking (DOOR) - have emerged in the last decade, each aiming to offer a comprehensive approach to evaluating treatment effects. A new methodology, referred to here as the Markov ordinal state transition model (MOST), has recently been introduced without focusing on an explicit link with GPC nor DOOR. This paper seeks to fill this gap by offering a comprehensive and comparative analysis of the three approaches. Through examples and an exploration of the structural and philosophical differences between the methods, our aim is to provide guidance and encourage lines of research in the rapidly-evolving landscape of hierarchical multi-component statistical methodologies.

stat.ME↗

State-of-the-art in selection of variables and functional forms in multivariable analysis -- outstanding issues

How to select variables and identify functional forms for continuous variables is a key concern when creating a multivariable model. Ad hoc 'traditional' approaches to variable selection have been in use for at least 50 years. Similarly, methods for determining functional forms for continuous variables were first suggested many years ago. More recently, many alternative approaches to address these two challenges have been proposed, but knowledge of their properties and meaningful comparisons between them are scarce. To define a state-of-the-art and to provide evidence-supported guidance to researchers who have only a basic level of statistical knowledge many outstanding issues in multivariable modelling remain. Our main aims are to identify and illustrate such gaps in the literature and present them at a moderate technical level to the wide community of practitioners, researchers and students of statistics. We briefly discuss general issues in building descriptive regression models, strategies for variable selection, different ways of choosing functional forms for continuous variables, and methods for combining the selection of variables and functions. We discuss two examples, taken from the medical literature, to illustrate problems in the practice of modelling. Our overview revealed that there is not yet enough evidence on which to base recommendations for the selection of variables and functional forms in multivariable analysis. Such evidence may come from comparisons between alternative methods. In particular, we highlight seven important topics that require further investigation and make suggestions for the direction of further research.

stat.ME↗