SearcharxivSearch

arXiv subjects

Mark Baillie

Publications and source records attributed to Mark Baillie.

7 recordsLinked to original sources

WATCH: A Workflow to Assess Treatment Effect Heterogeneity in Drug Development for Clinical Trial Sponsors

This paper proposes a Workflow for Assessing Treatment effeCt Heterogeneity (WATCH) in clinical drug development targeted at clinical trial sponsors. WATCH is designed to address the challenges of investigating treatment effect heterogeneity (TEH) in randomized clinical trials, where sample size and multiplicity limit the reliability of findings. The proposed workflow includes four steps: Analysis Planning, Initial Data Analysis and Analysis Dataset Creation, TEH Exploration, and Multidisciplinary Assessment. The workflow offers a general overview of how treatment effects vary by baseline covariates in the observed data, and guides interpretation of the observed findings based on external evidence and best scientific understanding. The workflow is exploratory and not inferential/confirmatory in nature, but should be pre-planned before data-base lock and analysis start. It is focused on providing a general overview rather than a single specific finding or subgroup with differential effect.

stat.AP

Efficiency of nonparametric superiority tests based on restricted mean survival time versus the log-rank test under proportional hazards

Background: For RCTs with time-to-event endpoints, proportional hazard (PH) models are typically used to estimate treatment effects and logrank tests are commonly used for hypothesis testing. There is growing support for replacing this approach with a model-free estimand and assumption-lean analysis method. One alternative is to base the analysis on the difference in restricted mean survival time (RMST) at a specific time, a single-number summary measure that can be defined without any restrictive assumptions on the outcome model. In a simple setting without covariates, an assumption-lean analysis can be achieved using nonparametric methods such as Kaplan Meier estimation. The main advantage of moving to a model-free summary measure and assumption-lean analysis is that the validity and interpretation of conclusions do not depend on the PH assumption. The potential disadvantage is that the nonparametric analysis may lose efficiency under PH. There is disagreement in recent literature on this issue. Methods: Asymptotic results and simulations are used to compare the efficiency of a log-rank test against a nonparametric analysis of the difference in RMST in a superiority trial under PH. Previous studies have separately examined the effect of event rates and the censoring distribution on relative efficiency. This investigation clarifies conflicting results from earlier research by exploring the joint effect of event rate and censoring distribution together. Several illustrative examples are provided. Results: In scenarios with high event rates and/or substantial censoring across a large proportion of the study window, and when both methods make use of the same amount of data, relative efficiency is close to unity. However, in cases with low event rates but when censoring is concentrated at the end of the study window, the PH analysis has a considerable efficiency advantage.

stat.ME

Predicting subgroup treatment effects for a new study: Motivations, results and learnings from running a data challenge in a pharmaceutical corporation

We present the motivation, experience and learnings from a data challenge conducted at a large pharmaceutical corporation on the topic of subgroup identification. The data challenge aimed at exploring approaches to subgroup identification for future clinical trials. To mimic a realistic setting, participants had access to 4 Phase III clinical trials to derive a subgroup and predict its treatment effect on a future study not accessible to challenge participants. 30 teams registered for the challenge with around 100 participants, primarily from Biostatistics organisation. We outline the motivation for running the challenge, the challenge rules and logistics. Finally, we present the results of the challenge, the participant feedback as well as the learnings, and how these learnings can be translated into statistical practice.

stat.AP

Why we should respect analysis results as data

The development and approval of new treatments generates large volumes of results, such as summaries of efficacy and safety. However, it is commonly overlooked that analyzing clinical study data also produces data in the form of results. For example, descriptive statistics and model predictions are data. Although integrating and putting findings into context is a cornerstone of scientific work, analysis results are often neglected as a data source. Results end up stored as "data products" such as PDF documents that are not machine readable or amenable to future analysis. We propose a solution to "calculate once, use many times" by combining analysis results standards with a common data model. This analysis results data model re-frames the target of analyses from static representations of the results (e.g., tables and figures) to a data model with applications in various contexts, including knowledge discovery. Further, we provide a working proof of concept detailing how to approach analyses standardization and construct a schema to store and query analysis results.

cs.CY

A Deep Learning Approach to Private Data Sharing of Medical Images Using Conditional GANs

Sharing data from clinical studies can facilitate innovative data-driven research and ultimately lead to better public health. However, sharing biomedical data can put sensitive personal information at risk. This is usually solved by anonymization, which is a slow and expensive process. An alternative to anonymization is sharing a synthetic dataset that bears a behaviour similar to the real data but preserves privacy. As part of the collaboration between Novartis and the Oxford Big Data Institute, we generate a synthetic dataset based on COSENTYX (secukinumab) Ankylosing Spondylitis clinical study. We apply an Auxiliary Classifier GAN to generate synthetic MRIs of vertebral units. The images are conditioned on the VU location (cervical, thoracic and lumbar). In this paper, we present a method for generating a synthetic dataset and conduct an in-depth analysis on its properties along three key metrics: image fidelity, sample diversity and dataset privacy.

cs.LG

Tutorial: Effective visual communication for the quantitative scientist

Effective visual communication is a core competency for pharmacometricians, statisticians, and more generally any quantitative scientist. It is essential in every step of a quantitative workflow, from scoping to execution and communicating results and conclusions. With this competency, we can better understand data and influence decisions towards appropriate actions. Without it, we can fool ourselves and others and pave the way to wrong conclusions and actions. The goal of this tutorial is to convey this competency. We posit three laws of effective visual communication for the quantitative scientist: have a clear purpose, show the data clearly, and make the message obvious. A concise "Cheat Sheet", available on https://graphicsprinciples.github.io, distills more granular recommendations for everyday practical use. Finally, these laws and recommendations are illustrated in four case studies.

stat.AP

Folksonomic Tag Clouds as an Aid to Content Indexing

Social tagging systems have recently developed as a popular method of data organisation on the Internet. These systems allow users to organise their content in a way that makes sense to them, rather than forcing them to use a pre-determined and rigid set of categorisations. These folksonomies provide well populated sources of unstructured tags describing web resources which could potentially be used as semantic index terms for these resources. However getting people to agree on what tags best describe a resource is a difficult problem, therefore any feature which increases the consistency and stability of terms chosen would be extremely beneficial. We investigate how the provision of a tag cloud, a weighted list of terms commonly used to assist in browsing a folksonomy, during the tagging process itself influences the tags produced and how difficult the user perceived the task to be. We show that illustrating the most popular tags to users assists in the tagging process and encourages a stable and consistent folksonomy to form.

cs.IR