SearcharxivSearch

arXiv subjects

Karen Marder

Publications and source records attributed to Karen Marder.

7 recordsLinked to original sources

PRESCCO: Efficient Prediction Intervals under a Right-Censored Covariate

In clinical studies, a patient's outcome, e.g., a cognitive test score, is typical or atypical depending on how it compares with the outcomes of patients at a similar point in a neurodegenerative disease, measured by how far they are from a common disease event. A prediction interval based on the time to that event provides that comparison, but that time is right-censored for most patients. Yet no method had computed such a prediction interval in this right-censored covariate setting. We first adapt three conformal prediction methods to this setting, and show that the estimated half-length, the distance the interval runs either side of its center, varies from one study to the next, so the same outcome can be judged typical in one study and atypical in another. We then develop the PRESCCO method, whose estimator of the half-length is semiparametrically efficient and doubly robust, staying consistent when one of its two models is misspecified, while the method loses no coverage. In simulations its standard deviation is four to thirty times smaller than under any conformal prediction method, and in a Huntington disease study with 77.2% right-censoring, an outcome typical in one study is no longer atypical in another.

stat.ME

SPYCE: A Doubly Robust Estimator for Trials Targeting Early Huntington Disease under Outcome-Dependent Censoring

Clinical trials for neurodegenerative diseases must identify sensitive endpoints -- outcomes that change rapidly enough to detect treatment effects. In Huntington disease, this requires measuring how outcomes change as participants approach Stage 1. Yet many participants exit studies before reaching this stage, making their time to Stage 1 right-censored. Estimating how outcomes change requires models for both time to Stage 1 and time to study exit. When participants with worse outcomes exit earlier, this outcome-dependent censoring causes existing estimators to produce contradictory results: for the same cognitive outcome, one estimator suggests improvement while another shows decline. Existing estimators either ignore outcome-dependent censoring or require one model to be correctly specified, with no protection when it is not. We introduce SPYCE, a doubly robust estimator (consistent when either model is correctly specified) that achieves the smallest possible variance and allows both models to be estimated nonparametrically without sacrificing efficiency. Applied to data from PREDICT-HD, an observational Huntington disease study, SPYCE resolves current contradictions, identifies caudate and putamen volume ratios as the most promising sensitive endpoints, and shows that as few as 241 participants per arm are needed to detect treatment effects, versus hundreds of thousands under estimators that cannot handle outcome-dependent censoring.

stat.ME

Robust Estimation under Outcome Dependent Right Censoring in Huntington Disease: Estimators for Low and High Censoring Rates

Across health applications, researchers model outcomes as a function of time to an event, but the event time is right-censored for participants who exit the study or otherwise do not experience the event during follow-up. When censoring depends on the outcome-as in neurodegenerative disease studies where dropout is potentially related to disease severity-standard regression estimators produce biased estimates. We develop three consistent estimators for this outcome-dependent censoring setting: two augmented inverse probability weighted (AIPW) estimators and one maximum likelihood estimator (MLE). We establish their asymptotic properties and derive their robust sandwich variance estimators that account for nuisance parameter estimation. A key contribution is demonstrating that the choice of estimator to use depends on the censoring rate-the MLE performs best under low censoring rates, while the AIPW estimators yield lower bias and a higher nominal coverage under high censoring rates. We apply our estimators to Huntington disease data to characterize health decline leading up to mild cognitive impairment onset. The AIPW estimator with robustness matrix provided clinically-backed estimates with improved precision over inverse probability weighting, while MLE exhibited bias. Our results provide practical guidance for estimator selection based on censoring rate.

stat.ME

Super doubly robust and efficient estimator for informative covariate censoring

Early intervention in neurodegenerative diseases requires identifying periods before diagnosis when decline is rapid enough to detect whether a therapy is slowing progression. Since rapid decline typically occurs close to diagnosis, identifying these periods requires knowing each patient's time of diagnosis. Yet many patients exit studies before diagnosis, making time of diagnosis right-censored by time of study exit -- creating a right-censored covariate problem when estimating decline. Existing estimators either assume noninformative covariate censoring, where time of study exit is independent of time of diagnosis, or allow informative covariate censoring, but require correctly specifying how these times are related. We developed SPIRE (Semi-Parametric Informative Right-censored covariate Estimator), a super doubly robust estimator that remains consistent without correctly specifying densities governing time of diagnosis or time of study exit. Typical double robustness requires at least one density to be correct; SPIRE requires neither. When both densities are correctly specified, SPIRE achieves semiparametric efficiency. We also developed a test for detecting informative covariate censoring. Simulations with 85% right-censoring demonstrated SPIRE's robustness, efficiency and reliable detection of informative covariate censoring. Applied to Huntington disease data, SPIRE handled informative covariate censoring appropriately and remained consistent regardless of density specification, providing a reliable tool for early intervention.

math.ST

Establishing the Parallels and Differences Between Right-Censored and Missing Covariates

While right-censored time-to-event outcomes have been studied for decades, handling time-to-event covariates, also known as right-censored covariates, is now of growing interest. So far, the literature has treated right-censored covariates as distinct from missing covariates, overlooking the potential applicability of estimators to both scenarios. We bridge this gap by establishing connections between right-censored and missing covariates under various assumptions about censoring and missingness, allowing us to identify parallels and differences to determine when estimators can be used in both contexts. These connections reveal adaptations to five estimators for right-censored covariates in the unexplored area of informative covariate right-censoring and to formulate a new estimator for this setting, where the event time depends on the censoring time. We establish the asymptotic properties of the six estimators, evaluate their robustness under incorrect distributional assumptions, and establish their comparative efficiency. We conducted a simulation study to confirm our theoretical results, and then applied all estimators to a Huntington disease observational study to analyze cognitive impairments as a function of time to clinical diagnosis.

stat.ME

Estimation of Genetic Risk Function with Covariates in the Presence of Missing Genotypes

In genetic epidemiological studies, family history data are collected on relatives of study participants and used to estimate the age-specific risk of disease for individuals who carry a causal mutation. However, a family member's genotype data may not be collected due to the high cost of in-person interview to obtain blood sample or death of a relative. Previously, efficient nonparametric genotype-specific risk estimation in censored mixture data has been proposed without considering covariates. With multiple predictive risk factors available, risk estimation requires a multivariate model to account for additional covariates that may affect disease risk simultaneously. Therefore, it is important to consider the role of covariates in the genotype-specific distribution estimation using family history data. We propose an estimation method that permits more precise risk prediction by controlling for individual characteristics and incorporating interaction effects with missing genotypes in relatives, and thus gene-gene interactions and gene-environment interactions can be handled within the framework of a single model. We examine performance of the proposed methods by simulations and apply them to estimate the age-specific cumulative risk of Parkinson's disease (PD) in carriers of LRRK2 G2019S mutation using first-degree relatives who are at genetic risk for PD. The utility of estimated carrier risk is demonstrated through designing a future clinical trial under various assumptions. Such sample size estimation is seen in the Huntington's disease literature using the length of abnormal expansion of a CAG repeat in the HTT gene, but is less common in the PD literature.

stat.ME

Combining isotonic regression and EM algorithm to predict genetic risk under monotonicity constraint

In certain genetic studies, clinicians and genetic counselors are interested in estimating the cumulative risk of a disease for individuals with and without a rare deleterious mutation. Estimating the cumulative risk is difficult, however, when the estimates are based on family history data. Often, the genetic mutation status in many family members is unknown; instead, only estimated probabilities of a patient having a certain mutation status are available. Also, ages of disease-onset are subject to right censoring. Existing methods to estimate the cumulative risk using such family-based data only provide estimation at individual time points, and are not guaranteed to be monotonic or nonnegative. In this paper, we develop a novel method that combines Expectation-Maximization and isotonic regression to estimate the cumulative risk across the entire support. Our estimator is monotonic, satisfies self-consistent estimating equations and has high power in detecting differences between the cumulative risks of different populations. Application of our estimator to a Parkinson's disease (PD) study provides the age-at-onset distribution of PD in PARK2 mutation carriers and noncarriers, and reveals a significant difference between the distribution in compound heterozygous carriers compared to noncarriers, but not between heterozygous carriers and noncarriers.

stat.AP