SearcharxivSearch

arXiv subjects

Tanayott Thaweethai

Publications and source records attributed to Tanayott Thaweethai.

4 recordsLinked to original sources

Inverse probability weighting for auxiliary variable dependent sampling in observational studies of Long COVID

Selective testing based on values of auxiliary variables is an increasingly popular design strategy in observational studies and is ubiquitous in electronic health record data. Ignoring this underlying sampling mechanism can lead to biased estimation and erroneous scientific conclusions. Yet, rigorous analytic methods for accounting for two-phase sampling designs in observational settings remain under-utilized. Motivated by the Researching COVID to Enhance Recovery (RECOVER) Adult and Pediatric observational cohort studies, we describe common pitfalls and an approach for analysis of data collected via auxiliary variable dependent sampling.

stat.ME

Statistical Considerations in Long COVID Research

Long COVID is a condition characterized by ongoing or relapsing symptoms attributable to SARS-CoV-2 infection that are present three or more months after infection. It represents a major clinical and public health concern as an estimated 5-10\% of individuals with a history of SARS-CoV-2 infection present with long term sequelae that range from mild to debilitating with profound impacts on quality of life. Clinical research studies of Long COVID have emerged rapidly over the past few years, and with them we are seeing several new data analytic challenges. In this manuscript, we highlight statistical challenges arising from the defining features of LC and associated study design strategies. This work is motivated by the Researching COVID to Enhance Recovery (RECOVER) Adult and Pediatric observational meta-cohort studies.

stat.AP

Sensitivity analysis for nonignorable missing values in blended analysis framework: a study on the effect of bariatric surgery via electronic health records

This paper establishes a series of sensitivity analyses to investigate the impact of missing values in the electronic health records (EHR) that are possibly missing not at random (MNAR). EHRs have gained tremendous interest due to their cost-effectiveness, but their employment for research involves numerous challenges, such as selection bias due to missing data. The blended analysis has been suggested to overcome such challenges, which decomposes the data provenance into a sequence of sub-mechanisms and uses a combination of inverse-probability weighting (IPW) and multiple imputation (MI) under missing at random assumption (MAR). In this paper, we expand the blended analysis under the MNAR assumption and present a sensitivity analysis framework to investigate the effect of MNAR missing values on the analysis results. We illustrate the performance of my proposed framework via numerical studies and conclude with strategies for interpreting the results of sensitivity analyses. In addition, we present an application of our framework to the DURABLE data set, an EHR from a study examining long-term outcomes of patients who underwent bariatric surgery.

stat.ME

Penalized regression with negative-unlabeled data: An approach to developing a long COVID research index

Moderate to severe post-acute sequelae of SARS-CoV-2 infection (PASC), also called long COVID, is estimated to impact as many as 10% of SARS-CoV-2 infected individuals, representing a chronic condition with a substantial global public health burden. An expansive literature has identified over 200 long-term and persistent symptoms associated with a history of SARS-CoV-2 infection; yet, there remains to be a clear consensus on a syndrome definition. Such a definition is a critical first step in future studies of risk and resiliency factors, mechanisms of disease, and interventions for both treatment and prevention. We recently applied a strategy for defining a PASC research index based on a Lasso-penalized logistic regression on history of SARS-CoV-2 infection. In the current paper we formalize and evaluate this approach through theoretical derivations and simulation studies. We demonstrate that this approach appropriately selects symptoms associated with PASC and results in a score that has high discriminatory power for detecting PASC. An application to data on participants enrolled in the RECOVER (Researching COVID to Enhance Recovery) Adult Cohort is presented to illustrate our findings.

stat.ME