SearcharxivSearch

arXiv subjects

Tyler Wagner

Publications and source records attributed to Tyler Wagner.

5 recordsLinked to original sources

Accounting for variable detection functions in temporal abundance modeling via transfer learning

Relative abundance, measured as the number of animals caught per unit of sampling effort (CPUE), is commonly used to monitor fish and wildlife populations, largely because sampling methods are cost-effective to implement. Modeling relative abundance, however, requires the assumption that the detection probability is constant across sampling events. This assumption is likely not valid, as the probability of detection often varies as a function of several factors, including the characteristics of individual animals and environmental conditions at the time of sampling. In contrast, methods to estimate absolute abundance, such as capture-recapture (CR), account for variable detection, but are often infeasible to implement across large spatiotemporal scales. Despite this, CR data are sometimes available for species of interest, albeit at smaller spatiotemporal extents. Leveraging information on detection probabilities from CR data to help inform estimates of widely available CPUE data could strengthen inferences about the status of fish and wildlife populations. We propose an approach to (i) learn the effect of environmental covariates on detection probabilities from CR data and (ii) transfer these detection functions to CPUE models for improved inference. Shown empirically through a simulation study, this approach improves estimates of abundance and the ability to detect temporal trends. We apply our transfer learning method using CR and CPUE data to recreationally important smallmouth bass (\textit{Micropterus dolomieu}) fisheries in Pennsylvania, USA rivers.

stat.AP

Longitudinal laboratory testing tied to PCR diagnostics in COVID-19 patients reveals temporal evolution of distinctive coagulopathy signatures

Temporal inference from laboratory testing results and their triangulation with clinical outcomes as described in the associated unstructured text from the providers notes in the Electronic Health Record (EHR) is integral to advancing precision medicine. Here, we studied 181 COVIDpos and 7,775 COVIDneg patients subjected to 1.3 million laboratory tests across 194 assays during a two-month observation period centered around their SARS-CoV-2 PCR testing dates. We found that compared to COVIDneg at the time of clinical presentation and diagnostic testing, COVIDpos patients tended to have higher plasma fibrinogen levels and similarly low platelet counts, with approximately 25% of patients in both cohorts showing outright thrombocytopenia. However, these measures show opposite longitudinal trends as the infection evolves, with declining fibrinogen and increasing platelet counts to levels that are lower and higher compared to the COVIDneg cohort, respectively. Our EHR augmented curation efforts suggest a minority of patients develop thromboembolic events after the PCR testing date, including rare cases with disseminated intravascular coagulopathy (DIC), with most patients lacking the platelet reductions typically observed in consumptive coagulopathies. These temporal trends present, for the first time, fine-grained resolution of COVID-19 associated coagulopathy (CAC), via a digital framework that synthesizes longitudinal lab measurements with structured medication data and neural network-powered extraction of outcomes from the unstructured EHR. This study demonstrates how a precision medicine platform can help contextualize each patients specific coagulation profile over time, towards the goal of informing better personalization of thromboprophylaxis regimen.

physics.med-ph

Augmented Curation of Unstructured Clinical Notes from a Massive EHR System Reveals Specific Phenotypic Signature of Impending COVID-19 Diagnosis

Understanding the temporal dynamics of COVID-19 patient phenotypes is necessary to derive fine-grained resolution of pathophysiology. Here we use state-of-the-art deep neural networks over an institution-wide machine intelligence platform for the augmented curation of 15.8 million clinical notes from 30,494 patients subjected to COVID-19 PCR diagnostic testing. By contrasting the Electronic Health Record (EHR)-derived clinical phenotypes of COVID-19-positive (COVIDpos, n=635) versus COVID-19-negative (COVIDneg, n=29,859) patients over each day of the week preceding the PCR testing date, we identify anosmia/dysgeusia (37.4-fold), myalgia/arthralgia (2.6-fold), diarrhea (2.2-fold), fever/chills (2.1-fold), respiratory difficulty (1.9-fold), and cough (1.8-fold) as significantly amplified in COVIDpos over COVIDneg patients. The specific combination of cough and diarrhea has a 3.2-fold amplification in COVIDpos patients during the week prior to PCR testing, and along with anosmia/dysgeusia, constitutes the earliest EHR-derived signature of COVID-19 (4-7 days prior to typical PCR testing date). This study introduces an Augmented Intelligence platform for the real-time synthesis of institutional knowledge captured in EHRs. The platform holds tremendous potential for scaling up curation throughput, with minimal need for retraining underlying neural networks, thus promising EHR-powered early diagnosis for a broad spectrum of diseases.

cs.LG

Knowledge synthesis from 100 million biomedical documents augments the deep expression profiling of coronavirus receptors

The COVID-19 pandemic demands assimilation of all available biomedical knowledge to decode its mechanisms of pathogenicity and transmission. Despite the recent renaissance in unsupervised neural networks for decoding unstructured natural languages, a platform for the real-time synthesis of the exponentially growing biomedical literature and its comprehensive triangulation with deep omic insights is not available. Here, we present the nferX platform for dynamic inference from over 45 quadrillion possible conceptual associations extracted from unstructured biomedical text, and their triangulation with Single Cell RNA-sequencing based insights from over 25 tissues. Using this platform, we identify intersections between the pathologic manifestations of COVID-19 and the comprehensive expression profile of the SARS-CoV-2 receptor ACE2. We find that tongue keratinocytes and olfactory epithelial cells are likely under-appreciated targets of SARS-CoV-2 infection, correlating with reported loss of sense of taste and smell as early indicators of COVID-19 infection, including in otherwise asymptomatic patients. Airway club cells, ciliated cells and type II pneumocytes in the lung, and enterocytes of the gut also express ACE2. This study demonstrates how a holistic data science platform can leverage unprecedented quantities of structured and unstructured publicly available data to accelerate the generation of impactful biological insights and hypotheses.

q-bio.GN

Identifying and characterizing extrapolation in multivariate response data

Extrapolation is defined as making predictions beyond the range of the data used to estimate a statistical model. In ecological studies, it is not always obvious when and where extrapolation occurs because of the multivariate nature of the data. Previous work on identifying extrapolation has focused on univariate response data, but these methods are not directly applicable to multivariate response data, which are more and more common in ecological investigations. In this paper, we extend previous work that identified extrapolation by applying the predictive variance from the univariate setting to the multivariate case. We illustrate our approach through an analysis of jointly modeled lake nutrients and indicators of algal biomass and water clarity in over 7000 inland lakes from across the Northeast and Mid-west US. In addition, we illustrate novel exploratory approaches for identifying regions of covariate space where extrapolation is more likely to occur using classification and regression trees.

stat.AP