SearcharxivSearch

arXiv subjects

Christopher D. Heaney

Publications and source records attributed to Christopher D. Heaney.

3 recordsLinked to original sources

Source apportionment of air pollution burden using geometric non-negative matrix factorization and high-throughput multi-pollutant air sensor data in Curtis Bay, Baltimore, USA

Air sensor networks provide hyperlocal, high-frequency data on multiple pollutants, but unlike speciated particulate matter (PM) measurements, they lack direct chemical signatures for source identification. High temporal resolution and multiple spatial locations nonetheless create new opportunities to interpret latent sources through their relationships with spatial proximity to known origins, temporal patterns, and meteorology. We analyze 451946 one-minute air sensor records from Curtis Bay (Baltimore, USA; October 2022 - June 2023), covering size-resolved PM, black carbon (BC), carbon monoxide (CO), nitric oxide (NO), and nitrogen dioxide (NO2), using a geometric non-negative matrix factorization (NMF) approach that scales to large datasets and yields provably unique source attribution percentages. Three stable latent sources emerge with converging evidence toward recognizable source categories: Source 1 explains $>$ 70% of fine and coarse PM and $\sim$30% of BC; Source 2 dominates CO and contributes $\sim$70% of BC, NO, and NO2; Source 3 is specific to the larger PM fractions, PM10 to PM40. Regression analyses and a case study on a known bulldozer incident link Sources 1 and 3 to a nearby coal terminal. Extreme-intensity episodes from Sources 1 and 3 averaged $\sim$33 and $\sim$24 minutes per day at the site nearest the terminal, attenuating with distance. Source 2 reflects diurnal traffic patterns. Together, these results show that dense air sensor networks paired with the geometric NMF method can move community air monitoring beyond pollution detection toward identifying likely source categories and informing actionable mitigation strategies.

stat.AP

Modeling in higher dimensions to improve diagnostic testing accuracy: theory and examples for multiplex saliva-based SARS-CoV-2 antibody assays

The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) pandemic has emphasized the importance and challenges of correctly interpreting antibody test results. Identification of positive and negative samples requires a classification strategy with low error rates, which is hard to achieve when the corresponding measurement values overlap. Additional uncertainty arises when classification schemes fail to account for complicated structure in data. We address these problems through a mathematical framework that combines high dimensional data modeling and optimal decision theory. Specifically, we show that appropriately increasing the dimension of data better separates positive and negative populations and reveals nuanced structure that can be described in terms of mathematical models. We combine these models with optimal decision theory to yield a classification scheme that better separates positive and negative samples relative to traditional methods such as confidence intervals (CIs) and receiver operating characteristics. We validate the usefulness of this approach in the context of a multiplex salivary SARS-CoV-2 immunoglobulin G assay dataset. This example illustrates how our analysis: (i) improves the assay accuracy (e.g. lowers classification errors by up to 42 % compared to CI methods); (ii) reduces the number of indeterminate samples when an inconclusive class is permissible (e.g. by 40 % compared to the original analysis of the example multiplex dataset); and (iii) decreases the number of antigens needed to classify samples. Our work showcases the power of mathematical modeling in diagnostic classification and highlights a method that can be adopted broadly in public health and clinical settings.

q-bio.QM

Optimal Decision Theory for Diagnostic Testing: Minimizing Indeterminate Classes with Applications to Saliva-Based SARS-CoV-2 Antibody Assays

In diagnostic testing, establishing an indeterminate class is an effective way to identify samples that cannot be accurately classified. However, such approaches also make testing less efficient and must be balanced against overall assay performance. We address this problem by reformulating data classification in terms of a constrained optimization problem that (i) minimizes the probability of labeling samples as indeterminate while (ii) ensuring that the remaining ones are classified with an average target accuracy X. We show that the solution to this problem is expressed in terms of a bathtub principle that holds out those samples with the lowest local accuracy up to an X-dependent threshold. To illustrate the usefulness of this analysis, we apply it to a multiplex, saliva-based SARS-CoV-2 antibody assay and demonstrate up to a 30 % reduction in the number of indeterminate samples relative to more traditional approaches.

stat.ME