SearcharxivSearch

arXiv subjects

Paul N. Patrone

Publications and source records attributed to Paul N. Patrone.

At least 19 recordsLinked to original sources

Wafer-Level Prototyping Tools for CMOS Bioelectronic Sensors

Integrating biology with complementary metal-oxide-semiconductor (CMOS) sensors can enable highly parallel measurements with minimal parasitic effects, significantly enhancing sensitivity. However, realizing this potential often requires overcoming substantial barriers related to design, fabrication, and heterogeneous integration. In this context, we present a comprehensive suite of tools and methods designed for wafer-scale biosensor prototyping that is sensitive, highly parallelizable, and manufacturable. A central component of our approach is a new initiative that allows for open-source multi-project wafers (MPW), giving all participants access to the designs submitted by others. We demonstrate that this strategy not only promotes design reuse but also facilitates advanced back-end-of-line (BEOL) fabrication techniques, improving the manufacturability and process yield of CMOS biosensors. Developing CMOS-based biosensors also involves the challenge of heterogeneous integration, which includes external electrical, mechanical, and fluid layers. We demonstrate simple modular designs that enable such integration for sample delivery and signal readout. Finally, we showcase the effectiveness of our approach in measuring the hybridization of DNA molecules by focusing on data acquisition and machine learning (ML) methods that leverage the parallelism of the sensors to enable robust classification of desirable analyte interactions.

physics.ins-det

Inequalities for Optimization of Classification Algorithms: A Perspective Motivated by Diagnostic Testing

Motivated by canonical problems in medical diagnostics, we propose and study properties of an objective function that uniformly bounds uncertainties in quantities of interest extracted from classifiers and related data analysis tools. We begin by adopting a set-theoretic perspective to show how two main tasks in diagnostics -- classification and prevalence estimation -- can be recast in terms of a variation on the confusion (or error) matrix ${\boldsymbol {\rm P}}$ typically considered in supervised learning. We then combine arguments from conditional probability with the Gershgorin circle theorem to demonstrate that the largest Gershgorin radius $\boldsymbol ρ_m$ of the matrix $\mathbb I-\boldsymbol {\rm P}$ (where $\mathbb I$ is the identity) yields uniform error bounds for both classification and prevalence estimation. In a two-class setting, $\boldsymbol ρ_m$ is minimized via a measure-theoretic ``water-leveling'' argument that optimizes an appropriately defined partition $U$ generating the matrix ${\boldsymbol {\rm P}}$. We also consider an example that illustrates the difficulty of generalizing the binary solution to a multi-class setting and deduce relevant properties of the confusion matrix.

stat.ML

Per-event Uncertainty Quantification for Flow Cytometry using Calibration Beads

Flow cytometry measurements are widely used in diagnostics and medical decision making. Incomplete understanding of sources of measurement uncertainty can make it difficult to distinguish autofluorescence and background sources from signals of interest. Moreover, established methods for modeling uncertainty overlook the fact that the apparent distribution of measurements is a convolution of the inherent the population variability (e.g., associated with calibration beads or cells) and instrument induced-effects. Such issues make it difficult, for example, to identify signals from small objects such as extracellular vesicles. To overcome such limitations, we formulate an explicit probabilistic measurement model that accounts for volume and labeling variation, background signals and fluorescence shot noise. Using raw data from routine per-event calibration measurements, we use this model to separate the aforementioned sources of uncertainty and demonstrate how such information can be used to facilitate decision-making and instrument characterization.

q-bio.QM

Uncertainty Quantification of Antibody Measurements: Physical Principles and Implications for Standardization

Harmonizing serology measurements is critical for identifying reference materials that permit standardization and comparison of results across different diagnostic platforms. However, the theoretical foundations of such tasks have yet to be fully explored in the context of antibody thermodynamics and uncertainty quantification (UQ). This has restricted the usefulness of standards currently deployed and limited the scope of materials considered as viable reference material. To address these problems, we develop rigorous theories of antibody normalization and harmonization, as well as formulate a probabilistic framework for defining correlates of protection. We begin by proposing a mathematical definition of harmonization equipped with structure needed to quantify uncertainty associated with the choice of standard, assay, etc. We then show how a thermodynamic description of serology measurements (i) relates this structure to the Gibbs free-energy of antibody binding, and thereby (ii) induces a regression analysis that directly harmonizes measurements. We supplement this with a novel, optimization-based normalization (not harmonization!) method that checks for consistency between reference and sample dilution curves. Last, we relate these analyses to uncertainty propagation techniques to estimate correlates of protection. A key result of these analyses is that under physically reasonable conditions, the choice of reference material does not increase uncertainty associated with harmonization or correlates of protection. We provide examples and validate main ideas in the context of an interlab study that lays the foundation for using monoclonal antibodies as a reference for SARS-CoV-2 serology measurements.

physics.bio-ph

Analysis of Diagnostics (Part I): Prevalence, Uncertainty Quantification, and Machine Learning

Diagnostic testing provides a unique setting for studying and developing tools in classification theory. In such contexts, the concept of prevalence, i.e. the number of individuals with a given condition, is fundamental, both as an inherent quantity of interest and as a parameter that controls classification accuracy. This manuscript is the first in a two-part series that studies deeper connections between classification theory and prevalence, showing how the latter establishes a more complete theory of uncertainty quantification (UQ) for certain types of machine learning (ML). We motivate this analysis via a lemma demonstrating that general classifiers minimizing a prevalence-weighted error contain the same probabilistic information as Bayes-optimal classifiers, which depend on conditional probability densities. This leads us to study relative probability level-sets $B^\star (q)$, which are reinterpreted as both classification boundaries and useful tools for quantifying uncertainty in class labels. To realize this in practice, we also propose a numerical, homotopy algorithm that estimates the $B^\star (q)$ by minimizing a prevalence-weighted empirical error. The successes and shortcomings of this method motivate us to revisit properties of the level sets, and we deduce the corresponding classifiers obey a useful monotonicity property that stabilizes the numerics and points to important extensions to UQ of ML. Throughout, we validate our methods in the context of synthetic data and a research-use-only SARS-CoV-2 enzyme-linked immunosorbent (ELISA) assay.

stat.ML

Analysis of Diagnostics (Part II): Prevalence, Linear Independence, and Unsupervised Learning

This is the second manuscript in a two-part series that uses diagnostic testing to understand the connection between prevalence (i.e. number of elements in a class), uncertainty quantification (UQ), and classification theory. Part I considered the context of supervised machine learning (ML) and established a duality between prevalence and the concept of relative conditional probability. The key idea of that analysis was to train a family of discriminative classifiers by minimizing a sum of prevalence-weighted empirical risk functions. The resulting outputs can be interpreted as relative probability level-sets, which thereby yield uncertainty estimates in the class labels. This procedure also demonstrated that certain discriminative and generative ML models are equivalent. Part II considers the extent to which these results can be extended to tasks in unsupervised learning through recourse to ideas in linear algebra. We first observe that the distribution of an impure population, for which the class of a corresponding sample is unknown, can be parameterized in terms of a prevalence. This motivates us to introduce the concept of linearly independent populations, which have different but unknown prevalence values. Using this, we identify an isomorphism between classifiers defined in terms of impure and pure populations. In certain cases, this also leads to a nonlinear system of equations whose solution yields the prevalence values of the linearly independent populations, fully realizing unsupervised learning as a generalization of supervised learning. We illustrate our methods in the context of synthetic data and a research-use-only SARS-CoV-2 enzyme-linked immunosorbent assay (ELISA).

stat.ML

Modeling in higher dimensions to improve diagnostic testing accuracy: theory and examples for multiplex saliva-based SARS-CoV-2 antibody assays

The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) pandemic has emphasized the importance and challenges of correctly interpreting antibody test results. Identification of positive and negative samples requires a classification strategy with low error rates, which is hard to achieve when the corresponding measurement values overlap. Additional uncertainty arises when classification schemes fail to account for complicated structure in data. We address these problems through a mathematical framework that combines high dimensional data modeling and optimal decision theory. Specifically, we show that appropriately increasing the dimension of data better separates positive and negative populations and reveals nuanced structure that can be described in terms of mathematical models. We combine these models with optimal decision theory to yield a classification scheme that better separates positive and negative samples relative to traditional methods such as confidence intervals (CIs) and receiver operating characteristics. We validate the usefulness of this approach in the context of a multiplex salivary SARS-CoV-2 immunoglobulin G assay dataset. This example illustrates how our analysis: (i) improves the assay accuracy (e.g. lowers classification errors by up to 42 % compared to CI methods); (ii) reduces the number of indeterminate samples when an inconclusive class is permissible (e.g. by 40 % compared to the original analysis of the example multiplex dataset); and (iii) decreases the number of antigens needed to classify samples. Our work showcases the power of mathematical modeling in diagnostic classification and highlights a method that can be adopted broadly in public health and clinical settings.

q-bio.QM

Optimal classification and generalized prevalence estimates for diagnostic settings with more than two classes

An accurate multiclass classification strategy is crucial to interpreting antibody tests. However, traditional methods based on confidence intervals or receiver operating characteristics lack clear extensions to settings with more than two classes. We address this problem by developing a multiclass classification based on probabilistic modeling and optimal decision theory that minimizes the convex combination of false classification rates. The classification process is challenging when the relative fraction of the population in each class, or generalized prevalence, is unknown. Thus, we also develop a method for estimating the generalized prevalence of test data that is independent of classification. We validate our approach on serological data with severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) naïve, previously infected, and vaccinated classes. Synthetic data are used to demonstrate that (i) prevalence estimates are unbiased and converge to true values and (ii) our procedure applies to arbitrary measurement dimensions. In contrast to the binary problem, the multiclass setting offers wide-reaching utility as the most general framework and provides new insight into prevalence estimation best practices.

q-bio.QM

Prevalence Estimation and Optimal Classification Methods to Account for Time Dependence in Antibody Levels

Serology testing can identify past infection by quantifying the immune response of an infected individual providing important public health guidance. Individual immune responses are time-dependent, which is reflected in antibody measurements. Moreover, the probability of obtaining a particular measurement changes due to prevalence as the disease progresses. Taking into account these personal and population-level effects, we develop a mathematical model that suggests a natural adaptive scheme for estimating prevalence as a function of time. We then combine the estimated prevalence with optimal decision theory to develop a time-dependent probabilistic classification scheme that minimizes error. We validate this analysis by using a combination of real-world and synthetic SARS-CoV-2 data and discuss the type of longitudinal studies needed to execute this scheme in real-world settings.

q-bio.QM

Reproducibility in Cytometry: Signals Analysis and its Connection to Uncertainty Quantification

Signals analysis for cytometry remains a challenging task that has a significant impact on uncertainty. Conventional cytometers assume that individual measurements are well characterized by simple properties such as the signal area, width, and height. However, these approaches have difficulty distinguishing inherent biological variability from instrument artifacts and operating conditions. As a result, it is challenging to quantify uncertainty in the properties of individual cells and perform tasks such as doublet deconvolution. We address these problems via signals analysis techniques that use scale transformations to: (I) separate variation in biomarker expression from effects due to flow conditions and particle size; (II) quantify reproducibility associated with a given laser interrogation region; (III) estimate uncertainty in measurement values on a per-event basis; and (IV) extract the singlets that make up a multiplet. The key idea behind this approach is to model how variable operating conditions deform the signal shape and then use constrained optimization to "undo" these deformations for measured signals; residuals to this process characterize reproducibility. Using a recently developed microfluidic cytometer, we demonstrate that these techniques can account for instrument and measurand induced variability with a residual uncertainty of less than 2.5% in the signal shape and less than 1% in integrated area.

physics.bio-ph

Optimal Decision Theory for Diagnostic Testing: Minimizing Indeterminate Classes with Applications to Saliva-Based SARS-CoV-2 Antibody Assays

In diagnostic testing, establishing an indeterminate class is an effective way to identify samples that cannot be accurately classified. However, such approaches also make testing less efficient and must be balanced against overall assay performance. We address this problem by reformulating data classification in terms of a constrained optimization problem that (i) minimizes the probability of labeling samples as indeterminate while (ii) ensuring that the remaining ones are classified with an average target accuracy X. We show that the solution to this problem is expressed in terms of a bathtub principle that holds out those samples with the lowest local accuracy up to an X-dependent threshold. To illustrate the usefulness of this analysis, we apply it to a multiplex, saliva-based SARS-CoV-2 antibody assay and demonstrate up to a 30 % reduction in the number of indeterminate samples relative to more traditional approaches.

stat.ME

Classification Under Uncertainty: Data Analysis for Diagnostic Antibody Testing

Formulating accurate and robust classification strategies is a key challenge of developing diagnostic and antibody tests. Methods that do not explicitly account for disease prevalence and uncertainty therein can lead to significant classification errors. We present a novel method that leverages optimal decision theory to address this problem. As a preliminary step, we develop an analysis that uses an assumed prevalence and conditional probability models of diagnostic measurement outcomes to define optimal (in the sense of minimizing rates of false positives and false negatives) classification domains. Critically, we demonstrate how this strategy can be generalized to a setting in which the prevalence is unknown by either: (i) defining a third class of hold-out samples that require further testing; or (ii) using an adaptive algorithm to estimate prevalence prior to defining classification domains. We also provide examples for a recently published SARS-CoV-2 serology test and discuss how measurement uncertainty (e.g. associated with instrumentation) can be incorporated into the analysis. We find that our new strategy decreases classification error by up to a decade relative to more traditional methods based on confidence intervals. Moreover, it establishes a theoretical foundation for generalizing techniques such as receiver operating characteristics (ROC) by connecting them to the broader field of optimization.

stat.ME

Improving Baseline Subtraction for Increased Sensitivity of Quantitative PCR Measurements

Motivated by the current COVID-19 health-crisis, we examine the task of baseline subtraction for quantitative polymerase chain-reaction (qPCR) measurements. In particular, we present an algorithm that leverages information obtained from non-template and/or DNA extraction-control experiments to remove systematic bias from amplification curves. We recast this problem in terms of mathematical optimization, i.e. by finding the amount of control signal that, when subtracted from an amplification curve, minimizes background noise. We demonstrate that this approach can yield a decade improvement in sensitivity relative to standard approaches, especially for data exhibiting late-cycle amplification. Critically, this increased sensitivity and accuracy promises more effective screening of viral DNA and a reduction in the rate of false-negatives in diagnostic settings.

q-bio.QM

Towards a priori uncertainty quantification in coarse-grained molecular dynamics: Generalized multipole potentials

In computational materials science, coarse-graining approaches often lack a priori uncertainty quantification (UQ) tools that estimate the accuracy of a reduced-order model before it is calibrated or deployed. This is especially the case in coarse-grained (CG) molecular dynamics (MD), where "bottom-up" methods need to run expensive atomistic simulations as part of the calibration process. As a result, scientists have been slow to adopt CG techniques in many settings because they do not know in advance whether the cost of developing the CG model is justified. To address this problem, we present an analytical method of coarse-graining rigid-body systems that yields corresponding intermolecular potentials with controllable levels of accuracy relative to their atomistic counterparts. Critically, this analysis: (i) provides a mathematical foundation for assessing the quality of a CG force field without running simulations, and (ii) provides a tool for understanding how atomistic systems can be viewed as appropriate limits of reduced-order models. Simulated results confirm the validity of this approach at the trajectory level and point to issues that must be addressed in coarse-graining fully non-rigid systems.

physics.comp-ph

Uncertainty Quantification for Molecular Dynamics

The goals of this chapter are twofold. First, we wish to introduce molecular dynamics (MD) and uncertainty quantification (UQ) in a common setting in order to demonstrate how the latter can increase confidence in the former. In some cases, this discussion culminates in our providing practical, mathematical tools that can be used to answer the question, "is this simulation reliable?" However, many questions remain unanswered. Thus, a second goal of this work is to highlight open problems where progress would aid the larger community.

physics.comp-ph

The Role of Data Analysis in Uncertainty Quantification: Case Studies for Materials Modeling

In computational materials science, mechanical properties are typically extracted from simulations by means of analysis routines that seek to mimic their experimental counterparts. However, simulated data often exhibit uncertainties that can propagate into final predictions in unexpected ways. Thus, modelers require data analysis tools that (i) address the problems posed by simulated data, and (ii) facilitate uncertainty quantification. In this manuscript, we discuss three case studies in materials modeling where careful data analysis can be leveraged to address specific instances of these issues. As a unifying theme, we highlight the idea that attention to physical and mathematical constraints surrounding the generation of computational data can significantly enhance its analysis.

physics.data-an

Estimating yield-strain via deformation-recovery simulations

In computational materials science, predicting the yield strain of crosslinked polymers remains a challenging task. A common approach is to identify yield as the first critical point of stress-strain curves simulated by molecular dynamics (MD). However, in such cases the underlying data can be excessively noisy, making it difficult to extract meaningful results. In this work, we propose an alternate method for identifying yield on the basis of deformation-recovery simulations. Notably, the corresponding raw data (i.e. residual strains) produce a sharper signal for yield via a transition in their global behavior. We analyze this transition by non- linear regression of computational data to a hyperbolic model. As part of this analysis, we also propose uncertainty quantification techniques for assessing when and to what extent the simulated data is informative of yield. Moreover, we show how the method directly tests for yield via the onset of permanent deformation and discuss recent experimental results, which compare favorably with our predictions.

cond-mat.mtrl-sci

Beyond histograms: efficiently estimating radial distribution functions via spectral Monte Carlo

Despite more than 40 years of research in condensed-matter physics, state-of-the-art approaches for simulating the radial distribution function (RDF) g(r) still rely on binning pair-separations into a histogram. Such methods suffer from undesirable properties, including subjectivity, high uncertainty, and slow rates of convergence. Moreover, such problems go undetected by the metrics often used to assess RDFs. To address these issues, we propose (I) a spectral Monte Carlo (SMC) method that yields g(r) as an analytical series expansion; and (II) a Sobolev norm that assesses the quality of RDFs by quantifying their fluctuations. Using the latter, we show that, relative to histogram-based approaches, SMC reduces by orders of magnitude both the noise in g(r) and the number of pair separations needed for acceptable convergence. Moreover, SMC reduces subjectivity and yields simple, differentiable formulas for the RDF, which are useful for tasks such as coarse-grained force-field calibration via iterative Boltzmann inversion.

cond-mat.mtrl-sci