SearcharxivSearch

arXiv subjects

Nathaniel T. Stevens

Publications and source records attributed to Nathaniel T. Stevens.

9 recordsLinked to original sources

Flexible Method Comparison with the Probability of Agreement

The comparison of methods of measurement is a common problem in clinical practice; as novel methods are developed, establishing their agreement with existing methods is crucial. The probability of agreement (PoA) has previously been proposed as an intuitive and informative means of assessing agreement between two methods of measurement. It straightforwardly quantifies the likelihood that two measurements by different methods on the same subject are clinically indistinguishable. In this paper, we overhaul and extend the PoA methodology by developing an inference framework that relaxes several restrictive assumptions made in previous implementations, ultimately increasing its utility in a wider range of applications. We illustrate this more flexible methodology in an example that compares methods of measuring total Prostatic Specific Antigen (tPSA). And we thoroughly investigate its performance via simulation. This work dramatically increases the flexibility, availability, and hence impact of the PoA approach for method comparison.

stat.ME

Design of Bayesian A/B Tests Controlling False Discovery Rates and Power

Businesses frequently run online controlled experiments (i.e., A/B tests) to learn about the effect of an intervention on multiple business metrics. To account for multiple hypothesis testing, multiple metrics are commonly aggregated into a single composite measure, losing valuable information, or strict family-wise error rate adjustments are imposed, leading to reduced power. In this paper, we propose an economical framework to design Bayesian A/B tests while controlling both power and the false discovery rate (FDR). Selecting optimal decision thresholds to control power and the FDR typically relies on intensive simulation at each sample size considered. Our framework efficiently recommends optimal sample sizes and decision thresholds for Bayesian A/B tests that satisfy criteria for the FDR and average power. Our approach is efficient because we leverage new theoretical results to obtain these recommendations using simulations conducted at only two sample sizes. Our methodology is illustrated using an example based on a real A/B test involving several metrics.

stat.ME

An Economical Approach to Design Posterior Analyses

To design Bayesian studies, criteria for the operating characteristics of posterior analyses - such as power and the type I error rate - are often assessed by estimating sampling distributions of posterior probabilities via simulation. In this paper, we propose an economical method to determine optimal sample sizes and decision criteria for such studies. Using our theoretical results that model posterior probabilities as a function of the sample size, we assess operating characteristics throughout the sample size space given simulations conducted at only two sample sizes. These theoretical results are used to construct bootstrap confidence intervals for the optimal sample sizes and decision criteria that reflect the stochastic nature of simulation-based design. We also repurpose the simulations conducted in our approach to efficiently investigate various sample sizes and decision criteria using contour plots. The broad applicability and wide impact of our methodology is illustrated using two clinical examples.

stat.ME

Fast Power Curve Approximation for Posterior Analyses

Bayesian hypothesis tests leverage posterior probabilities, Bayes factors, or credible intervals to inform data-driven decision making. We propose a framework for power curve approximation with such hypothesis tests. We present a fast approach to explore the approximate sampling distribution of posterior probabilities when the conditions for the Bernstein-von Mises theorem are satisfied. We extend that approach to consider segments of such sampling distributions in a targeted manner for each sample size explored. These sampling distribution segments are used to construct power curves for various types of posterior analyses. Our resulting method for power curve approximation is orders of magnitude faster than conventional power curve estimation for Bayesian hypothesis tests. We also prove the consistency of the corresponding power estimates and sample size recommendations under certain conditions.

stat.ME

Bioequivalence Design with Sampling Distribution Segments

In bioequivalence design, power analyses dictate how much data must be collected to detect the absence of clinically important effects. Power is computed as a tail probability in the sampling distribution of the pertinent test statistics. When these test statistics cannot be constructed from pivotal quantities, their sampling distributions are approximated via repetitive, time-intensive computer simulation. We propose a novel simulation-based method to quickly approximate the power curve for many such bioequivalence tests by efficiently exploring segments (as opposed to the entirety) of the relevant sampling distributions. Despite not estimating the entire sampling distribution, this approach prompts unbiased sample size recommendations. We illustrate this method using two-group bioequivalence tests with unequal variances and overview its broader applicability in clinical design. All methods proposed in this work can be implemented using the developed dent package in R.

stat.ME

Posterior Ramifications of Prior Dependence Structures

Prior elicitation methods for Bayesian analyses transfigure prior information into quantifiable prior distributions. Recently, methods that leverage copulas have been proposed to accommodate more flexible dependence structures when eliciting multivariate priors. We show that the posterior cannot retain many of these flexible prior dependence structures in large-sample settings, and we emphasize that it is our responsibility as statisticians to communicate this to practitioners. We therefore overview objectives for prior specification that guide conversations between statisticians and practitioners to promote alignment between the flexibility in the prior dependence structure and the objectives for posterior analysis. Because correctly specifying the dependence structure a priori can be difficult, we consider how the choice of prior copula impacts the posterior distribution in terms of asymptotic convergence of the posterior mode. Our resulting recommendations clarify when it is useful to elicit intricate prior dependence structures and when it is not.

stat.ME

An Economical Approach to Design with Precision Criteria

Estimation frameworks for statistical inference are preferred to hypothesis testing when quantifying uncertainty and precise estimation are more valuable than binary decisions about statistical significance. Study design for estimation-based investigations often uses precision criteria to select sample sizes that control the length of interval estimates with respect to a sampling distribution. In this paper, we formally define the length probability distribution which characterizes the probability of obtaining a sufficiently narrow interval estimate as a function of the sample size. This distribution can then be used to determine the smallest sample size needed to ensure an interval estimate is sufficiently narrow. We prove that this distribution is approximately normal in large-sample settings for many data generation processes. However, this approximate normality may not hold for studies with moderate sample sizes, particularly when incorporating prior information or obtaining asymmetric interval estimates. Thus, we also propose an efficient simulation-based approach that estimates the sampling distribution of interval estimate lengths at only two sample sizes. Our methodology provides a unified framework for design with precision criteria in Bayesian and frequentist settings with parametric, semiparametric, and nonparametric inference. We illustrate the broad applicability of this framework with various examples.

stat.ME

Monitoring dynamic networks: a simulation-based strategy for comparing monitoring methods and a comparative study

Recently there has been a lot of interest in monitoring and identifying changes in dynamic networks, which has led to the development of a variety of monitoring methods. Unfortunately, these methods have not been systematically compared; moreover, new methods are often designed for a specialized use case. In light of this, we propose the use of simulation to compare the performance of network monitoring methods over a variety of dynamic network changes. Using our family of simulated dynamic networks, we compare the performance of several state-of-the-art social network monitoring methods in the literature. We compare their performance over a variety of types of change; we consider both increases in communication levels, node propensity change as well as changes in community structure. We show that there does not exist one method that is uniformly superior to the others; the best method depends on the context and the type of change one wishes to detect. As such, we conclude that a variety of methods is needed for network monitoring and that it is important to understand in which scenarios a given method is appropriate.

stat.CO

Modeling and detecting change in temporal networks via a dynamic degree corrected stochastic block model

In many applications it is of interest to identify anomalous behavior within a dynamic interacting system. Such anomalous interactions are reflected by structural changes in the network representation of the system. We propose and investigate the use of a dynamic version of the degree corrected stochastic block model (DCSBM) to model and monitor dynamic networks that undergo a significant structural change. We apply statistical process monitoring techniques to the estimated parameters of the DCSBM to identify significant structural changes in the network. Application of our surveillance strategy to the dynamic U.S. Senate co-voting network reveals that we are able to detect significant changes in the network that reflect both times of cohesion and times of polarization among Republican and Democratic party members. These findings provide valuable insight about the evolution of the bipartisan political system in the United States. Our analysis demonstrates that the dynamic DCSBM monitoring procedure effectively detects local and global structural changes in dynamic networks. The DCSBM approach is an example of a more general framework that combines parametric random graph models and statistical process monitoring techniques for network surveillance.

stat.ME