SearcharxivSearch

arXiv subjects

Michael J. Higgins

Publications and source records attributed to Michael J. Higgins.

9 recordsLinked to original sources

Detecting Treatment Interference under the K-Nearest-Neighbors Interference Model

We propose a model of treatment interference where the response of a unit depends only on its treatment status and the statuses of units within its K-neighborhood. Current methods for detecting interference include carefully designed randomized experiments and conditional randomization tests on a set of focal units. We give guidance on how to choose focal units under this model of interference. We then conduct a simulation study to evaluate the efficacy of existing methods for detecting network interference. We show that this choice of focal units leads to powerful tests of treatment interference which outperform current experimental methods.

stat.ME

Demystifying Statistical Matching Algorithms for Big Data

Statistical matching is an effective method for estimating causal effects in which treated units are paired with control units with ``similar'' values of confounding covariates prior to performing estimation. In this way, matching helps isolate the effect of treatment on response from effects due to the confounding covariates. While there are a large number of software packages to perform statistical matching, the algorithms and techniques used to solve statistical matching problems -- especially matching without replacement -- are not widely understood. In this paper, we describe in detail commonly-used algorithms and techniques for solving statistical matching problems. We focus in particular on the efficiency of these algorithms as the number of observations grow large. We advocate for the further development of statistical matching methods that impose and exploit ``sparsity'' -- by greatly restricting the available matches for a given treated unit -- as this may be critical to ensure scalability of matching methods as data sizes grow large.

stat.ME

Estimation of Causal Effects Under K-Nearest Neighbors Interference

Considerable recent work has focused on methods for analyzing experiments which exhibit treatment interference -- that is, when the treatment status of one unit may affect the response of another unit. Such settings are common in experiments on social networks. We consider a model of treatment interference -- the K-nearest neighbors interference model (KNNIM) -- for which the response of one unit depends not only on the treatment status given to that unit, but also the treatment status of its $K$ ``closest'' neighbors. We derive causal estimands under KNNIM in a way that allows us to identify how each of the $K$-nearest neighbors contributes to the indirect effect of treatment. We propose unbiased estimators for these estimands and derive conservative variance estimates for these unbiased estimators. We then consider extensions of these estimators under an assumption of no weak interaction between direct and indirect effects. We perform a simulation study to determine the efficacy of these estimators under different treatment interference scenarios. We apply our methodology to an experiment designed to assess the impact of a conflict-reducing program in middle schools in New Jersey, and we give evidence that the effect of treatment propagates primarily through a unit's closest connection.

stat.ME

From one environment to many: The problem of replicability of statistical inferences

Among plausible causes for replicability failure, one that has not received sufficient attention is the environment in which the research is conducted. Consisting of the population, equipment, personnel, and various conditions such as location, time, and weather, the research environment can affect treatments and outcomes, and changes in the research environment that occur when an experiment is redone can affect replicability. We examine the extent to which such changes contribute to replicability failure. Our framework is that of an initial experiment that generates the data and a follow-up experiment that is done the same way except for a change in the research environment. We assume that the initial experiment satisfies the assumptions of the two-sample t-statistic and that the follow-up experiment is described by a mixed model which includes environmental parameters. We derive expressions for the effect that the research environment has on power, sample size selection, p-values, and confidence levels. We measure the size of the environmental effect with the environmental effect ratio EER which is the ratio of the standard deviations of environment by treatment interaction and error. By varying EER, it is possible to determine conditions that favor replicability and those that do not.

stat.ME

The Benefits of Probability-Proportional-to-Size Sampling in Cluster-Randomized Experiments

In a cluster-randomized experiment, treatment is assigned to clusters of individual units of interest--households, classrooms, villages, etc.--instead of the units themselves. The number of clusters sampled and the number of units sampled within each cluster is typically restricted by a budget constraint. Previous analysis of cluster randomized experiments under the Neyman-Rubin potential outcomes model of response have assumed a simple random sample of clusters. Estimators of the population average treatment effect (PATE) under this assumption are often either biased or not invariant to location shifts of potential outcomes. We demonstrate that, by sampling clusters with probability proportional to the number of units within a cluster, the Horvitz-Thompson estimator (HT) is invariant to location shifts and unbiasedly estimates PATE. We derive standard errors of HT and discuss how to estimate these standard errors. We also show that results hold for stratified random samples when samples are drawn proportionally to cluster size within each stratum. We demonstrate the efficacy of this sampling scheme using a simulation based on data from an experiment measuring the efficacy of the National Solidarity Programme in Afghanistan.

stat.ME

A new method for quantifying network cyclic structure to improve community detection

A distinguishing property of communities in networks is that cycles are more prevalent within communities than across communities. Thus, the detection of these communities may be aided through the incorporation of measures of the local "richness" of the cyclic structure. In this paper, we introduce renewal non-backtracking random walks (RNBRW) as a way of quantifying this structure. RNBRW gives a weight to each edge equal to the probability that a non-backtracking random walk completes a cycle with that edge. Hence, edges with larger weights may be thought of as more important to the formation of cycles. Of note, since separate random walks can be performed in parallel, RNBRW weights can be estimated very quickly, even for large graphs. We give simulation results showing that pre-weighting edges through RNBRW may substantially improve the performance of common community detection algorithms. Our results suggest that RNBRW is especially efficient for the challenging case of detecting communities in sparse graphs.

cs.SI

Generalized full matching and extrapolation of the results from a large-scale voter mobilization experiment

Matching is an important tool in causal inference. The method provides a conceptually straightforward way to make groups of units comparable on observed characteristics. The use of the method is, however, limited to situations where the study design is fairly simple and the sample is moderately sized. We illustrate the issue by revisiting a large-scale voter mobilization experiment that took place in Michigan for the 2006 election. We ask what the causal effects would have been if the treatments in the experiment were scaled up to the full population. Matching could help us answer this question, but no existing matching method can accommodate the six treatment arms and the 6,762,701 observations involved in the study. To offer a solution this and similar empirical problems, we introduce a generalization of the full matching method and an associated algorithm. The method can be used with any number of treatment conditions, and it is shown to produce near-optimal matchings. The worst case maximum within-group dissimilarity is no worse than four times the optimal solution, and simulation results indicate that its performance is considerably closer to the optimal solution on average. Despite its performance, the algorithm is fast and uses little memory. It terminates, on average, in linearithmic time using linear space. This enables investigators to construct well-performing matchings within minutes even in complex studies with samples of several million units.

stat.ME

A virtual instrument to standardise the calibration of atomic force microscope cantilevers

Atomic force microscope (AFM) users often calibrate the spring constants of cantilevers using functionality built into individual instruments. This is performed without reference to a global standard, which hinders robust comparison of force measurements reported by different laboratories. In this article, we describe a virtual instrument (an internet-based initiative) whereby users from all laboratories can instantly and quantitatively compare their calibration measurements to those of others - standardising AFM force measurements - and simultaneously enabling non-invasive calibration of AFM cantilevers of any geometry. This global calibration initiative requires no additional instrumentation or data processing on the part of the user. It utilises a single website where users upload currently available data. A proof-of-principle demonstration of this initiative is presented using measured data from five independent laboratories across three countries, which also allows for an assessment of current calibration.

physics.ins-det