SearcharxivSearch

arXiv subjects

Pallavi Basu

Publications and source records attributed to Pallavi Basu.

10 recordsLinked to original sources

Randomization tests for model specification in causal inference under network interference

Analysis of experimental data becomes challenging when the underlying population is connected by a network. Exposure mapping is a common tool in the literature for defining and estimating spillover effects. These mappings reduce the dimensionality of the estimand, thereby facilitating identifiability. It is assumed that this mapping is correctly specified, leaving the choice of the exposure mapping to the analyst. This makes estimators of the spillover effect, such as the Horvitz-Thompson estimator, vulnerable to bias from model misspecification. Although these estimators have been shown to be robust to certain forms of controlled misspecification, there has been relatively little methodological progress in empirically investigating appropriate exposure mappings. In this paper, we propose a novel design-based model specification framework for causal inference. Building on this, we develop a randomization-testing procedure to assess the correct specification of an exposure-mapping model in the presence of network interference. We provide theoretical guarantees for the asymptotic validity of the proposed testing procedure. We establish the favorable power properties of our method through an extensive simulation study and illustrate it in a field experiment investigating the effect of anti-conflict norms among adolescents.

stat.ME

Ranking by Lifts: A Cost-Benefit Approach to Large-Scale A/B Tests

A/B testing is a core tool for decision-making in business experimentation, particularly in digital platforms and marketplaces. Practitioners often prioritize lift in performance metrics while seeking to control the costs of false discoveries. This paper develops a decision-theoretic framework for maximizing expected profit subject to a constraint on the cost-weighted false discovery rate (FDR). We propose an empirical Bayes approach that uses a greedy knapsack algorithm to rank experiments based on the ratio of expected lift to cost, incorporating the local false discovery rate (lfdr) as a key statistic. The resulting oracle rule is valid and rank-optimal. In large-scale settings, we establish the asymptotic validity of a data-driven implementation and demonstrate superior finite-sample performance over existing FDR-controlling methods. An application to A/B tests run on the Optimizely platform highlights the business value of the approach.

stat.ME

Quasi-randomization tests for network interference: a random graph approach

Network interference occurs when the treatment status of one unit affects the potential outcomes of other units, giving rise to spillover effects that are difficult to test for. We propose treating the network as a random variable rather than a fixed quantity to address this challenge. This overcomes a key challenge of non-imputability of potential outcomes under the null and avoids the computational intractability of existing conditional randomization tests. Our quasi-randomization test builds the null distribution of no spillover effects using random graph null models, is exactly valid in finite samples under mild assumptions on the network-generating process, and offers substantially improved power over existing methods, particularly in cluster-randomized trials. We validate our approach via simulation and illustrate it by testing for interference in a weather insurance adoption experiment in rural China.

stat.ME

Exact confidence intervals for the mixing distribution from binomial mixture distribution samples

We present methodology for constructing pointwise confidence intervals for the cumulative distribution function and the quantiles of mixing distributions on the unit interval from binomial mixture distribution samples. No assumptions are made on the shape of the mixing distribution. The confidence intervals are constructed by inverting exact tests of composite null hypotheses regarding the mixing distribution. Our method may be applied to any deconvolution approach that produces test statistics whose distribution is stochastically monotone for stochastic increase of the mixing distribution. We propose a hierarchical Bayes approach, which uses finite Polya Trees for modelling the mixing distribution, that provides stable and accurate deconvolution estimates without the need for additional tuning parameters. Our main technical result establishes the stochastic monotonicity property of the test statistics produced by the hierarchical Bayes approach. Leveraging the need for the stochastic monotonicity property, we explicitly derive the smallest asymptotic confidence intervals that may be constructed using our methodology. Raising the question whether it is possible to construct smaller confidence intervals for the mixing distribution without making parametric assumptions on its shape.

stat.ME

An Empirical Bayes Approach to Controlling the False Discovery Exceedance

In large-scale multiple hypothesis testing problems, the false discovery exceedance (FDX) provides a desirable alternative to the widely used false discovery rate (FDR) when the false discovery proportion (FDP) is highly variable. We develop an empirical Bayes approach to control the FDX. We show that, for independent hypotheses from a two-group model and dependent hypotheses from a Gaussian model fulfilling the exchangeability condition, an oracle decision rule based on ranking and thresholding the local false discovery rate (lfdr) is optimal in the sense that the power is maximized subject to the FDX constraint. We propose a data-driven FDX procedure that uses carefully designed computational shortcuts to emulate the oracle rule. We investigate the empirical performance of the proposed method using both simulated and real data and study the merits of FDX control through an application for identifying abnormal stock trading strategies.

stat.ME

Constructing a More Closely Matched Control Group in a Difference-in-Differences Analysis: Its Effect on History Interacting with Group Bias

Difference-in-differences analysis with a control group that differs considerably from a treated group is vulnerable to bias from historical events that have different effects on the groups. Constructing a more closely matched control group by matching a subset of the overall control group to the treated group may result in less bias. We study this phenomenon in simulation studies. We study the effect of mountaintop removal mining (MRM) on mortality using a difference-in-differences analysis that makes use of the increase in MRM following the 1990 Clean Air Act Amendments. For a difference-in-differences analysis of the effect of MRM on mortality, we constructed a more closely matched control group and found a 95\% confidence interval that contains substantial adverse effects along with no effect and small beneficial effects.

stat.ME

Large-Scale Model Selection with Misspecification

Model selection is crucial to high-dimensional learning and inference for contemporary big data applications in pinpointing the best set of covariates among a sequence of candidate interpretable models. Most existing work assumes implicitly that the models are correctly specified or have fixed dimensionality. Yet both features of model misspecification and high dimensionality are prevalent in practice. In this paper, we exploit the framework of model selection principles in misspecified models originated in Lv and Liu (2014) and investigate the asymptotic expansion of Bayesian principle of model selection in the setting of high-dimensional misspecified models. With a natural choice of prior probabilities that encourages interpretability and incorporates Kullback-Leibler divergence, we suggest the high-dimensional generalized Bayesian information criterion with prior probability (HGBIC_p) for large-scale model selection with misspecification. Our new information criterion characterizes the impacts of both model misspecification and high dimensionality on model selection. We further establish the consistency of covariance contrast matrix estimation and the model selection consistency of HGBIC_p in ultra-high dimensions under some mild regularity conditions. The advantages of our new method are supported by numerical studies.

stat.ME

Optimal design for high-throughput screening via false discovery rate control

High-throughput screening (HTS) is a large-scale hierarchical process in which a large number of chemicals are tested in multiple stages. Conventional statistical analyses of HTS studies often suffer from high testing error rates and soaring costs in large-scale settings. This article develops new methodologies for false discovery rate control and optimal design in HTS studies. We propose a two-stage procedure that determines the optimal numbers of replicates at different screening stages while simultaneously controlling the false discovery rate in the confirmatory stage subject to a constraint on the total budget. The merits of the proposed methods are illustrated using both simulated and real data. We show that the proposed screening procedure effectively controls the error rate and the design leads to improved detection power. This is achieved at the expense of a limited budget.

stat.AP

Weighted False Discovery Rate Control in Large-Scale Multiple Testing

The use of weights provides an effective strategy to incorporate prior domain knowledge in large-scale inference. This paper studies weighted multiple testing in a decision-theoretic framework. We develop oracle and data-driven procedures that aim to maximize the expected number of true positives subject to a constraint on the weighted false discovery rate. The asymptotic validity and optimality of the proposed methods are established. The results demonstrate that incorporating informative domain knowledge enhances the interpretability of results and precision of inference. Simulation studies show that the proposed method controls the error rate at the nominal level, and the gain in power over existing methods is substantial in many settings. An application to genome-wide association study is discussed.

stat.ME

Model Selection in High-Dimensional Misspecified Models

Model selection is indispensable to high-dimensional sparse modeling in selecting the best set of covariates among a sequence of candidate models. Most existing work assumes implicitly that the model is correctly specified or of fixed dimensions. Yet model misspecification and high dimensionality are common in real applications. In this paper, we investigate two classical Kullback-Leibler divergence and Bayesian principles of model selection in the setting of high-dimensional misspecified models. Asymptotic expansions of these principles reveal that the effect of model misspecification is crucial and should be taken into account, leading to the generalized AIC and generalized BIC in high dimensions. With a natural choice of prior probabilities, we suggest the generalized BIC with prior probability which involves a logarithmic factor of the dimensionality in penalizing model complexity. We further establish the consistency of the covariance contrast matrix estimator in a general setting. Our results and new method are supported by numerical studies.

math.ST