Searcharxiv⌕ Search

arXiv subjects

Thomas S. Richardson

Publications and source records attributed to Thomas S. Richardson.

At least 19 recordsLinked to original sources

Online activity prediction via generalized Indian buffet process models

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism: how many users to expose and for how long. This often requires forecasting user engagement, i.e., whether enough users will trigger, and when a target participation level will be reached, from limited pilot data. We introduce a Bayesian nonparametric model for predicting both new-user counts and total triggers, accommodating the heavy-tailed engagement patterns typical of web experiments. All predictive quantities can be computed without intensive numerical procedures such as Markov chain Monte Carlo (MCMC) or variational inference. We evaluate on three public datasets (over 450 public benchmark evaluations) and a proprietary benchmark drawn from 759 production A/B tests comprising 1,774 arms. Across the benchmark analyses, our models are competitive and frequently improve accuracy in forecasting new users, total triggers, and time to reach a target sample size compared with state-of-the-art competitors, especially when only a few pilot days are observed.

stat.AP↗

The Categorical Instrumental Variable Model: Characterization, Partial Identification, and Statistical Inference

We study categorical instrumental variable (IV) models with instrument, treatment and outcome taking finitely many values. We derive a simple closed-form characterization of the set of joint distributions of potential outcomes that are compatible with a given observed data distribution in terms of a minimal set of inequalities. These inequalities unify several different IV models defined by versions of the independence and exclusion restriction assumptions. They lead to sharp bounds on causal functionals and provide a sharp criterion for model falsification. For linear functionals of the joint counterfactual distribution, such as pairwise average treatment effects and probabilities of potential outcomes, we construct confidence intervals with simultaneous finite-sample coverage, using a tail bound on the Kullback--Leibler divergence. We illustrate our method using data from the Minneapolis Domestic Violence Experiment.

math.ST↗

Demonstration Experiments

Adaptive experiments are used extensively in online platforms, healthcare and biotechnology, and the social sciences. Often, the primary goal is not to precisely estimate a treatment effect but to demonstrate that at least one candidate intervention yields a positive effect, for some subpopulation and on some measured outcome. We formalize this objective as testing the global null in a threshold bandit framework, and develop two inference procedures that are valid under general adaptive sampling: one that pools information across promising arms, and one based on time-uniform multiple testing of individual arm means. To support the latter, we establish a moderate-deviations principle for the sequential $t$-statistic, justifying asymptotic confidence sequences in settings where the number of arms is large relative to the sample size. To illustrate how adaptive designs can target the proposed statistics, we recast experimental design as bandit optimization with an arm's reward given by its signal-to-noise ratio, and analyze an allocation rule for which we establish a logarithmic regret bound. We apply the methods in a simulation study of targeting unconditional cash transfer programs.

math.ST↗

Individual Treatment Effect: Prediction Intervals and Sharp Bounds

Individual treatment effect (ITE) is often regarded as the ideal target of inference in causal analyses and has been the focus of several recent studies. In this paper, we describe the intrinsic limits regarding what can be learned concerning ITEs given data from large randomized experiments. We consider when a valid prediction interval for the ITE is informative and when it can be bounded away from zero. The joint distribution over potential outcomes is only partially identified from a randomized trial. Consequently, to be valid, an ITE prediction interval must be valid for all joint distribution consistent with the observed data and hence will in general be wider than that resulting from knowledge of this joint distribution. We characterize prediction intervals in the binary treatment and outcome setting, and extend these insights to models with continuous and ordinal outcomes. We derive sharp bounds on the probability mass function (pmf) of the individual treatment effect (ITE). Finally, we contrast prediction intervals for the ITE and confidence intervals for the average treatment effect (ATE). This also leads to the consideration of Fisher versus Neyman null hypotheses. While confidence intervals for the ATE shrink with increasing sample size due to its status as a population parameter, prediction intervals for the ITE generally do not vanish, leading to scenarios where one may reject the Neyman null yet still find evidence consistent with the Fisher null, highlighting the challenges of individualized decision-making under partial identification.

stat.ME↗

Bounds on the Distribution of a Sum of Two Random Variables: Revisiting a problem of Kolmogorov with application to Individual Treatment Effects

We revisit the following problem, proposed by Kolmogorov: given prescribed marginal distributions $F$ and $G$ for random variables $X,Y$ respectively, characterize the set of compatible distribution functions for the sum $Z=X+Y$. Bounds on the distribution function for $Z$ were first given by Markarov (1982) and Rüschendorf (1982) independently. Frank et al. (1987) provided a solution to the same problem using copula theory. However, though these authors obtain the same bounds, they make different assertions concerning their sharpness. In addition, their solutions leave some open problems in the case when the given marginal distribution functions are discontinuous. These issues have led to some confusion and erroneous statements in subsequent literature, which we correct. Kolmogorov's problem is closely related to inferring possible distributions for individual treatment effects $Y_1 - Y_0$ given the marginal distributions of $Y_1$ and $Y_0$; the latter being identified from a randomized experiment. We use our new insights to sharpen and correct the results due to Fan and Park (2010) concerning individual treatment effects, and to fill some other logical gaps.

math.ST↗

Assumptions and Bounds in the Instrumental Variable Model

In this note we give proofs for results relating to the Instrumental Variable (IV) model with binary response $Y$ and binary treatment $X$, but with an instrument $Z$ with $K$ states. These results were originally stated in Richardson & Robins (2014), "ACE Bounds; SEMS with Equilibrium Conditions," arXiv:1410.0470.

math.ST↗

A Nonparametric Bayes Approach to Online Activity Prediction

Accurately predicting the onset of specific activities within defined timeframes holds significant importance in several applied contexts. In particular, accurate prediction of the number of future users that will be exposed to an intervention is an important piece of information for experimenters running online experiments (A/B tests). In this work, we propose a novel approach to predict the number of users that will be active in a given time period, as well as the temporal trajectory needed to attain a desired user participation threshold. We model user activity using a Bayesian nonparametric approach which allows us to capture the underlying heterogeneity in user engagement. We derive closed-form expressions for the number of new users expected in a given period, and a simple Monte Carlo algorithm targeting the posterior distribution of the number of days needed to attain a desired number of users; the latter is important for experimental planning. We illustrate the performance of our approach via several experiments on synthetic and real world data, in which we show that our novel method outperforms existing competitors.

stat.ME↗

Nested Markov Properties for Acyclic Directed Mixed Graphs

Conditional independence models associated with directed acyclic graphs (DAGs) may be characterized in at least three different ways: via a factorization, the global Markov property (given by the d-separation criterion), and the local Markov property. Marginals of DAG models also imply equality constraints that are not conditional independences; the well-known ``Verma constraint'' is an example. Constraints of this type are used for testing edges, and in a computationally efficient marginalization scheme via variable elimination. We show that equality constraints like the ``Verma constraint'' can be viewed as conditional independences in kernel objects obtained from joint distributions via a fixing operation that generalizes conditioning and marginalization. We use these constraints to define, via ordered local and global Markov properties, and a factorization, a graphical model associated with acyclic directed mixed graphs (ADMGs). We prove that marginal distributions of DAG models lie in this model, and that a set of these constraints given by Tian provides an alternative definition of the model. Finally, we show that the fixing operation used to define the model leads to a particularly simple characterization of identifiable causal effects in hidden variable causal DAG models.

stat.ME↗

Potential Outcome and Decision Theoretic Foundations for Statistical Causality

In a recent paper published in the Journal of Causal Inference, Philip Dawid has described a graphical causal model based on decision diagrams. This article describes how single-world intervention graphs (SWIGs) relate to these diagrams. In this way, a correspondence is established between Dawid's approach and those based on potential outcomes such as Robins' Finest Fully Randomized Causally Interpreted Structured Tree Graphs. In more detail, a reformulation of Dawid's theory is given that is essentially equivalent to his proposal and isomorphic to SWIGs.

math.ST↗

Coherent modeling of longitudinal causal effects on binary outcomes

Analyses of biomedical studies often necessitate modeling longitudinal causal effects. The current focus on personalized medicine and effect heterogeneity makes this task even more challenging. Towards this end, structural nested mean models (SNMMs) are fundamental tools for studying heterogeneous treatment effects in longitudinal studies. However, when outcomes are binary, current methods for estimating multiplicative and additive SNMM parameters suffer from variation dependence between the causal parameters and the non-causal nuisance parameters. This leads to a series of difficulties in interpretation, estimation and computation. These difficulties have hindered the uptake of SNMMs in biomedical practice, where binary outcomes are very common. We solve the variation dependence problem for the binary multiplicative SNMM via a reparametrization of the non-causal nuisance parameters. Our novel nuisance parameters are variation independent of the causal parameters, and hence allow for coherent modeling of heterogeneous effects from longitudinal studies with binary outcomes. Our parametrization also provides a key building block for flexible doubly robust estimation of the causal parameters. Along the way, we prove that an additive SNMM with binary outcomes does not admit a variation independent parametrization, thereby justifying the restriction to multiplicative SNMMs.

stat.ME↗

The m-connecting imset and factorization for ADMG models

Directed acyclic graph (DAG) models have become widely studied and applied in statistics and machine learning -- indeed, their simplicity facilitates efficient procedures for learning and inference. Unfortunately, these models are not closed under marginalization, making them poorly equipped to handle systems with latent confounding. Acyclic directed mixed graph (ADMG) models characterize margins of DAG models, making them far better suited to handle such systems. However, ADMG models have not seen wide-spread use due to their complexity and a shortage of statistical tools for their analysis. In this paper, we introduce the m-connecting imset which provides an alternative representation for the independence models induced by ADMGs. Furthermore, we define the m-connecting factorization criterion for ADMG models, characterized by a single equation, and prove its equivalence to the global Markov property. The m-connecting imset and factorization criterion provide two new statistical tools for learning and inference with ADMG models. We demonstrate the usefulness of these tools by formulating and evaluating a consistent scoring criterion with a closed form solution.

stat.ML↗

Estimation of local treatment effects under the binary instrumental variable model

Instrumental variables are widely used to deal with unmeasured confounding in observational studies and imperfect randomized controlled trials. In these studies, researchers often target the so-called local average treatment effect as it is identifiable under mild conditions. In this paper, we consider estimation of the local average treatment effect under the binary instrumental variable model. We discuss the challenges for causal estimation with a binary outcome, and show that surprisingly, it can be more difficult than the case with a continuous outcome. We propose novel modeling and estimating procedures that improve upon existing proposals in terms of model congeniality, interpretability, robustness or efficiency. Our approach is illustrated via simulation studies and a real data analysis.

stat.ME↗

Multiplicative Effect Modeling: The General Case

Generalized linear models, such as logistic regression, are widely used to model the association between a treatment and a binary outcome as a function of baseline covariates. However, the coefficients of a logistic regression model correspond to log odds ratios, while subject-matter scientists are often interested in relative risks. Although odds ratios are sometimes used to approximate relative risks, this approximation is appropriate only when the outcome of interest is rare for all levels of the covariates. Poisson regressions do measure multiplicative treatment effects including relative risks, but with a binary outcome not all combinations of parameters lead to fitted means that are between zero and one. Enforcing this constraint makes the parameters variation dependent, which is undesirable for modeling, estimation and computation. Focusing on the special case where the treatment is also binary, Richardson2017 propose a novel binomial regression model, that allows direct modeling of the relative risk. The model uses a log odds-product nuisance model leading to variation independent parameter spaces. Building on this we present general approaches to modeling the multiplicative effect of a continuous or categorical treatment on a binary outcome. Monte Carlo simulations demonstrate the desirable performance of our proposed methods. A data analysis further exemplifies our methods.

stat.ME↗

Robust Estimation of Propensity Score Weights via Subclassification

Weighting estimators based on propensity scores are widely used for causal estimation in a variety of contexts, such as observational studies, marginal structural models and interference. They enjoy appealing theoretical properties such as consistency and possible efficiency under correct model specification. However, this theoretical appeal may be diminished in practice by sensitivity to misspecification of the propensity score model. To improve on this, we borrow an idea from an alternative approach to causal effect estimation in observational studies, namely subclassification estimators. It is well known that compared to weighting estimators, subclassification methods are usually more robust to model misspecification. In this paper, we first discuss an intrinsic connection between the seemingly unrelated weighting and subclassification estimators, and then use this connection to construct robust propensity score weights via subclassification. We illustrate this idea by proposing so-called full-classification weights and accompanying estimators for causal effect estimation in observational studies. Our novel estimators are both consistent and robust to model misspecification, thereby combining the strengths of traditional weighting and subclassification estimators for causal effect estimation from observational studies. Numerical studies show that the proposed estimators perform favorably compared to existing methods.

stat.ME↗

Multivariate Counterfactual Systems And Causal Graphical Models

Among Judea Pearl's many contributions to Causality and Statistics, the graphical d-separation} criterion, the do-calculus and the mediation formula stand out. In this chapter we show that d-separation} provides direct insight into an earlier causal model originally described in terms of potential outcomes and event trees. In turn, the resulting synthesis leads to a simplification of the do-calculus that clarifies and separates the underlying concepts, and a simple counterfactual formulation of a complete identification algorithm in causal models with hidden variables.

stat.ME↗

An Interventionist Approach to Mediation Analysis

Judea Pearl's insight that, when errors are assumed independent, the Pure (aka Natural) Direct Effect (PDE) is non-parametrically identified via the Mediation Formula was `path-breaking' in more than one sense! In the same paper Pearl described a thought-experiment as a way to motivate the PDE. Analysis of this experiment led Robins \& Richardson to a novel way of conceptualizing direct effects in terms of interventions on an expanded graph in which treatment is decomposed into multiple separable components. We further develop this novel theory here, showing that it provides a self-contained framework for discussing mediation without reference to cross-world (nested) counterfactuals or interventions on the mediator. The theory preserves the dictum `no causation without manipulation' and makes questions of mediation empirically testable in future Randomized Controlled Trials. Even so, we prove the interventionist and nested counterfactual approaches remain tightly coupled under a Non-Parametric Structural Equation Model except in the presence of a `recanting witness.' In fact, our analysis also leads to a simple sound and complete algorithm for determining identification in the (non-interventionist) theory of path-specific counterfactuals.

stat.ME↗

Chernoff-type Concentration of Empirical Probabilities in Relative Entropy

We study the relative entropy of the empirical probability vector with respect to the true probability vector in multinomial sampling of $k$ categories, which, when multiplied by sample size $n$, is also the log-likelihood ratio statistic. We generalize a recent result and show that the moment generating function of the statistic is bounded by a polynomial of degree $n$ on the unit interval, uniformly over all true probability vectors. We characterize the family of polynomials indexed by $(k,n)$ and obtain explicit formulae. Consequently, we develop Chernoff-type tail bounds, including a closed-form version from a large sample expansion of the bound minimizer. Our bound dominates the classic method-of-types bound and is competitive with the state of the art. We demonstrate with an application to estimating the proportion of unseen butterflies.

math.ST↗

Discussion of 'Estimating time-varying causal excursion effect in mobile health with binary outcomes' by T. Qian et al

We discuss the recent paper on "excursion effect" by T. Qian et al. (2020). We show that the methods presented have close relationships to others in the literature, in particular to a series of papers by Robins, Hernán and collaborators on analyzing observational studies as a series of randomized trials. There is also a close relationship to the history-restricted and the history-adjusted marginal structural models (MSM). Important differences and their methodological implications are clarified. We also demonstrate that the excursion effect can depend on the design and discuss its suitability for modifying the treatment protocol.

stat.ME↗