Searcharxiv⌕ Search

arXiv subjects

Aparajithan Venkateswaran

Publications and source records attributed to Aparajithan Venkateswaran.

3 recordsLinked to original sources

Towards Complete Causal Explanation with Expert Knowledge

We study the problem of restricting a Markov equivalence class of maximal ancestral graphs (MAGs) to only those MAGs that contain certain edge marks, which we refer to as expert or orientation knowledge. Such a restriction of the Markov equivalence class can be uniquely represented by a restricted essential ancestral graph. Our contributions are several-fold. First, we prove certain properties for the entire Markov equivalence class including a conjecture from Ali et al. (2009). Second, we present several new sound graphical orientation rules for adding orientation knowledge to an essential ancestral graph. We also show that some orientation rules of Zhang (2008b) are not needed for restricting the Markov equivalence class with orientation knowledge. Third, we provide an algorithm for including this orientation knowledge and show that in certain settings the output of our algorithm is a restricted essential ancestral graph. Finally, outside of the specified settings, we provide an algorithm for checking whether a graph is a restricted essential graph and discuss its runtime. This work can be seen as a generalization of Meek (1995) to settings which allow for latent confounding.

stat.ML↗

Robustly estimating heterogeneity in factorial data using Rashomon Partitions

In both observational data and randomized control trials, researchers select statistical models to articulate how the outcome of interest varies with combinations of observable covariates. Choosing a model that is too simple can obfuscate important heterogeneity in outcomes between covariate groups, while too much complexity risks identifying spurious patterns. In this paper, we propose a novel Bayesian framework for model uncertainty called Rashomon Partition Sets (RPSs). The RPS consists of all models that have posterior density close to the maximum a posteriori (MAP) model. We construct the RPS by enumeration, rather than sampling, which ensures that we explore all models with high evidence in the data, even if they offer dramatically different substantive explanations. We use a l0 prior, which allows the allows us to capture complex heterogeneity without imposing strong assumptions about the associations between effects, showing this prior is minimax optimal from an information-theoretic perspective. We characterize the approximation error of (functions of) parameters computed conditional on being in the RPS relative to the entire posterior. We propose an algorithm to enumerate the RPS from the class of models that are interpretable and unique, then provide bounds on the size of the RPS. We give simulation evidence along with three empirical examples: price effects on charitable giving, heterogeneity in chromosomal structure, and the introduction of microfinance.

stat.ME↗

Feasible contact tracing

Contact tracing is one of the most important tools for preventing the spread of infectious diseases, but as the experience of COVID-19 showed, it is also next-to-impossible to implement when the disease is spreading rapidly. We show how to substantially improve the efficiency of contact tracing by combining standard microeconomic tools that measure heterogeneity in how infectious a sick person is with ideas from machine learning about sequential optimization. Our contributions are twofold. First, we incorporate heterogeneity in individual infectiousness in a multi-armed bandit to establish optimal algorithms. At the heart of this strategy is a focus on learning. In the typical conceptualization of contact tracing, contacts of an infected person are tested to find more infections. Under a learning-first framework, however, contacts of infected persons are tested to ascertain whether the infected person is likely to be a "high infector" and to find additional infections only if it is likely to be highly fruitful. Second, we demonstrate using three administrative contact tracing datasets from India and Pakistan during COVID-19 that this strategy improves efficiency. Using our algorithm, we find 80% of infections with just 40% of contacts while current approaches test twice as many contacts to identify the same number of infections. We further show that a simple strategy that can be easily implemented in the field performs at nearly optimal levels, allowing for, what we call, feasible contact tracing. These results are immediately transferable to contact tracing in any epidemic.

stat.ME↗