SearcharxivSearch

arXiv subjects

Louis Cammarata

Publications and source records attributed to Louis Cammarata.

3 recordsLinked to original sources

Active Learning for Optimal Intervention Design in Causal Models

Sequential experimental design to discover interventions that achieve a desired outcome is a key problem in various domains including science, engineering and public policy. When the space of possible interventions is large, making an exhaustive search infeasible, experimental design strategies are needed. In this context, encoding the causal relationships between the variables, and thus the effect of interventions on the system, is critical for identifying desirable interventions more efficiently. Here, we develop a causal active learning strategy to identify interventions that are optimal, as measured by the discrepancy between the post-interventional mean of the distribution and a desired target mean. The approach employs a Bayesian update for the causal model and prioritizes interventions using a carefully designed, causally informed acquisition function. This acquisition function is evaluated in closed form, allowing for fast optimization. The resulting algorithms are theoretically grounded with information-theoretic bounds and provable consistency results for linear causal models with known causal graph. We apply our approach to both synthetic data and single-cell transcriptomic data from Perturb-CITE-seq experiments to identify optimal perturbations that induce a specific cell state transition. The causally informed acquisition function generally outperforms existing criteria allowing for optimal intervention design with fewer but carefully selected samples.

cs.LG

Power Enhancement and Phase Transitions for Global Testing of the Mixed Membership Stochastic Block Model

The mixed-membership stochastic block model (MMSBM) is a common model for social networks. Given an $n$-node symmetric network generated from a $K$-community MMSBM, we would like to test $K=1$ versus $K>1$. We first study the degree-based $χ^2$ test and the orthodox Signed Quadrilateral (oSQ) test. These two statistics estimate an order-2 polynomial and an order-4 polynomial of a "signal" matrix, respectively. We derive the asymptotic null distribution and power for both tests. However, for each test, there exists a parameter regime where its power is unsatisfactory. It motivates us to propose a power enhancement (PE) test to combine the strengths of both tests. We show that the PE test has a tractable null distribution and improves the power of both tests. To assess the optimality of PE, we consider a randomized setting, where the $n$ membership vectors are independently drawn from a distribution on the standard simplex. We show that the success of global testing is governed by a quantity $β_n(K,P,h)$, which depends on the community structure matrix $P$ and the mean vector $h$ of memberships. For each given $(K, P, h)$, a test is called $\textit{ optimal}$ if it distinguishes two hypotheses when $β_n(K, P,h)\to\infty$. A test is called $\textit{optimally adaptive}$ if it is optimal for all $(K, P, h)$. We show that the PE test is optimally adaptive, while many existing tests are only optimal for some particular $(K, P, h)$, hence, not optimally adaptive.

math.ST

Causal Network Models of SARS-CoV-2 Expression and Aging to Identify Candidates for Drug Repurposing

Given the severity of the SARS-CoV-2 pandemic, a major challenge is to rapidly repurpose existing approved drugs for clinical interventions. While a number of data-driven and experimental approaches have been suggested in the context of drug repurposing, a platform that systematically integrates available transcriptomic, proteomic and structural data is missing. More importantly, given that SARS-CoV-2 pathogenicity is highly age-dependent, it is critical to integrate aging signatures into drug discovery platforms. We here take advantage of large-scale transcriptional drug screens combined with RNA-seq data of the lung epithelium with SARS-CoV-2 infection as well as the aging lung. To identify robust druggable protein targets, we propose a principled causal framework that makes use of multiple data modalities. Our analysis highlights the importance of serine/threonine and tyrosine kinases as potential targets that intersect the SARS-CoV-2 and aging pathways. By integrating transcriptomic, proteomic and structural data that is available for many diseases, our drug discovery platform is broadly applicable. Rigorous in vitro experiments as well as clinical trials are needed to validate the identified candidate drugs.

q-bio.MN