SearcharxivSearch

arXiv subjects

Jichun Xie

Publications and source records attributed to Jichun Xie.

8 recordsLinked to original sources

HaploPerturb: Low-rank copula construction of haplotype perturbations improves sequence-to-function analysis of Alzheimer's disease loci

Sequence-to-function models predict molecular phenotypes from complete sequence windows. At genome-wide association study loci, however, the prevailing design perturbs only the lead variant on the reference genome, even though the lead is often correlated with nearby variants through linkage disequilibrium. This single-variant perturbation implicitly fixes all linked alleles at their reference-genome states and may therefore create an uncommon or unobserved population haplotype. We study this input-construction problem at 38 Alzheimer's disease loci. We introduce HaploPerturb, which uses phased ROS/MAP genotypes or the European 1000 Genomes panel to fit a fixed-margin latent Gaussian factor model and rank partner configurations conditional on each lead allele. The leading public-panel construction agrees with the donor-panel construction at all loci under a strict linkage-disequilibrium threshold and 36 of 37 loci under a broader threshold after restricting to shared partners. Known-truth simulations show exact recovery of the dominant configuration under strong linkage disequilibrium and expose persistent residual correlation under a misspecified one-factor model. In an AlphaGenome benchmark against cell-type-specific ROS/MAP eQTLs, broader-set public-panel haplotypes yield microglial enrichment of 2.07 (95\% whole-locus bootstrap percentile interval 1.48--3.60), compared with 1.43 (0.69--2.29) for a lead-only edit. Empirical-mode and LD-sign backgrounds yield 2.20 (1.60--3.64), with no detectable advantage or loss relative to the HaploPerturb top configuration. Thus population-informed sequence construction matters in this application, while the choice among reasonable leading haplotype rules is less consequential than the choice between a haplotype and a lead-only reference background.

stat.AP

Finding Non-Redundant Simpson's Paradox from Multidimensional Data

Simpson's paradox, a long-standing statistical phenomenon, describes the reversal of an observed association when data are disaggregated into sub-populations. It has critical implications across statistics, epidemiology, economics, and causal inference. Existing methods for detecting Simpson's paradox overlook a key issue: many paradoxes are redundant, arising from equivalent selections of data subsets, identical partitioning of sub-populations, and correlated outcome variables, which obscure essential patterns and inflate computational cost. In this paper, we present the first framework for discovering non-redundant Simpson's paradoxes. We formalize three types of redundancy - sibling child, separator, and statistic equivalence - and show that redundancy forms an equivalence relation. Leveraging this insight, we propose a concise representation framework for systematically organizing redundant paradoxes and design efficient algorithms that integrate depth-first materialization of the base table with redundancy-aware paradox discovery. Experiments on real-world datasets and synthetic benchmarks show that redundant paradoxes are widespread, on some real datasets constituting over 40% of all paradoxes, while our algorithms scale to millions of records, reduce run time by up to 60%, and discover paradoxes that are structurally robust under data perturbation. These results demonstrate that Simpson's paradoxes can be efficiently identified, concisely summarized, and meaningfully interpreted in large multidimensional datasets.

cs.DB

DART2: a robust multiple testing method to smartly leverage helpful or misleading ancillary information

In many applications of multiple testing, ancillary information is available, reflecting the hypothesis null or alternative status. Several methods have been developed to leverage this ancillary information to enhance testing power, typically requiring the ancillary information is helpful enough to ensure favorable performance. In this paper, we develop a robust and effective distance-assisted multiple testing procedure named DART2, designed to be powerful and robust regardless of the quality of ancillary information. When the ancillary information is helpful, DART2 can asymptotically control FDR while improving power; otherwise, DART2 can still control FDR and maintain power at least as high as ignoring the ancillary information. We demonstrated DART2's superior performance compared to existing methods through numerical studies under various settings. In addition, DART2 has been applied to a gene association study where we have shown its superior accuracy and robustness under two different types of ancillary information.

stat.ML

Distance Assisted Recursive Testing

In many applications, a large number of features are collected with the goal to identify a few important ones. Sometimes, these features lie in a metric space with a known distance matrix, which partially reflects their co-importance pattern. Proper use of the distance matrix will boost the power of identifying important features. Hence, we develop a new multiple testing framework named the Distance Assisted Recursive Testing (DART). DART has two stages. In stage 1, we transform the distance matrix into an aggregation tree, where each node represents a set of features. In stage 2, based on the aggregation tree, we set up dynamic node hypotheses and perform multiple testing on the tree. All rejections are mapped back to the features. Under mild assumptions, the false discovery proportion of DART converges to the desired level in high probability converging to one. We illustrate by theory and simulations that DART has superior performance under various models compared to the existing methods. We applied DART to a clinical trial in the allogeneic stem cell transplantation study to identify the gut microbiota whose abundance will be impacted by the after-transplant care.

stat.ME

TEAM: A Multiple Testing Algorithm on the Aggregation Tree for Flow Cytometry Analysis

In immunology studies, flow cytometry is a commonly used multivariate single-cell assay. One key goal in flow cytometry analysis is to pinpoint the immune cells responsive to certain stimuli. Statistically, this problem can be translated into comparing two protein expression probability density functions (PDFs) before and after the stimulus; the goal is to pinpoint the regions where these two pdfs differ. In this paper, we model this comparison as a multiple testing problem. First, we partition the sample space into small bins. In each bin we form a hypothesis to test the existence of differential pdfs. Second, we develop a novel multiple testing method, called TEAM (Testing on the Aggregation tree Method), to identify those bins that harbor differential pdfs while controlling the false discovery rate (FDR) under the desired level. TEAM embeds the testing procedure into an aggregation tree to test from fine- to coarse-resolution. The procedure achieves the statistical goal of pinpointing differential pdfs to the smallest possible regions. TEAM is computationally efficient, capable of analyzing large flow cytometry data sets in much shorter time compared with competing methods. We applied TEAM and competing methods on a flow cytometry data set to identify T cells responsive to the cytomeglovirus (CMV)-pp65 antigen stimulation. TEAM successfully identified the monofunctional, bifunctional, and polyfunctional T cells while the competing methods either did not finish in a reasonable time frame or provided less interpretable results. Numerical simulations and theoretical justifications demonstrate that TEAM has asymptotically valid, powerful, and robust performance. Overall, TEAM is a computationally efficient and statistically powerful algorithm that can yield meaningful biological insights in flow cytometry studies.

stat.ME

False Discovery Rate Control for High-Dimensional Networks of Quantile Associations Conditioning on Covariates

Motivated by the gene co-expression pattern analysis, we propose a novel sample quantile-based contingency (squac) statistic to infer quantile associations conditioning on covariates. It features enhanced flexibility in handling variables with both arbitrary distributions and complex association patterns conditioning on covariates. We first derive its asymptotic null distribution, and then develop a multiple testing procedure based on squac to simultaneously test the independence between one pair of variables conditioning on covariates for all $p(p-1)/2$ pairs. Here, $p$ is the length of the outcomes and could exceed the sample size. The testing procedure does not require resampling or perturbation, and thus is computationally efficient. We prove by theory and numerical experiments that this testing method asymptotically controls the false discovery rate (\FDR). It outperforms all alternative methods when the complex association panterns exist. Applied to a gastric cancer data, this testing method successfully inferred the gene co-expression networks of early and late stage patients. It identified more changes in the networks which are associated with cancer survivals. We extend our method to the case that both the length of the outcomes and the length of covariates exceed the sample size, and show that the asymptotic theory still holds.

stat.ME

High Dimensional Tests for Functional Networks of Brain Anatomic Regions

There has been increasing interests in learning resting-state brain functional connectivity of autism disorders using functional magnetic resonance imaging (fMRI) data. The data in a standard brain template consist of over 200,000 voxel specific time series for each single subject. Such an ultra-high dimensionality of data makes the voxel-level functional connectivity analysis (involving four billion voxel pairs) lack of power and extremely inefficient. In this work, we introduce a new framework to identify functional brain network at brain anatomic region-level for each individual. We propose two pairwise tests to detect region dependence, and one multiple testing procedure to identify global structures of the network. The limiting null distributions of the test statistics are derived. It is also shown that the tests are rate optimal when the alternative networks are sparse. The numerical studies show the proposed tests are valid and powerful. We apply our method to a resting-state fMRI study on autism and identify patient-unique and control-unique hub regions. These findings are consistent with autism clinical symptoms.

stat.ME

PenPC: A Two-step Approach to Estimate the Skeletons of High Dimensional Directed Acyclic Graphs

Estimation of the skeleton of a directed acyclic graph (DAG) is of great importance for understanding the underlying DAG and causaleffects can be assessed from the skeleton when the DAG is notidentifiable. We propose a novel method named PenPC toestimate the skeleton of a high-dimensional DAG by a two-stepapproach. We first estimate the non-zero entries of a concentrationmatrix using penalized regression, and then fix the differencebetween the concentration matrix and the skeleton by evaluating aset of conditional independence hypotheses. For high dimensionalproblems where the number of vertices $p$ is in polynomial orexponential scale of sample size $n$, we study the asymptoticproperty of PenPC on two types of graphs: traditionalrandom graphs where all the vertices have the same expected numberof neighbors, and scale-free graphs where a few vertices may have alarge number of neighbors. As illustrated by extensive simulationsand applications on gene expression data of cancer patients, PenPChas higher sensitivity and specificity than the standard-of-the-artmethod, the PC-stable algorithm.

stat.ME