SearcharxivSearch

arXiv subjects

Siyu Heng

Publications and source records attributed to Siyu Heng.

At least 19 recordsLinked to original sources

Spectroscopic fingerprints of a ferroaxial charge density wave

Unconventional charge density waves (CDWs) with complex order parameters can host exotic collective modes and non-trivial topologies. They have emerged as a new frontier in the study of quantum matter. Recent experiments on rare-earth tritellurides have reported evidence for a ferroaxial CDW through the detection of characteristic Raman modes. This phase, often regarded as a hidden order, has been recognized to arise from the coupling between charge and orbital degrees of freedom in these materials. Yet, spectroscopic insight into its underlying electronic structure and the explicit form of its order parameter symmetry has remained elusive. Here, we present results from linearly polarized angle-resolved photoemission spectroscopy (ARPES) and scanning tunneling microscopy (STM) measurements of the CDW phase in LaTe$_3$. Our ARPES measurements reveal a complex landscape of spectral gaps across the reconstructed Fermi surface, while our STM-based quasiparticle interference (QPI) mapping, enhanced through the selective deposition of atomic scattering centers, directly reveals an inter-orbital CDW with mixed $p_x$-$p_z$ orbital character. The detailed analysis of the QPI characteristics in terms of the order parameter symmetry within the orbital subspace of the Fermi surface suggests a mixed CDW phase with substantial ferroaxial component, which breaks all vertical mirror symmetries. More broadly, our work establishes a powerful spectroscopic pathway, based on scattering off individual atoms, for identifying and characterizing hidden, multi-component electronic orders in quantum materials using STM and ARPES measurements.

cond-mat.str-el

Propensity Score Propagation: A General Framework for Design-Based Inference with Unknown Propensity Scores

Design-based inference, also known as randomization-based or finite-population inference, provides a principled framework for trustworthy statistical inference. It attributes randomness solely to the design mechanism, such as treatment assignment, survey sampling, or missingness, without imposing super-population distributional or modeling assumptions on the outcome data. From the seminal work of Fisher and Neyman to its recent resurgence, design-based inference has played a central role in causal inference, survey sampling, and missing data analysis. However, its use in many modern applications has been limited by a fundamental obstacle: existing design-based inference theory typically assumes that propensity scores (i.e., design probabilities) are known, whereas they are usually unknown in observational studies, real-world surveys, and missing data problems. We propose propensity score propagation, a general framework for valid design-based inference with unknown propensity scores. The framework uses a regeneration-and-union procedure to propagate uncertainty from propensity score estimation into downstream design-based inference, without introducing super-population assumptions about the outcomes. It accommodates both parametric and nonparametric propensity score settings, integrates seamlessly with existing design-based methods developed for known propensity scores, and applies broadly across design-based problems. Theoretical and simulation results show that the proposed framework achieves nominal coverage, even when existing approaches exhibit substantial under-coverage.

stat.ME

A Universal Framework for Factorial Matched Observational Studies with General Treatment Types: Design, Analysis, and Applications

Matching is one of the most widely used causal inference frameworks in observational studies. However, all the existing matching-based causal inference methods are designed for either a single treatment with general treatment types (e.g., binary, ordinal, or continuous) or factorial (multiple) treatments with binary treatments only. To our knowledge, no existing matching-based causal methods can handle factorial treatments with general treatment types. This critical gap substantially hinders the applicability of matching in many real-world problems, in which there are often multiple, potentially non-binary (e.g., continuous) treatment components. To address this critical gap, this work develops a universal framework for the design and analysis of factorial matched observational studies with general treatment types (e.g., binary, ordinal, or continuous). We first propose a two-stage non-bipartite matching algorithm that constructs matched sets of units with similar covariates but distinct combinations of treatment doses, thereby enabling valid estimation of both main and interaction effects. We then introduce a new class of generalized factorial Neyman-type estimands that provide model-free, finite-population-valid definitions of marginal and interaction causal effects under factorial treatments with general treatment types. Randomization-based Fisher-type and Neyman-type inference procedures are developed, including unbiased estimators, asymptotically valid variance estimators, and variance adjustments incorporating covariate information for improved efficiency. Finally, we illustrate the proposed framework through a county-level application that evaluates the causal impacts of work- and non-work-trip reductions (social distancing practices) on COVID-19-related and drug-related outcomes during the COVID-19 pandemic in the United States.

stat.ME

A Non-Bipartite Matching Framework for Difference-in-Differences with General Treatment Types

Difference-in-differences (DID) is one of the most widely used causal inference frameworks in observational studies. However, most existing DID methods are designed for binary treatments and cannot be readily applied to non-binary treatment settings. Although recent work has begun to extend DID to non-binary (e.g., continuous) treatments, these approaches typically require strong additional assumptions, including parametric outcome models or the presence of idealized comparison units with (nearly) static treatment levels over time (commonly called ``stayers'' or ``quasi-stayers''). In this technical note, we introduce a new non-bipartite matching framework for DID that naturally accommodates general treatment types (e.g., binary, ordinal, or continuous). Our framework makes three main contributions. First, we develop an optimal non-bipartite matching design for DID that jointly balances baseline covariates across comparable units (reducing bias) and maximizes contrasts in treatment trajectories over time (improving efficiency). Second, we establish a post-matching randomization condition, the design-based counterpart to the traditional parallel-trends assumption, which enables valid design-based inference. Third, we introduce the sample average DID ratio, a finite-population-valid and fully nonparametric causal estimand applicable to arbitrary treatment types. Our design-based approach that preserves the full treatment-dose information, avoids parametric assumptions, does not rely on the existence of stayers or quasi-stayers, and operates entirely within a finite-population framework, without appealing to hypothetical super-populations or outcome distributions.

stat.ME

Determining vaccine responders in the presence of baseline immunity using single-cell assays and paired control samples

A key objective in vaccine studies is to evaluate vaccine-induced immunogenicity and determine whether participants have mounted a response to the vaccine. Cellular immune responses are essential for assessing vaccine-induced immunogenicity, and single-cell assays, such as intracellular cytokine staining (ICS) are commonly employed to profile individual immune cell phenotypes and the cytokines they produce after stimulation. In this article, we introduce a novel statistical framework for identifying vaccine responders using ICS data collected before and after vaccination. This framework incorporates paired control data to account for potential unintended variations between assay runs, such as batch effects, that could lead to misclassification of participants as vaccine responders. To formally integrate paired control data for accounting for assay variation across different time points (i.e., before and after vaccination), our proposed framework calculates and reports two p-values, both adjusting for paired control data but in distinct ways: (i) the maximally adjusted p-value, which applies the most conservative adjustment to the unadjusted p-value, ensuring validity over all plausible batch effects consistent with the paired control samples' data, and (ii) the minimally adjusted p-value, which imposes only the minimal adjustment to the unadjusted p-value, such that the adjusted p-value cannot be falsified by the paired control samples' data. We apply this framework to analyze ICS data collected at baseline and 4 weeks post-vaccination from the COVID-19 Prevention Network (CoVPN) 3008 study. Our analysis helps address two clinical questions: 1) which participants exhibited evidence of an incident Omicron infection, and 2) which participants showed vaccine-induced T cell responses against the Omicron BA.4/5 Spike protein.

stat.ME

Re-evaluating the impact of reduced malaria prevalence on birthweight in sub-Saharan Africa: A pair-of-pairs study via two-stage bipartite and non-bipartite matching

According to the WHO, in 2021, about 32% of pregnant women in sub-Saharan Africa were infected with malaria during pregnancy. Malaria infection during pregnancy can cause various adverse birth outcomes such as low birthweight. Over the past two decades, while some sub-Saharan African areas have experienced a large reduction in malaria prevalence due to improved malaria control and treatments, others have observed little change. Individual-level interventional studies have shown that preventing malaria infection during pregnancy can improve birth outcomes such as birthweight; however, it is still unclear whether natural reductions in malaria prevalence may help improve community-level birth outcomes. We conduct an observational study using 203,141 children's records in 18 sub-Saharan African countries from 2000 to 2018. Using heterogeneity of changes in malaria prevalence, we propose and apply a novel pair-of-pairs design via two-stage bipartite and non-bipartite matching to conduct a difference-in-differences study with a continuous measure of malaria prevalence, namely the Plasmodium falciparum parasite rate among children aged 2 to 10 ($\text{PfPR}_{2-10}$). The proposed novel statistical methodology allows us to apply difference-in-differences without dichotomizing $\text{PfPR}_{2-10}$, which can substantially increase the effective sample size, improve covariate balance, and facilitate the dose-response relationship during analysis. Our outcome analysis finds that among the pairs of clusters we study, the largest reduction in $\text{PfPR}_{2-10}$ over early and late years is estimated to increase the average birthweight by 98.899 grams (95% CI: $[39.002, 158.796]$), which is associated with reduced risks of several adverse birth or life-course outcomes. The proposed novel statistical methodology can be replicated in many other disease areas.

stat.AP

Bridging the Gap Between Design and Analysis: Randomization Inference and Sensitivity Analysis for Matched Observational Studies with Treatment Doses

Matching is a commonly used causal inference study design in observational studies. Through matching on measured confounders between different treatment groups, valid randomization inferences can be conducted under the no unmeasured confounding assumption, and sensitivity analysis can be further performed to assess sensitivity of randomization inference results to potential unmeasured confounding. However, for many common matching designs, there is still a lack of valid downstream randomization inference and sensitivity analysis approaches. Specifically, in matched observational studies with treatment doses (e.g., continuous or ordinal treatments), with the exception of some special cases such as pair matching, there is no existing randomization inference or sensitivity analysis approach for studying analogs of the sample average treatment effect (Neyman-type weak nulls), and no existing valid sensitivity analysis approach for testing the sharp null of no effect for any subject (Fisher's sharp null) when the outcome is non-binary. To fill these gaps, we propose new methods for randomization inference and sensitivity analysis that can work for general matching designs with treatment doses, applicable to general types of outcome variables (e.g., binary, ordinal, or continuous), and cover both Fisher's sharp null and Neyman-type weak nulls. We illustrate our approaches via comprehensive simulation studies and a real-data application.

stat.ME

Bias Mitigation in Matched Observational Studies with Continuous Treatments: Calipered Non-Bipartite Matching and Bias-Corrected Estimation and Inference

In matched observational studies with continuous treatments, individuals with different treatment doses but the same or similar covariate values are paired for causal inference. While inexact covariate matching (i.e., covariate imbalance after matching) is common in practice, previous matched studies with continuous treatments have often overlooked this issue as long as post-matching covariate balance meets certain criteria. Through re-analyzing a matched observational study on the social distancing effect on COVID-19 case counts, we show that this routine practice can introduce severe bias for causal inference. Motivated by this finding, we propose a general framework for mitigating bias due to inexact matching in matched observational studies with continuous treatments, covering the matching, estimation, and inference stages. In the matching stage, we propose a carefully designed caliper that incorporates both covariate and treatment dose information to improve matching for downstream treatment effect estimation and inference. For the estimation and inference, we introduce a bias-corrected Neyman estimator paired with a corresponding bias-corrected variance estimator. The effectiveness of our proposed framework is demonstrated through numerical studies and a re-analysis of the aforementioned observational study on the effect of social distancing on COVID-19 case counts. An open-source R package for implementing our framework has also been developed.

stat.ME

DELTA: Dual Consistency Delving with Topological Uncertainty for Active Graph Domain Adaptation

Graph domain adaptation has recently enabled knowledge transfer across different graphs. However, without the semantic information on target graphs, the performance on target graphs is still far from satisfactory. To address the issue, we study the problem of active graph domain adaptation, which selects a small quantitative of informative nodes on the target graph for extra annotation. This problem is highly challenging due to the complicated topological relationships and the distribution discrepancy across graphs. In this paper, we propose a novel approach named Dual Consistency Delving with Topological Uncertainty (DELTA) for active graph domain adaptation. Our DELTA consists of an edge-oriented graph subnetwork and a path-oriented graph subnetwork, which can explore topological semantics from complementary perspectives. In particular, our edge-oriented graph subnetwork utilizes the message passing mechanism to learn neighborhood information, while our path-oriented graph subnetwork explores high-order relationships from sub-structures. To jointly learn from two subnetworks, we roughly select informative candidate nodes with the consideration of consistency across two subnetworks. Then, we aggregate local semantics from its K-hop subgraph based on node degrees for topological uncertainty estimation. To overcome potential distribution shifts, we compare target nodes and their corresponding source nodes for discrepancy scores as an additional component for fine selection. Extensive experiments on benchmark datasets demonstrate that DELTA outperforms various state-of-the-art approaches. The code implementation of DELTA is available at https://github.com/goose315/DELTA.

cs.LG

A Comprehensive Graph Pooling Benchmark: Effectiveness, Robustness and Generalizability

Graph pooling has gained attention for its ability to obtain effective node and graph representations for various downstream tasks. Despite the recent surge in graph pooling approaches, there is a lack of standardized experimental settings and fair benchmarks to evaluate their performance. To address this issue, we have constructed a comprehensive benchmark that includes 17 graph pooling methods and 28 different graph datasets. This benchmark systematically assesses the performance of graph pooling methods in three dimensions, i.e., effectiveness, robustness, and generalizability. We first evaluate the performance of these graph pooling approaches across different tasks including graph classification, graph regression and node classification. Then, we investigate their performance under potential noise attacks and out-of-distribution shifts in real-world scenarios. We also involve detailed efficiency analysis, backbone analysis, parameter analysis and visualization to provide more evidence. Extensive experiments validate the strong capability and applicability of graph pooling approaches in various scenarios, which can provide valuable insights and guidance for deep geometric learning research. The source code of our benchmark is available at https://github.com/goose315/Graph_Pooling_Benchmark.

cs.LG

Towards Robust Matched Observational Studies with General Treatment Types: Consistency, Efficiency, and Adaptivity

To ensure reliable causal conclusions from observational studies, researchers routinely conduct sensitivity analysis to assess robustness to unmeasured confounding. In matched observational studies (one of the most popular observational study designs), two foundational concepts, design sensitivity and Bahadur-Rosenbaum efficiency, are used to quantify the robustness of test statistics and study designs in sensitivity analyses. Unfortunately, these measures of robustness are unavailable for non-binary treatments (e.g., continuous treatments) and consequently, prevailing recommendations about robust tests may be misleading. In this work, we provide a unified framework to quantify the robustness of test statistics and study designs for any treatment type. We first present a negative result about a popular, ad-hoc approach based on dichotomizing the treatment variable. Next, we consider two complementary parameterizations of the sensitivity parameter that apply to arbitrary treatment types and establish a one-to-one correspondence between them. Using these equivalent parameterizations, we generalize design sensitivity and Bahadur-Rosenbaum efficiency and derive unified formulas for both quantities under general treatment types. We also propose a general data-adaptive approach that combines candidate test statistics to enhance robustness against unmeasured confounding. Our results yield new insights into the robustness of tests and study designs in matched observational studies with non-binary treatments.

stat.ME

Sensitivity Analysis for Matched Observational Studies with Continuous Exposures and Binary Outcomes

Matching is one of the most widely used study designs for adjusting for measured confounders in observational studies. However, unmeasured confounding may exist and cannot be removed by matching. Therefore, a sensitivity analysis is typically needed to assess a causal conclusion's sensitivity to unmeasured confounding. Sensitivity analysis frameworks for binary exposures have been well-established for various matching designs and are commonly used in various studies. However, unlike the binary exposure case, there still lacks valid and general sensitivity analysis methods for continuous exposures, except in some special cases such as pair matching. To fill this gap in the binary outcome case, we develop a sensitivity analysis framework for general matching designs with continuous exposures and binary outcomes. First, we use probabilistic lattice theory to show our sensitivity analysis approach is finite-population-exact under Fisher's sharp null. Second, we prove a novel design sensitivity formula as a powerful tool for asymptotically evaluating the performance of our sensitivity analysis approach. Third, to allow effect heterogeneity with binary outcomes, we introduce a framework for conducting asymptotically exact inference and sensitivity analysis on generalized attributable effects with binary outcomes via mixed-integer programming. Fourth, for the continuous outcomes case, we show that conducting an asymptotically exact sensitivity analysis in matched observational studies when both the exposures and outcomes are continuous is generally NP-hard, except in some special cases such as pair matching. As a real data application, we apply our new methods to study the effect of early-life lead exposure on juvenile delinquency. We also develop a publicly available R package for implementation of the methods in this work.

stat.ME

Reconciling Overt Bias and Hidden Bias in Sensitivity Analysis for Matched Observational Studies

Matching is one of the most widely used causal inference designs in observational studies, but post-matching confounding bias remains a critical concern. This bias includes overt bias from inexact matching on measured confounders and hidden bias from unmeasured confounders. Researchers routinely apply the famous Rosenbaum-type sensitivity analysis after matching to assess the impact of these biases on causal conclusions. In this work, we show that this approach is often conservative and may overstate sensitivity to confounding bias because the classical solution to the Rosenbaum sensitivity model may allocate hypothetical hidden bias in ways that contradict the overt bias observed in the matched dataset. To address this problem, we propose a new approach to Rosenbaum-type sensitivity analysis by ensuring compatibility between hidden and overt biases. Our approach does not need to add any additional assumptions (beyond mild regularity conditions) to Rosenbaum-type sensitivity analysis, and can produce uniformly more informative sensitivity analysis results than the conventional Rosenbaum-type sensitivity analysis. Computationally, our approach can be solved efficiently via iterative convex programming. Extensive simulations and a real data application demonstrate substantial gains in statistical power of sensitivity analysis. Importantly, our approach can also be applied to many other sensitivity analysis frameworks.

stat.ME

Design-Based Causal Inference with Missing Outcomes: Missingness Mechanisms, Imputation-Assisted Randomization Tests, and Covariate Adjustment

Design-based causal inference, also known as randomization-based or finite-population causal inference, is one of the most widely used causal inference frameworks, largely due to the merit that its validity can be guaranteed by study design (e.g., randomized experiments) and does not require assuming specific outcome-generating distributions or super-population models. Despite its advantages, design-based causal inference can still suffer from other issues, among which outcome missingness is a prevalent and significant challenge. This work systematically studies the outcome missingness problem in design-based causal inference. First, we propose a general and flexible outcome missingness mechanism that can facilitate finite-population-exact randomization tests of no treatment effect. Second, under this general missingness mechanism, we propose a general framework called ``imputation and re-imputation" for conducting randomization tests in design-based causal inference with missing outcomes. We prove that our framework can still ensure finite-population-exact type-I error rate control even when the imputation model was misspecified or when unobserved covariates or interference exist in the missingness mechanism. Third, we extend our framework to conduct covariate adjustment in randomization tests and construct finite-population-valid confidence regions with missing outcomes. Our framework is evaluated via extensive simulation studies and applied to a large-scale randomized experiment.

stat.ME

Randomization-Based Inference for Average Treatment Effects in Inexactly Matched Observational Studies

Matching is a widely used causal inference design that aims to approximate a randomized experiment using observational data by forming matched sets of treated and control units based on similarities in their covariates. Ideally, treated units are exactly matched with controls on these covariates, enabling randomization-based inference for treatment effects as in a randomized experiment, under the assumption of no unobserved covariates. However, inexact matching often occurs, leading to residual covariate imbalance after matching. Previous matched studies have typically overlooked this issue and relied on conventional randomization-based inference, assuming that some covariate balance criteria are met. Recent research, however, has shown that this approach can introduce significant bias and proposed methods to correct for bias arising from inexact matching in randomization-based inference. These methods, however, are primarily focused on the constant treatment effect and its extensions (i.e., Fisher's sharp null) and do not apply to average treatment effects (i.e., Neyman's weak null). To address this gap, we introduce a new method--inverse post-matching probability weighting--for conducting randomization-based inference for average treatment effects under inexact matching. Our theoretical and simulation results indicate that, compared to conventional randomization-based inference methods, our approach significantly reduces bias and improves coverage rates in the presence of inexact matching.

stat.ME

Testing Biased Randomization Assumptions and Quantifying Imperfect Matching and Residual Confounding in Matched Observational Studies

One central goal of design of observational studies is to embed non-experimental data into an approximate randomized controlled trial using statistical matching. Despite empirical researchers' best intention and effort to create high-quality matched samples, residual imbalance due to observed covariates not being well matched often persists. Although statistical tests have been developed to test the randomization assumption and its implications, few provide a means to quantify the level of residual confounding due to observed covariates not being well matched in matched samples. In this article, we develop two generic classes of exact statistical tests for a biased randomization assumption. One important by-product of our testing framework is a quantity called residual sensitivity value (RSV), which provides a means to quantify the level of residual confounding due to imperfect matching of observed covariates in a matched sample. We advocate taking into account RSV in the downstream primary analysis. The proposed methodology is illustrated by re-examining a famous observational study concerning the effect of right heart catheterization (RHC) in the initial care of critically ill patients. Code implementing the method can be found in the supplementary materials.

stat.ME

Sensitivity Analysis for Binary Outcome Misclassification in Randomization Tests via Integer Programming

Conducting a randomization test is a common method for testing causal null hypotheses in randomized experiments. The popularity of randomization tests is largely because their statistical validity only depends on the randomization design, and no distributional or modeling assumption on the outcome variable is needed. However, randomization tests may still suffer from other sources of bias, among which outcome misclassification is a significant one. We propose a model-free and finite-population sensitivity analysis approach for binary outcome misclassification in randomization tests. A central quantity in our framework is ``warning accuracy," defined as the threshold such that a randomization test result based on the measured outcomes may differ from that based on the true outcomes if the outcome measurement accuracy did not surpass that threshold. We show how learning the warning accuracy and related concepts can amplify analyses of randomization tests subject to outcome misclassification without adding additional assumptions. We show that the warning accuracy can be computed efficiently for large data sets by adaptively reformulating a large-scale integer program with respect to the randomization design. We apply the proposed approach to the Prostate Cancer Prevention Trial (PCPT). We also developed an open-source R package for implementation of our approach.

stat.ME

Re-Evaluating Strengthened-IV Designs: Asymptotic Efficiency, Bias Formula, and the Validity and Power of Sensitivity Analyses

Instrumental variables (IVs) are extensively used to estimate treatment effects when the treatment and outcome are confounded by unmeasured confounders; however, weak IVs are often encountered in empirical studies and may cause problems. Many studies have considered building a stronger IV from the original, possibly weak, IV in the design stage of a matched study at the cost of not using some of the samples in the analysis. It is widely accepted that strengthening an IV tends to render nonparametric tests more powerful and will increase the power of sensitivity analyses in large samples. In this article, we re-evaluate this conventional wisdom to bring new insights into this topic. We consider matched observational studies from three perspectives. First, we evaluate the trade-off between IV strength and sample size on nonparametric tests assuming the IV is valid and exhibit conditions under which strengthening an IV increases power and conversely conditions under which it decreases power. Second, we derive a necessary condition for a valid sensitivity analysis model with continuous doses. We show that the $Γ$ sensitivity analysis model, which has been previously used to come to the conclusion that strengthening an IV increases the power of sensitivity analyses in large samples, does not apply to the continuous IV setting and thus this previously reached conclusion may be invalid. Third, we quantify the bias of the Wald estimator with a possibly invalid IV under an oracle and leverage it to develop a valid sensitivity analysis framework; under this framework, we show that strengthening an IV may amplify or mitigate the bias of the estimator, and may or may not increase the power of sensitivity analyses. We also discuss how to better adjust for the observed covariates when building an IV in matched studies.

stat.ME