SearcharxivSearch

arXiv subjects

Yikang Chen

Publications and source records attributed to Yikang Chen.

8 recordsLinked to original sources

DAG-FM: A Foundation Model for Causal Discovery under Heterogeneous Causal Mechanisms

Causal discovery from observational tabular data remains fundamentally challenging, primarily due to the heterogeneity of underlying causal mechanisms and the high-dimensional combinatorial search space of Directed Acyclic Graphs (DAGs). In this paper, we propose \textbf{DAG-FM}, a novel foundation model architecture that amortizes causal discovery. Unlike direct matrix prediction, DAG-FM decomposes the causal discovery process into two auto-regressive stages using two specialized Transformer-based sub-modules: a leaf-node predictor and a parent-node predictor. To effectively model complex row-column interactions, we adopt a robust tabular interaction block to output feature-wise representations. Crucially, to handle diverse and unknown Functional Causal Model (FCM) assumptions in real-world scenarios, we introduce Mixture-of-Leaf-Experts (MoLE), allowing the model to dynamically route and adapt to identifiable mechanism families. Through an iterative inference algorithm, DAG-FM seamlessly extracts causal orderings and constructs valid DAGs. Extensive experiments demonstrate that DAG-FM achieves state-of-the-art performance on both synthetic benchmarks and complex real-world datasets, significantly outperforming traditional classical algorithms and recent foundation models in both accuracy and scalability.

cs.LG

DCD-PFN: A Decoupling-Aware Foundation Model for Causal Discovery

Causal discovery is critical for understanding complex data-generating mechanisms, yet traditional algorithms often struggle with highly non-linear and noisy systems, or suffer from severe computational bottlenecks. Recent tabular foundation models based on Prior-Data Fitted Networks (PFNs) have demonstrated remarkable zero-shot inference capabilities, but their potential for explicit structural causal discovery remains underexplored. To bridge this gap, we propose DCD-PFN, a decoupling-aware foundation model for causal discovery. Instead of directly amortizing global graph reconstruction, DCD-PFN focuses on local causal discovery through a decoupling-based paradigm. Through pre-training on diverse synthetic Structural Causal Models (SCMs), the model learns sample-wise decoupling weights that enable Markov boundary (MB) identification. Furthermore, by leveraging parallelized local discovery, DCD-PFN efficiently reconstructs global causal graphs while remaining grounded in the theoretical foundations of decoupling-based causal discovery. Experiments demonstrate that our foundation model achieves robust zero-shot generalization.

cs.LG

Causal Discovery via Quantile Partial Effect

Quantile Partial Effect (QPE) is a statistic associated with conditional quantile regression, measuring the effect of covariates at different levels. Our theory demonstrates that when the QPE of cause on effect is assumed to lie in a finite linear span, cause and effect are identifiable from their observational distribution. This generalizes previous identifiability results based on Functional Causal Models (FCMs) with additive, heteroscedastic noise, etc. Meanwhile, since QPE resides entirely at the observational level, this parametric assumption does not require considering mechanisms, noise, or even the Markov assumption, but rather directly utilizes the asymmetry of shape characteristics in the observational distribution. By performing basis function tests on the estimated QPE, causal directions can be distinguished, which is empirically shown to be effective in experiments on a large number of bivariate causal discovery datasets. For multivariate causal discovery, leveraging the close connection between QPE and score functions, we find that Fisher Information is sufficient as a statistical measure to determine causal order when assumptions are made about the second moment of QPE. We validate the feasibility of using Fisher Information to identify causal order on multiple synthetic and real-world multivariate causal discovery datasets.

cs.LG

Exogenous Isomorphism for Counterfactual Identifiability

This paper investigates $\sim_{\mathcal{L}_3}$-identifiability, a form of complete counterfactual identifiability within the Pearl Causal Hierarchy (PCH) framework, ensuring that all Structural Causal Models (SCMs) satisfying the given assumptions provide consistent answers to all causal questions. To simplify this problem, we introduce exogenous isomorphism and propose $\sim_{\mathrm{EI}}$-identifiability, reflecting the strength of model identifiability required for $\sim_{\mathcal{L}_3}$-identifiability. We explore sufficient assumptions for achieving $\sim_{\mathrm{EI}}$-identifiability in two special classes of SCMs: Bijective SCMs (BSCMs), based on counterfactual transport, and Triangular Monotonic SCMs (TM-SCMs), which extend $\sim_{\mathcal{L}_2}$-identifiability. Our results unify and generalize existing theories, providing theoretical guarantees for practical applications. Finally, we leverage neural TM-SCMs to address the consistency problem in counterfactual reasoning, with experiments validating both the effectiveness of our method and the correctness of the theory.

cs.LG

Exogenous Matching: Learning Good Proposals for Tractable Counterfactual Estimation

We propose an importance sampling method for tractable and efficient estimation of counterfactual expressions in general settings, named Exogenous Matching. By minimizing a common upper bound of counterfactual estimators, we transform the variance minimization problem into a conditional distribution learning problem, enabling its integration with existing conditional distribution modeling approaches. We validate the theoretical results through experiments under various types and settings of Structural Causal Models (SCMs) and demonstrate the outperformance on counterfactual estimation tasks compared to other existing importance sampling methods. We also explore the impact of injecting structural prior knowledge (counterfactual Markov boundaries) on the results. Finally, we apply this method to identifiable proxy SCMs and demonstrate the unbiasedness of the estimates, empirically illustrating the applicability of the method to practical scenarios.

cs.LG

CIER: A Novel Experience Replay Approach with Causal Inference in Deep Reinforcement Learning

In the training process of Deep Reinforcement Learning (DRL), agents require repetitive interactions with the environment. With an increase in training volume and model complexity, it is still a challenging problem to enhance data utilization and explainability of DRL training. This paper addresses these challenges by focusing on the temporal correlations within the time dimension of time series. We propose a novel approach to segment multivariate time series into meaningful subsequences and represent the time series based on these subsequences. Furthermore, the subsequences are employed for causal inference to identify fundamental causal factors that significantly impact training outcomes. We design a module to provide feedback on the causality during DRL training. Several experiments demonstrate the feasibility of our approach in common environments, confirming its ability to enhance the effectiveness of DRL training and impart a certain level of explainability to the training process. Additionally, we extended our approach with priority experience replay algorithm, and experimental results demonstrate the continued effectiveness of our approach.

cs.LG

On the possibility to detect gravitational waves from post-merger super-massive neutron stars with a kilohertz detector

The detection of a secular post-merger gravitational wave (GW) signal in a binary neutron star (BNS) merger serves as strong evidence for the formation of a long-lived post-merger neutron star (NS), which can help constrain the maximum mass of NSs and differentiate NS equation of states. We specifically focus on the detection of GW emissions from rigidly rotating NSs formed through BNS mergers, using several kilohertz GW detectors that have been designed. We simulate the BNS mergers within the detecting limit of LIGO-Virgo-KARGA O4 and attempt to find out on what fraction the simulated sources may have a detectable secular post-merger GW signal. For kilohertz detectors designed in the same configuration of LIGO A+, we find that the design with peak sensitivity at approximately $2{\rm kHz}$ is most appropriate for such signals. The fraction of sources that have a detectable secular post-merger GW signal would be approximately $0.94\% - 11\%$ when the spindowns of the post-merger rigidly rotating NSs are dominated by GW radiation, while be approximately $0.46\% - 1.6\%$ when the contribution of electromagnetic (EM) radiation to the spin-down processes is non-negligible. We also estimate this fraction based on other well-known proposed kilohertz GW detectors and find that, with advanced design, it can reach approximately $12\% - 45\%$ for the GW-dominated spindown case and $4.7\% - 16\%$ when both the GW and EM radiations are considered.

gr-qc

Towards observing the neutron star collapse with gravitational wave detectors

Gravitational waves from binary neutron star inspirals have been detected along with the electromagnetic transients coming from the aftermath of the merger in GW170817. However, much is still unknown about the post-merger dynamics that connects these two sets of observables. This includes if, and when, the post-merger remnant star collapses to a black hole, and what are the necessary conditions to power a short gamma-ray burst, and other observed electromagnetic counterparts. Observing the collapse of the post-merger neutron star would shed led on these questions, constraining models for the short gamma-ray burst engine and the hot neutron star equation of state. In this work, we explore the scope of using gravitational wave detectors to measure the timing of the collapse either indirectly, by establishing the shut-off of the post-merger gravitational emission, or---more challengingly---directly, by detecting the collapse signal. For the indirect approach, we consider a kilohertz high-frequency detector design that utilises a previously studied coupled arm cavity and signal recycling cavity resonance. This design would give a signal-to-noise ratio of 0.5\,-\,8.6 (depending on the variation of waveform parameters) for a collapse gravitational wave signal occurring at 10\,ms post-merger of a binary at 50\,Mpc and with total mass $2.7 M_\odot$. For the direct approach, we propose a narrow-band detector design, utilising the sensitivity around the frequency of the arm cavity free spectral range. The proposed detector achieves a signal-to-noise ratio of 0.3\,-\,1.9, independent of the collapse time. This detector is limited by both the fundamental classical and quantum noise with the arm cavity power chosen as 10\,MW.

gr-qc