Searcharxiv⌕ Search

arXiv subjects

Byeongguk Kang

Publications and source records attributed to Byeongguk Kang.

2 recordsLinked to original sources

Local Fisher Information Enables Sparse Causal Discovery

Sparse causal discovery calls for methods that exploit graph structure without estimating high-dimensional densities. We introduce Fisher Information Completion Search (FiCS), a source-first algorithm for additive noise models that uses one local Fisher score for both ordering and parent selection. Under regularity and nonconstant-parent conditions, we prove that a node's local Fisher information equals the noise Fisher information exactly when the conditioning set contains all parents, provided that it contains no descendants. This Fisher parent completion identifies the parent set as the unique minimal Fisher completion. With a maximum conditioning set size $q$ at least the maximum indegree $d$, population FiCS queries marginals of at most $q+1$ variables and recovers the true directed acyclic graph under a positive ordering margin. Bounded conditioning also has a population advantage: reducing $q$ toward $d$ cannot decrease, and can strictly increase, the ordering margin. A growing non-Gaussian family separates local Fisher selection from conditional-variance and leaf-first Fisher ordering. For the regularized kernel Stein estimator, we establish high-dimensional DAG consistency under $q\{1+\log(p/q)\}+\log p=o(n)$, uniform Fisher separation, local approximation, and compatible ridge and parent penalty parameters. Experiments show the strongest gains when $n$ is small relative to $p$, quantify the effect of the conditioning size, and demonstrate competitive reference-graph recovery on three real-data benchmarks.

stat.ML↗

Hold-Out Scoring for Efficient Gaussian DAG Learning

High-dimensional Gaussian DAG learning faces a statistical-computational gap: methods with sharp sample complexity rely on computationally expensive subset search and a supplied indegree bound, whereas polynomial-time alternatives have less favorable sample complexity. We introduce HOST, an efficient DAG learning algorithm that replaces subset search with nodewise hold-out scoring and convex regression, without requiring a supplied indegree bound. Our key insight is that recovering a correct ordering does not require uniformly small estimation errors in ordering scores but only one-sided control of those errors. In the ordering step, HOST exploits the fact that score estimation using hold-out samples inflates ordering scores in expectation, which is the favorable direction for candidates that should not yet be selected. Given the ordering, HOST recovers parents by recursively removing indirect effects from total effects between two nodes. Under suitable conditions, HOST exactly recovers a $p$-node DAG of maximum indegree $d$ with sample complexity of order $d\log p$ in polynomial time. Experiments show that HOST achieves competitive graph recovery while exhibiting favorable runtime scaling.

stat.ML↗