SearcharxivSearch

arXiv subjects

Gyeonghun Kang

Publications and source records attributed to Gyeonghun Kang.

3 recordsLinked to original sources

Transformers Can Learn Posterior Predictive Distributions In-Context

Prior-data fitted networks (PFNs) have recently emerged as a powerful approach for Bayesian prediction tasks, approximating the posterior predictive distribution (PPD) through in-context learning. Despite their strong empirical performance and ability to go beyond point predictions, theoretical understandings of the algorithmic capability of transformers to learn distributions in context are still lacking. Focusing on Gaussian process regression problems, we show by construction that transformers can implement a gradient descent algorithm targeting the posterior predictive mean and variance, followed by nonlinear mappings that yield binned probabilities of PPD. We study the error bounds of the approximated PPD in terms of attention depth and bin resolution. Based on these results, we further demonstrate the key role of normalization and the choice of attention depth in enabling the extrapolation abilities of transformers beyond the pretraining sample size range. We conduct simulations that corroborate our findings, providing insight into the expressivity of PFNs targeting PPDs and how architectural choices may influence generalization capabilities.

stat.ML

Multiscale Cochran-Mantel-Haenszel Scanning for Conditional Dependency

We propose a nonparametric approach to testing conditional independence and estimating conditional association, generalizing the Cochran-Mantel-Haenszel (CMH) test and odds-ratio estimator to continuous sample spaces. It leverages a multiscale scanning approach to decompose the sample space into a cascade of $2\times 2 \times T$ tables. Following the CMH test, we condition on the marginal order statistics, which are "almost ancillary" regarding conditional dependency. This strategy helps overcome a key challenge faced by other methods that discretize the sample space: we achieve consistency without requiring stratum sample sizes to grow to infinity, a constraint often difficult to satisfy in practice. Our method produces easy-to-compute test statistics with a known asymptotic null distribution under the conditional sampling model, scaling almost linearly with the sample size. Our simulation results demonstrate reliable Type I error control, even with small samples and high-dimensional conditioning, and competitive power compared to state-of-the-art tests. Finally, a case study on Uber ride-share data highlights the method's unique dual capability, inherited from the CMH, to both test and identify the nature of the inferred conditional association. By providing summary statistics that capture the strength and direction of local associations, our method offers practitioners a useful tool for learning conditional dependencies.

stat.ME

Model selection-based estimation for generalized additive models using mixtures of g-priors: Towards systematization

We explore the estimation of generalized additive models using basis expansion in conjunction with Bayesian model selection. Although Bayesian model selection is useful for regression splines, it has traditionally been applied mainly to Gaussian regression owing to the availability of a tractable marginal likelihood. We extend this method to handle an exponential family of distributions by using the Laplace approximation of the likelihood. Although this approach works well with any Gaussian prior distribution, consensus has not been reached on the best prior for nonparametric regression with basis expansions. Our investigation indicates that the classical unit information prior may not be ideal for nonparametric regression. Instead, we find that mixtures of g-priors are more effective. We evaluate various mixtures of g-priors to assess their performance in estimating generalized additive models. Additionally, we compare several priors for knots to determine the most effective strategy. Our simulation studies demonstrate that model selection-based approaches outperform other Bayesian methods.

stat.ME