SearcharxivSearch

arXiv subjects

Seonghyun Jeong

Publications and source records attributed to Seonghyun Jeong.

16 recordsLinked to original sources

Adaptive Functional Clustering with Structured Dependence via Variational Inference

Functional clustering is an important tool for identifying latent heterogeneity in functional data and has been widely applied across various scientific fields. However, many existing methods are not fully adaptive, as they may require the number of clusters to be prespecified and may lack automatic control over the smoothness of the underlying functions. They also commonly assume independent and identically distributed errors, thereby overlooking additional within-curve dependence. We propose a fully adaptive Bayesian procedure for functional clustering that addresses these limitations through Dirichlet process priors, adaptive smoothness control, and flexible covariance modeling. For computational scalability, the proposed method employs variational inference as an efficient alternative to Markov chain Monte Carlo. Together, these features provide a unified Bayesian framework for functional clustering and cluster-specific mean function estimation in the presence of structured within-curve dependence.

stat.ME

Bayesian Triangulation Splines: Spatial Adaptation on Irregular Domains

Conventional nonparametric regression methods for two-dimensional non-rectangular domains often overlook domain geometry and allow smoothing across boundaries. In spatial and geostatistical applications, this assumption is frequently invalid because domain boundaries typically constrain interactions among observations. Accommodating spatially varying smoothness is also substantially more challenging than in the univariate setting, and most existing methods do not adequately capture this local structure of the target function. To address these challenges, we propose Bayesian triangulation splines, which constructs locally adaptive splines over a polygonal domain. The method employs constrained Delaunay triangulations to respect boundary geometry and adapt to heterogeneous smoothness. A carefully designed prior further improves empirical performance. Under a global Sobolev smoothness assumption, we show that the proposed method achieves the optimal posterior contraction rate and adapts to unknown smoothness. We also show that the method exhibits ideal spatial adaptation in the sense that it achieves the oracle rate for inhomogeneous or locally varying structural features. Crucially, this oracle guarantee is not specific to constrained Delaunay triangulations, but holds over any triangulation satisfying weak shape-regularity conditions. Simulation studies confirm that the proposed method outperforms existing approaches by achieving higher estimation accuracy while maintaining low model complexity.

stat.ME

Laplace Variational Inference for Dirichlet Process Mixtures of Marked Poisson Point Processes

Marked point process data arise when events occur in a space with event-level marks. We study clustering of replicated marked Poisson point processes and introduce Dirichlet process mixtures of marked Poisson point processes, a Bayesian nonparametric model that jointly infers latent cluster structure, the number of clusters, and continuous mark-specific intensity surfaces. We use a squared link intensity representation to obtain tractable continuous domain likelihood terms without gridding or thinning. For posterior inference, we develop an efficient variational Bayes algorithm with a constrained Laplace approximation for the nonconjugate basis-coefficient block. The resulting coefficient update is formulated as a constrained optimization problem, which avoids the sign ambiguity and nodal-line issue of squared-link models. We further establish theoretical guarantees for mode finding optimization. We demonstrate the performance of the proposed model and algorithm through synthetic experiments and real-data analysis.

stat.ME

Bayesian Nonparametric Modeling for Multivariate Conditional Copula Regression with Varying Coefficients

Multivariate mixed-type outcomes are difficult to model jointly, and additional complexity arises when both marginal effects and dependence structures vary with a covariate such as age or time. Existing approaches often impose restrictive dependence assumptions or lack sufficient flexibility to accommodate heterogeneous response types in a unified framework. To address this issue, we propose a Bayesian nonparametric framework for multivariate conditional copula regression with varying coefficients. The proposed model combines adaptive spline-based marginal regressions with an infinite mixture of Gaussian copulas whose weights vary with the covariate through a probit stick-breaking process. This construction provides flexible covariate-dependent dependence modeling while avoiding explicit global constraints on functional correlation matrices. We further establish approximation results for the proposed copula representation and develop a Markov chain Monte Carlo algorithm for posterior inference. Simulation studies show accurate recovery under correct specification and robust performance under copula misspecification. In an analysis of the BRFSS 2023 data, the proposed model reveals age-varying marginal effects and dependence patterns among multiple health outcomes, providing a coherent joint view of multimorbidity beyond separate marginal analyses.

stat.ME

Posterior Contraction for Sparse Neural Networks in Besov Spaces with Intrinsic Dimensionality

This work establishes that sparse Bayesian neural networks achieve optimal posterior contraction rates over anisotropic Besov spaces and their hierarchical compositions. These structures reflect the intrinsic dimensionality of the underlying function, thereby mitigating the curse of dimensionality. Our analysis shows that Bayesian neural networks equipped with either sparse or continuous shrinkage priors attain the optimal rates which are dependent on the intrinsic dimension of the true structures. Moreover, we show that these priors enable rate adaptation, allowing the posterior to contract at the optimal rate even when the smoothness level of the true function is unknown. The proposed framework accommodates a broad class of functions, including additive and multiplicative Besov functions as special cases. These results advance the theoretical foundations of Bayesian neural networks and provide rigorous justification for their practical effectiveness in high-dimensional, structured estimation problems.

stat.ML

$L_2$-norm posterior contraction in Gaussian models with unknown variance

The testing-based approach is a fundamental tool for establishing posterior contraction rates. Although the Hellinger metric is attractive owing to the existence of a desirable test function, it is not directly applicable in Gaussian models, because translating the Hellinger metric into more intuitive metrics typically requires strong boundedness conditions. When the variance is known, this issue can be addressed by directly constructing a test function relative to the $L_2$-metric using the likelihood ratio test. However, when the variance is unknown, existing results are limited and rely on restrictive assumptions. To overcome this limitation, we derive a test function tailored to an unknown variance setting with respect to the $L_2$-metric and provide sufficient conditions for posterior contraction based on the testing-based approach. We apply this result to analyze high-dimensional regression and nonparametric regression.

math.ST

Penalty-Induced Basis Exploration for Bayesian Splines

Spline basis exploration via Bayesian model selection is a widely employed strategy for determining the optimal set of basis terms in nonparametric regression. However, despite its widespread use, this approach often encounters performance limitations owing to the finite approximation of infinite-dimensional parameters. This limitation arises because Bayesian model selection tends to favor simpler models over more complex ones when the true model is not among the candidates. Drawing inspiration from penalized splines, one potential remedy is to incorporate an additional roughness penalty that directly regulates the smoothness of functions. This strategy mitigates underfitting by allowing the inclusion of more basis terms while preventing overfitting through explicit smoothness control. Motivated by this insight, we propose a novel penalty-induced prior distribution for Bayesian basis exploration. The proposed prior evaluates the complexity of spline functions based on a convex combination of a roughness penalty and a ridge-type penalty for model selection. Our method adapts to the unknown level of smoothness and achieves the minimax-optimal posterior contraction rate up to a logarithmic factor. We also provide an efficient Markov chain Monte Carlo algorithm for its implementation. Extensive simulation studies demonstrate that our method outperforms competing approaches in terms of performance metrics and model complexity. An application to real datasets further substantiates the validity of our proposed approach.

stat.ME

Model selection-based estimation for generalized additive models using mixtures of g-priors: Towards systematization

We explore the estimation of generalized additive models using basis expansion in conjunction with Bayesian model selection. Although Bayesian model selection is useful for regression splines, it has traditionally been applied mainly to Gaussian regression owing to the availability of a tractable marginal likelihood. We extend this method to handle an exponential family of distributions by using the Laplace approximation of the likelihood. Although this approach works well with any Gaussian prior distribution, consensus has not been reached on the best prior for nonparametric regression with basis expansions. Our investigation indicates that the classical unit information prior may not be ideal for nonparametric regression. Instead, we find that mixtures of g-priors are more effective. We evaluate various mixtures of g-priors to assess their performance in estimating generalized additive models. Additionally, we compare several priors for knots to determine the most effective strategy. Our simulation studies demonstrate that model selection-based approaches outperform other Bayesian methods.

stat.ME

Unsupervised Outlier Detection using Random Subspace and Subsampling Ensembles of Dirichlet Process Mixtures

Probabilistic mixture models are recognized as effective tools for unsupervised outlier detection owing to their interpretability and global characteristics. Among these, Dirichlet process mixture models stand out as a strong alternative to conventional finite mixture models for both clustering and outlier detection tasks. Unlike finite mixture models, Dirichlet process mixtures are infinite mixture models that automatically determine the number of mixture components based on the data. Despite their advantages, the adoption of Dirichlet process mixture models for unsupervised outlier detection has been limited by challenges related to computational inefficiency and sensitivity to outliers in the construction of outlier detectors. Additionally, Dirichlet process Gaussian mixtures struggle to effectively model non-Gaussian data with discrete or binary features. To address these challenges, we propose a novel outlier detection method that utilizes ensembles of Dirichlet process Gaussian mixtures. This unsupervised algorithm employs random subspace and subsampling ensembles to ensure efficient computation and improve the robustness of the outlier detector. The ensemble approach further improves the suitability of the proposed method for detecting outliers in non-Gaussian data. Furthermore, our method uses variational inference for Dirichlet process mixtures, which ensures both efficient and rapid computation. Empirical analyses using benchmark datasets demonstrate that our method outperforms existing approaches in unsupervised outlier detection.

cs.LG

A Bayesian Convolutional Neural Network-based Generalized Linear Model

Convolutional neural networks (CNNs) provide flexible function approximations for a wide variety of applications when the input variables are in the form of images or spatial data. Although CNNs often outperform traditional statistical models in prediction accuracy, statistical inference, such as estimating the effects of covariates and quantifying the prediction uncertainty, is not trivial due to the highly complicated model structure and overparameterization. To address this challenge, we propose a new Bayesian approach by embedding CNNs within the generalized linear models (GLMs) framework. We use extracted nodes from the last hidden layer of CNN with Monte Carlo (MC) dropout as informative covariates in GLM. This improves accuracy in prediction and regression coefficient inference, allowing for the interpretation of coefficients and uncertainty quantification. By fitting ensemble GLMs across multiple realizations from MC dropout, we can account for uncertainties in extracting the features. We apply our methods to biological and epidemiological problems, which have both high-dimensional correlated inputs and vector covariates. Specifically, we consider malaria incidence data, brain tumor image data, and fMRI data. By extracting information from correlated inputs, the proposed method can provide an interpretable Bayesian analysis. The algorithm can be broadly applicable to image regressions or correlated data analysis by enabling accurate Bayesian inference quickly.

stat.ME

The art of BART: Minimax optimality over nonhomogeneous smoothness in high dimension

Many asymptotically minimax procedures for function estimation often rely on somewhat arbitrary and restrictive assumptions such as isotropy or spatial homogeneity. This work enhances the theoretical understanding of Bayesian additive regression trees under substantially relaxed smoothness assumptions. We provide a comprehensive study of asymptotic optimality and posterior contraction of Bayesian forests when the regression function has anisotropic smoothness that possibly varies over the function domain. The regression function can also be possibly discontinuous. We introduce a new class of sparse {\em piecewise heterogeneous anisotropic} Hölder functions and derive their minimax lower bound of estimation in high-dimensional scenarios under the $L_2$-loss. We then find that the Bayesian tree priors, coupled with a Dirichlet subset selection prior for sparse estimation in high-dimensional scenarios, adapt to unknown heterogeneous smoothness, discontinuity, and sparsity. These results show that Bayesian forests are uniquely suited for more general estimation problems that would render other default machine learning tools, such as Gaussian processes, suboptimal. Our numerical study shows that Bayesian forests often outperform other competitors such as random forests and deep neural networks, which are believed to work well for discontinuous or complicated smooth functions. Beyond nonparametric regression, we also examined posterior contraction of Bayesian forests for density estimation and binary classification using the technique developed in this study.

math.ST

Functional clustering methods for binary longitudinal data with temporal heterogeneity

In the analysis of binary longitudinal data, it is of interest to model a dynamic relationship between a response and covariates as a function of time, while also investigating similar patterns of time-dependent interactions. We present a novel generalized varying-coefficient model that accounts for within-subject variability and simultaneously clusters varying-coefficient functions, without restricting the number of clusters nor overfitting the data. In the analysis of a heterogeneous series of binary data, the model extracts population-level fixed effects, cluster-level varying effects, and subject-level random effects. Various simulation studies show the validity and utility of the proposed method to correctly specify cluster-specific varying-coefficients when the number of clusters is unknown. The proposed method is applied to a heterogeneous series of binary data in the German Socioeconomic Panel (GSOEP) study, where we identify three major clusters demonstrating the different varying effects of socioeconomic predictors as a function of age on the working status.

stat.ME

Posterior contraction in group sparse logit models for categorical responses

This paper studies posterior contraction rates in multi-category logit models with priors incorporating group sparse structures. We consider a general class of logit models that includes the well-known multinomial logit models as a special case. Group sparsity is useful when predictor variables are naturally clustered and particularly useful for variable selection in the multinomial logit models. We provide a unified platform for posterior contraction rates of group-sparse logit models that include binary logistic regression under individual sparsity. No size restriction is directly imposed on the true signal in this study. In addition to establishing the first-ever contraction properties for multi-category logit models under group sparsity, this work also refines recent findings on the Bayesian theory of binary logistic regression.

math.ST

Bayesian model selection in additive partial linear models via locally adaptive splines

We provide a flexible framework for selecting among a class of additive partial linear models that allows both linear and nonlinear additive components. In practice, it is challenging to determine which additive components should be excluded from the model while simultaneously determining whether nonzero additive components should be represented as linear or non-linear components in the final model. In this paper, we propose a Bayesian model selection method that is facilitated by a carefully specified class of models, including the choice of a prior distribution and the nonparametric model used for the nonlinear additive components. We employ a series of latent variables that determine the effect of each variable among the three possibilities (no effect, linear effect, and nonlinear effect) and that simultaneously determine the knots of each spline for a suitable penalization of smooth functions. The use of a pseudo-prior distribution along with a collapsing scheme enables us to deploy well-behaved Markov chain Monte Carlo samplers, both for model selection and for fitting the preferred model. Our method and algorithm are deployed on a suite of numerical studies and are applied to a nutritional epidemiology study. The numerical results show that the proposed methodology outperforms previously available methods in terms of effective sample sizes of the Markov chain samplers and the overall misclassification rates.

stat.ME

Unified Bayesian theory of sparse linear regression with nuisance parameters

We study frequentist asymptotic properties of Bayesian procedures for high-dimensional Gaussian sparse regression when unknown nuisance parameters are involved. Nuisance parameters can be finite-, high-, or infinite-dimensional. A mixture of point masses at zero and continuous distributions is used for the prior distribution on sparse regression coefficients, and appropriate prior distributions are used for nuisance parameters. The optimal posterior contraction of sparse regression coefficients, hampered by the presence of nuisance parameters, is also examined and discussed. It is shown that the procedure yields strong model selection consistency. A Bernstein-von Mises-type theorem for sparse regression coefficients is also obtained for uncertainty quantification through credible sets with guaranteed frequentist coverage. Asymptotic properties of numerous examples are investigated using the theories developed in this study.

math.ST

Bayesian Linear Regression for Multivariate Responses Under Group Sparsity

We study frequentist properties of a Bayesian high-dimensional multivariate linear regression model with correlated responses. The predictors are separated into many groups and the group structure is pre-determined. Two features of the model are unique: (i) group sparsity is imposed on the predictors. (ii) the covariance matrix is unknown and its dimensions can also be high. We choose a product of independent spike-and-slab priors on the regression coefficients and a new prior on the covariance matrix based on its eigendecomposition. Each spike-and-slab prior is a mixture of a point mass at zero and a multivariate density involving a $\ell_{2,1}$-norm. We first obtain the posterior contraction rate, the bounds on the effective dimension of the model with high posterior probabilities. We then show that the multivariate regression coefficients can be recovered under certain compatibility conditions. Finally, we quantify the uncertainty for the regression coefficients with frequentist validity through a Bernstein-von Mises type theorem. The result leads to selection consistency for the Bayesian method. We derive the posterior contraction rate using the general theory by constructing a suitable test from the first principle using moment bounds for certain likelihood ratios. This leads to posterior concentration around the truth with respect to the average Rényi divergence of order 1/2. This technique of obtaining the required tests for posterior contraction rate could be useful in many other problems.

math.ST