Searcharxiv⌕ Search

arXiv subjects

Łukasz Rajkowski

Publications and source records attributed to Łukasz Rajkowski.

3 recordsLinked to original sources

A note on the geometry of the MAP partition in Conjugate Exponential Bayesian Mixture Models

We investigate the geometry of the maximal a posteriori (MAP) partition in the Bayesian Mixture Model where the component and the base distributions are chosen from conjugate exponential families. We prove that in this case the clusters are separated by the contour lines of a linear functional of the sufficient statistic. As a particular example, we describe Bayesian Mixture of Normals with Normal-inverse-Wishart prior on the component mean and covariance, in which the clusters in any MAP partition are separated by a quadratic surface. In connection with results of Rajkowski (2018), where the linear separability of clusters in the Bayesian Mixture Model with a fixed component covariance matrix was proved, it gives a nice Bayesian analogue of the geometric properties of Fisher Discriminant Analysis (LDA and QDA).

math.ST↗

A score function for Bayesian cluster analysis

We propose a score function for Bayesian clustering. The function is parameter free and captures the interplay between the within cluster variance and the between cluster entropy of a clustering. It can be used to choose the number of clusters in well-established clustering methods such as hierarchical clustering or $K$-means algorithm.

stat.OT↗

Analysis of the maximal posterior partition in the Dirichlet Process Gaussian Mixture Model

Mixture models are a natural choice in many applications, but it can be difficult to place an a priori upper bound on the number of components. To circumvent this, investigators are turning increasingly to Dirichlet process mixture models (DPMMs). It is therefore important to develop an understanding of the strengths and weaknesses of this approach. This work considers the MAP (maximum a posteriori) clustering for the Gaussian DPMM (where the cluster means have Gaussian distribution and, for each cluster, the observations within the cluster have Gaussian distribution). Some desirable properties of the MAP partition are proved: `almost disjointness' of the convex hulls of clusters (they may have at most one point in common) and (with natural assumptions) the comparability of sizes of those clusters that intersect any fixed ball with the number of observations (as the latter goes to infinity). Consequently, the number of such clusters remains bounded. Furthermore, if the data arises from independent identically distributed sampling from a given distribution with bounded support then the asymptotic MAP partition of the observation space maximises a function which has a straightforward expression, which depends only on the within-group covariance parameter. As the operator norm of this covariance parameter decreases, the number of clusters in the MAP partition becomes arbitrarily large, which may lead to the overestimation of the number of mixture components.

math.ST↗