SearcharxivSearch

arXiv subjects

Nikolay H. Balov

Publications and source records attributed to Nikolay H. Balov.

5 recordsLinked to original sources

Consistent Model Selection of Discrete Bayesian Networks from Incomplete Data

A maximum likelihood based model selection of discrete Bayesian networks is considered. The model selection is performed through scoring function $S$, which, for a given network $G$ and $n$-sample $D_n$, is defined to be the maximum log-likelihood $l$ minus a penalization term $λ_n h$ proportional to network complexity $h(G)$, $$ S(G|D_n) = l(G|D_n) - λ_n h(G). $$ The data is allowed to have missing values at random that has prompted, to improve the efficiency of estimation, a replacement of the standard log-likelihood with the sum of sample average node log-likelihoods. The latter avoids the exclusion of most partially missing data records and allows the comparison of models fitted to different samples. Provided that a discrete Bayesian network is identifiable for a given missing data distribution, we show that if the sequence $λ_n$ converges to zero at a slower rate than $n^{-{1/2}}$ then the estimation is consistent. Moreover, we establish that BIC model selection ($λ_n=0.5\log(n)/n$) applied to the node-average log-likelihood is in general not consistent. This is in contrast to the complete data case where BIC is known to be consistent. The conclusions are confirmed by numerical examples.

math.ST

On the Stochastic Rank of Metric Functions

For a class of integral operators with kernels metric functions on manifold we find some necessary and sufficient conditions to have finite rank. The problem we pose has a stochastic nature and boils down to the following alternative question. For a random sample of discrete points, what will be the probability the symmetric matrix of pairwise distances to have full rank? When the metric is an analytic function, the question finds full and satisfactory answer. As an important application, we consider a class of tensor systems of equations formulating the problem of recovering a manifold distribution from its covariance field and solve this problem for representing manifolds such as Euclidean space and unit sphere.

math.MG

Covariance fields

We introduce and study covariance fields of distributions on a Riemannian manifold. At each point on the manifold, covariance is defined to be a symmetric and positive definite (2,0)-tensor. Its product with the metric tensor specifies a linear operator on the respected tangent space. Collectively, these operators form a covariance operator field. We show that, in most circumstances, covariance fields are continuous. We also solve the inverse problem: recovering distribution from a covariance field. Surprisingly, this is not possible on Euclidean spaces. On non-Euclidean manifolds however, covariance fields are true distribution representations.

math.ST

Comparing and interpolating distributions on manifold

We are interested in comparing probability distributions defined on Riemannian manifold. The traditional approach to study a distribution relies on locating its mean point and finding the dispersion about that point. On a general manifold however, even if two distributions are sufficiently concentrated and have unique means, a comparison of their covariances is not possible due to the difference in local parametrizations. To circumvent the problem we associate a covariance field with each distribution and compare them at common points by applying a similarity invariant function on their representing matrices. In this way we are able to define distances between distributions. We also propose new approach for interpolating discrete distributions and derive some criteria that assure consistent results. Finally, we illustrate with some experimental results on the unit 2-sphere.

math.ST

Covariance of centered distributions on manifold

We define and study a family of distributions with domain complete Riemannian manifold. They are obtained by projection onto a fixed tangent space via the inverse exponential map. This construction is a popular choice in the literature for it makes it easy to generalize well known multivariate Euclidean distributions. However, most of the available solutions use coordinate specific definition that makes them less versatile. %We propose improvements in two directions. We define the distributions of interest in coordinate independent way by utilizing co-variant 2-tensors. Then we study the relation of these distributions to their Euclidean counterparts. In particular, we are interested in relating the covariance to the tensor that controls distribution concentration. We find approximating expression for this relation in general and give more precise formulas in case of manifolds of constant curvature, positive or negative. Results are confirmed by simulation studies of the standard normal distribution on the unit-sphere and hyperbolic plane.

math.ST