Searcharxiv⌕ Search

arXiv subjects

Natalia Bochkina

Publications and source records attributed to Natalia Bochkina.

11 recordsLinked to original sources

Covariate-dependent hierarchical Dirichlet processes

Bayesian hierarchical modeling is a natural framework to effectively integrate data and borrow information across groups. In this paper, we address problems related to density estimation and identifying clusters across related groups, by proposing a hierarchical Bayesian approach that incorporates additional covariate information. To achieve flexibility, our approach builds on ideas from Bayesian nonparametrics, combining the hierarchical Dirichlet process with dependent Dirichlet processes. The proposed model is widely applicable, accommodating multiple and mixed covariate types through appropriate kernel functions as well as different output types through suitable component-specific likelihoods. This extends our ability to discern the relationship between covariates and clusters, while also effectively borrowing information and quantifying differences across groups. By employing a data augmentation trick, we are able to tackle the intractable normalized weights and construct a Markov chain Monte Carlo algorithm for posterior inference. The proposed method is illustrated on simulated data and two real data sets on single-cell RNA sequencing (scRNA-seq) and calcium imaging. For scRNA-seq data, we show that the incorporation of cell dynamics facilitates the discovery of additional cell subgroups. On calcium imaging data, our method identifies interpretable clusters of time frames with similar neural activity, aligning with the observed behavior of the animal.

stat.ME↗

Shared Differential Clustering across Single-cell RNA Sequencing Datasets with the Hierarchical Dirichlet Process

Single-cell RNA sequencing (scRNA-seq) is powerful technology that allows researchers to understand gene expression patterns at the single-cell level. However, analysing scRNA-seq data is challenging due to issues and biases in data collection. In this work, we construct an integrated Bayesian model that simultaneously addresses normalization, imputation and batch effects and also nonparametrically clusters cells into groups across multiple datasets. A Gibbs sampler based on a finite-dimensional approximation of the HDP is developed for posterior inference.

q-bio.GN↗

Bayesian Inverse Problems with Heterogeneous Variance

We consider inverse problems in Hilbert spaces under correlated Gaussian noise and use a Bayesian approach to find their regularised solution. We focus on mildly ill-posed inverse problems with the noise being generalised derivative of fractional Brownian motion, using a novel wavelet - based approach we call vaguelette-vaguelette. It allows us to apply sequence space methods without assuming that all operators are simultaneously diagonalisable. The results are proved for more general bases and covariance operators. Our primary aim is to study the posterior contraction rate in such inverse problems over Sobolev classes of true functions, comparing it to the derived minimax rate. Secondly, we study the effect of plugging in a consistent estimator of variances in sequence space on the posterior contraction rate, for instance where there are repeated observations. This result is also applied to the problem where the forward operator is observed with error. Thirdly, we show that an adaptive empirical Bayes posterior distribution contracts at the optimal rate, in the minimax sense, under a condition on prior smoothness, with a plugged in maximum marginal likelihood estimator of the prior scale. These theoretical results are illustrated on simulated data.

math.ST↗

Bernstein - von Mises theorem and misspecified models: a review

This is a review of asymptotic and non-asymptotic behaviour of Bayesian methods under model specification. In particular we focus on consistency, i.e. convergence of the posterior distribution to the point mass at the best parametric approximation to the true model, and conditions for it to be locally Gaussian around this point. For well specified regular models, variance of the Gaussian approximation coincides with the Fisher information, making Bayesian inference asymptotically efficient. In this review, we discuss how this is affected by model misspecification. We also discuss approaches to adjust Bayesian inference to make it asymptotically efficient under model misspecification.

math.ST↗

Testing for equal correlation matrices with application to paired gene expression data

We present a novel method for testing the hypothesis of equality of two correlation matrices using paired high-dimensional datasets. We consider test statistics based on the average of squares, maximum and sum of exceedances of Fisher transform sample correlations and we derive approximate null distributions using asymptotic and non-parametric distributions. Theoretical results on the power of the tests are presented and backed up by a range of simulation experiments. We apply the methodology to a case study of colorectal tumour gene expression data with the aim of discovering biological pathway lists of genes that present significantly different correlation matrices on healthy and tumour samples. We find strong evidence for a large part of the pathway lists correlation matrices to change among the two medical conditions.

stat.ME↗

Adaptive density estimation based on a mixture of Gammas

We consider the problem of Bayesian density estimation on the positive semiline for possibly unbounded densities. We propose a hierarchical Bayesian estimator based on the gamma mixture prior which can be viewed as a location mixture. We study convergence rates of Bayesian density estimators based on such mixtures. We construct approximations of the local Hölder densities, and of their extension to unbounded densities, to be continuous mixtures of gamma distributions, leading to approximations of such densities by finite mixtures. These results are then used to derive posterior concentration rates, with priors based on these mixture models. The rates are minimax (up to a log n term) and since the priors are independent of the smoothness the rates are adaptive to the smoothness.

math.ST↗

Joint Estimation of Sparse Networks with application to Paired Gene Expression data

We consider a method to jointly estimate sparse precision matrices and their underlying graph structures using dependent high-dimensional datasets. We present a penalized maximum likelihood estimator which encourages both sparsity and similarity in the estimated precision matrices where tuning parameters are automatically selected by controlling the expected number of false positive edges. We also incorporate an extra step to remove edges which represent an overestimation of triangular motifs. We conduct a simulation study to show that the proposed methodology presents consistent results for different combinations of sample size and dimension. Then, we apply the suggested approaches to a high-dimensional real case study of gene expression data with samples in two medical conditions, healthy and colon cancer tissues, to estimate a common network of genes as well as the differentially connected genes that are important to the disease. We find denser graph structures for healthy samples than for tumor samples, with groups of genes interacting together in the shape of clusters.

stat.AP↗

Selection of the Regularization Parameter in Graphical Models using Network Characteristics

Gaussian graphical models represent the underlying graph structure of conditional dependence between random variables which can be determined using their partial correlation or precision matrix. In a high-dimensional setting, the precision matrix is estimated using penalized likelihood by adding a penalization term which controls the amount of sparsity in the precision matrix and totally characterizes the complexity and structure of the graph. The most commonly used penalization term is the L1 norm of the precision matrix scaled by the regularization parameter which determines the trade-off between sparsity of the graph and fit to the data. In this paper we propose several procedures to select the regularization parameter in the estimation of graphical models that focus on recovering reliably the appropriate network structure of the graph. We conduct an extensive simulation study to show that the proposed methods produce useful results for different network topologies. The approaches are also applied in a high-dimensional real case study of gene expression data with the aim to discover the genes relevant to colon cancer. Using this data, we find graph structures which are verified to display significant biological gene associations.

stat.AP↗

Besov regularity of functions with sparse random wavelet coefficients

This paper addresses the problem of regularity properties of functions represented as an expansion in a wavelet basis with random coefficients in terms of finiteness of their Besov norm with probability 1. Such representations are used to specify a prior measure in Bayesian nonparametric wavelet regression. Investigating regularity of such functions is an important problem since the support of the posterior measure does not include functions that are not in the support of the prior measure, and therefore determines the functions that are possible to estimate using specified Bayesian model. We consider a more general parametrisation than has been studied previously which allows to study a priori regularity of functions under a wider class of priors. We also emphasise the difference between the abstract stochastic expansions that have been studied in the literature, and the expansions actually arising in nonparametric regression, and show that the latter cover a wider class of functions than the former. We also extend these results to stochastic expansions in an overcomplete wavelet dictionary.

math.ST↗

Consistency of the posterior distribution in generalised linear inverse problems

For ill-posed inverse problems, a regularised solution can be interpreted as a mode of the posterior distribution in a Bayesian framework. This framework enriches the set the solutions, as other posterior estimates can be used as a solution to the inverse problem, such as the posterior mean that can be easier to compute in practice. In this paper we prove consistency of Bayesian solutions of an ill-posed linear inverse problem in the Ky Fan metric for a general class of likelihoods and prior distributions in a finite dimensional setting. This result can be applied to study infinite dimensional problems by letting the dimension of the unknown parameter grow to infinity which can be viewed as discretisation on a grid or spectral approximation of an infinite dimensional problem. Likelihood and the prior distribution are assumed to be in an exponential form that includes distributions from the exponential family, and to be differentiable. The observations can be dependent. No assumption of finite moments of observations, such as expected value or the variance, is necessary thus allowing for possibly non-regular likelihoods, and allowing for non-conjugate and improper priors. If the variance exists, it may be heteroscedastic, namely, it may depend on the unknown function. We observe quite a surprising phenomenon when applying our result to the spectral approximation framework where it is possible to achieve the parametric rate of convergence, i.e the problem becomes self-regularised. We also consider a particular case of the unknown parameter being on the boundary of the parameter set, and show that the rate of convergence in this case is faster than for an interior point parameter.

math.ST↗

Sparse Empirical Bayes Analysis (SEBA)

We consider a joint processing of $n$ independent sparse regression problems. Each is based on a sample $(y_{i1},x_{i1})...,(y_{im},x_{im})$ of $m$ \iid observations from $y_{i1}=x_{i1}\tβ_i+\eps_{i1}$, $y_{i1}\in \R$, $x_{i 1}\in\R^p$, $i=1,...,n$, and $\eps_{i1}\dist N(0,\sig^2)$, say. $p$ is large enough so that the empirical risk minimizer is not consistent. We consider three possible extensions of the lasso estimator to deal with this problem, the lassoes, the group lasso and the RING lasso, each utilizing a different assumption how these problems are related. For each estimator we give a Bayesian interpretation, and we present both persistency analysis and non-asymptotic error bounds based on restricted eigenvalue - type assumptions.

stat.ML↗