SearcharxivSearch

arXiv subjects

A. Bekker

Publications and source records attributed to A. Bekker.

8 recordsLinked to original sources

Handling Missingness and Censoring in Dirichlet Models

Likelihood-based inference for compositional data generally requires fully observed compositions, hindering the direct treatment of missing or censored components on the simplex. In this paper, we develop an expectation-maximisation (EM)-type algorithm for maximum likelihood estimation of the Dirichlet parameters in the presence of missing and censored components under a unified coarsening framework. The Dirichlet distribution---the canonical probability model for compositional data, which plays a role analogous to that of the multivariate normal distribution for unconstrained multivariate data---provides the foundation for our methodology. Our methodology preserves the compositional structure of the data while simultaneously performing parameter estimation and model-based imputation. We evaluate the performance of our estimators and imputations through a simulation study under increasingly complex coarsening mechanisms, including both missing and censored data. We compare our method with an existing model-based approach and a nonparametric alternative. Finally, we illustrate the practical utility of our methodology using mercury speciation data, in which compositions are only partially observed because of detection limits and incomplete speciation. Our results indicate that the Dirichlet distribution provides a suitable model for these data and that our method yields imputations that better preserve the observed compositional structure than competing approaches.

stat.ME

A Computational Note on the Graphical Ridge in High-dimension

This article explores the estimation of precision matrices in high-dimensional Gaussian graphical models. We address the challenge of improving the accuracy of maximum likelihood-based precision estimation through penalization. Specifically, we consider an elastic net penalty, which incorporates both L1 and Frobenius norm penalties while accounting for the target matrix during estimation. To enhance precision matrix estimation, we propose a novel two-step estimator that combines the strengths of ridge and graphical lasso estimators. Through this approach, we aim to improve overall estimation performance. Our empirical analysis demonstrates the superior efficiency of our proposed method compared to alternative approaches. We validate the effectiveness of our proposal through numerical experiments and application on three real datasets. These examples illustrate the practical applicability and usefulness of our proposed estimator.

stat.ME

Soft computing for the posterior of a new matrix t graphical network

Modelling noisy data in a network context remains an unavoidable obstacle; fortunately, random matrix theory may comprehensively describe network environments effectively. Thus it necessitates the probabilistic characterisation of these networks (and accompanying noisy data) using matrix variate models. Denoising network data using a Bayes approach is not common in surveyed literature. This paper adopts the Bayesian viewpoint and introduces a new matrix variate t-model in a prior sense by relying on the matrix variate gamma distribution for the noise process, following the Gaussian graphical network for the cases when the normality assumption is violated. From a statistical learning viewpoint, such a theoretical consideration indubitably benefits the real-world comprehension of structures causing noisy data with network-based attributes as part of machine learning in data science. A full structural learning procedure is provided for calculating and approximating the resulting posterior of interest to assess the considered model's network centrality measures. Experiments with synthetic and real-world stock price data are performed not only to validate the proposed algorithm's capabilities but also to show that this model has wider flexibility than originally implied in Billio et al. (2021).

stat.ME

A naıve Bayesian graphical elastic net: driving advances in differential network analysis

Differential Networks (DNs), tools that encapsulate interactions within intricate systems, are brought under the Bayesian lens in this research. A novel naıve Bayesian adaptive graphical elastic net (BAE) prior is introduced to estimate the components of the DN. A heuristic structure determination mechanism and a block Gibbs sampler are derived. Performance is initially gauged on synthetic datasets encompassing various network topologies, aiming to assess and compare the flexibility to those of the Bayesian adaptive graphical lasso and ridge-type procedures. The naıve BAE estimator consistently ranks within the top two performers, highlighting its inherent adaptability. Finally, the BAE is applied to real-world datasets across diverse domains such as oncology, nephrology, and enology, underscoring its potential utility in comprehensive network analysis.

stat.ME

A Data Driven Bayesian Graphical Ridge Estimator

Bayesian methodologies prioritising accurate associations above sparsity in Gaussian graphical model (GGM) estimation remain relatively scarce in scientific literature. It is well accepted that the $\ell_2$ penalty enjoys a smaller computational footprint in GGM estimation, whilst the $\ell_1$ penalty encourages sparsity in the estimand. The Bayesian adaptive graphical lasso prior is used as a departure point in the formulation of a computationally efficient graphical ridge-type prior for events where accurate associations are prioritised over sparse representations. A novel block Gibbs sampler for simulating precision matrices is constructed using a ridge-type penalisation. The Bayesian graphical ridge-type prior is extended to a Bayesian adaptive graphical ridge-type prior. Synthetic experiments indicate that the graphical ridge-type estimators enjoy computational efficiency, in moderate dimensions, and numerical performance, for relatively non-sparse precision matrices, when compared to their lasso counterparts. The adaptive graphical ridge-type estimator is applied to cell signaling data to infer key associations between phosphorylated proteins in human T cell signalling. All computational workloads are carried out using the baygel R package.

stat.ME

Developing multivariate distributions using Dirichlet generator

There exist several endeavors proposing a new family of extended distributions using the beta-generating technique. This is a well-known mechanism in developing flexible distributions, by embedding the cumulative distribution function (cdf) of a baseline distribution within the beta distribution that acts as a generator. Univariate beta-generated distributions offer many fruitful and tractable properties and have applications in hydrology, biology and environmental sciences amongst other fields. In the univariate cases, this extension works well, however, for multivariate cases, the beta distribution generator delivers complex expressions. In this document, the proposed extension from the univariate to the multivariate domain addresses the need of flexible multivariate distributions that can model a wide range of multivariate data. This new family of multivariate distributions, whose marginals are beta-generated distributed, is constructed with the function H(x_{1},...,x_{p})=F(G_{1}(x_{1}),G_{2}(x_{2}),...,G_{p}(x_{p})), where $G_{i}(x_{i})$ are the cdfs of the gamma (baseline) distribution and F(.) as the cdf of the Dirichlet distribution. Hence as the main example, a general model having the support [0,1]^{p} (for p variates), using the Dirichlet as the generator, is developed together with some distributional properties, such as the moment generating function. The proposed Dirichlet-generated distributions can be applied to compositional data. The parameters of the model are estimated by using the maximum likelihood method. The effectiveness and prominence of the proposed family are illustrated by analyzing simulated as well as two real datasets. A new model testing technique is introduced to evaluate the performance of the multivariate models.

math.ST

Wishart Generator Distribution

The Wishart distribution and its generalizations are among the most prominent probability distributions in multivariate statistical analysis, arising naturally in applied research and as a basis for theoretical models. In this paper, we generalize the Wishart distribution utilizing a different approach that leads to the Wishart generator distribution with the Wishart distribution as a special case. It is not restricted, however some special cases are exhibited. Important statistical characteristics of the Wishart generator distribution are derived from the matrix theory viewpoint. Estimation is also touched upon as a guide for further research from the classical approach as well as from the Bayesian paradigm. The paper is concluded by giving applications of two special cases of this distribution in calculating the product of beta functions and astronomy.

math.ST

Kernel Oriented Generator Distribution

Matrix variate beta (MVB) distributions are used in different fields of hypothesis testing, multivariate correlation analysis, zero regression, canonical correlation analysis and etc. In this approach a unified methodology is proposed to generate matrix variate distributions by combining the kernel of MVB distributions of different types with an unknown Borel measurable function of trace operator over matrix space, called generator component. The latter component is a principal element of these newly defined generator type matrix variate distributions. The matrix variate Kummer beta distribution is amongst others a special case. Several statistical properties of this newly defined family of distributions are derived. In the conclusion other extensions and developments are discussed.

math.ST