SearcharxivSearch

arXiv subjects

Valeria Vitelli

Publications and source records attributed to Valeria Vitelli.

16 recordsLinked to original sources

Bayesian Plackett--Luce latent block models for ranked data

We introduce a Bayesian latent block model that jointly partitions assessors and items under a Plackett--Luce observation model. Assessors are assigned to $C$ clusters and items to $K$ blocks; items in a block share a common strength parameter within each assessor cluster, yielding a parsimonious $C\times K$ co-clustering representation. Independent Gnedin priors infer $C$ and $K$. Data augmentation gives conjugate Gibbs updates and a tractable MCMC sampler with split-merge moves. Simulations characterize recovery and posterior uncertainty as signal, ranking depth, and group balance vary. Applied to the cancer gene atlas (TCGA) pan-cancer top-500 gene-expression rankings, the model reveals tissue-driven sample structure while compressing gene-level heterogeneity into interpretable blocks. Rank-based GSEA of posterior gene scores supports biological interpretation.

stat.ME

Inverse Probability Weighting in a Post-Bayesian World

We present a justification of the use of Inverse Probability Weighting (IPW) in a post-Bayesian framework, in which the bias-correction provided by IPW in a frequentist context is reframed as a reweighting of the Kullback-Leibler (KL) divergence between the statistical model and the true data-generating parameter value. We provide a coherent argument in support of this approach, including theoretical results concerning convergence and properties of the generalised belief posteriors. We present examples demonstrating the utility of post-Bayesian IPW in practice: these include two simulated examples of inference under selection bias in the observed data, and a large-scale real-data example concerning systematic biases present in registry data when using prostate-specific antigen (PSA) to predict prostate cancer mortality. The empirical and theoretical results together show the utility of IPW to address classes of problems previously intractable within a Bayesian approach.

stat.ME

Bayesian nonparametric Mallows model for clustering preference data

Preference learning refers to the learning of latent patterns from ranking and preference data of different kinds. Typical aims of preference learning are to infer a shared consensus ranking, to learn individual-level preferences, and to perform unsupervised clustering. The Mallows model is among the few approaches that can achieve all these objectives jointly. Previous work has developed computationally tractable methods for Bayesian inference based on a MCMC Metropolis-Hastings scheme, where clustering is performed via a finite mixture of Mallows models. Inference on the number of clusters is then conducted a posteriori. Here we propose a Bayesian nonparametric Mallows model, based on a Dirichlet process mixture model. This allows joint inference on the number of non-empty clusters and on the clustering allocation, as well as posterior inference on cluster-specific parameters. The implementation of the proposed sampling algorithm is integrated into the existing R package BayesMallows, which also supports data in the form of incomplete rankings and pairwise comparisons. Simulated data show good performance of the nonparametric model compared to a finite mixture model in terms of recovery of the correct number of clusters, while empirical data on movie ratings show the model's effectiveness in providing personalized movie recommendations on discarded ratings.

stat.ME

Bayesian genome-wide clustering and variable selection of transcriptomic data via rank-based mixtures

With the increasing availability of ranking data, there has been a growing demand for appropriate unsupervised rank-based inferential frameworks capable of handling high-dimensional datasets and providing uncertainty quantification for all estimates. Rank-based methods have also seen a growing popularity in -omics pipelines, as ranking continuous measurements provides a robust means of handling non-normally distributed data. The Bayesian Mallows model (BMM) has emerged as a promising choice because of its adaptability to various types of ranking data and its flexible framework, integrating cluster-wise rank aggregation with inference at the individual level. However, the scalability of BMM to ultra-high-dimensional settings, such as -omics analyses, has remained limited. The present paper addresses this issue by introducing the first rank-based model generalizing BMM to jointly handle clustering and variable selection, namely the lower-dimensional Bayesian Mallows Model Mixture (lowBM3). The proposed method provides a novel Bayesian framework that simultaneously handles heterogeneity in the sample, unsupervised parameter estimation, and model selection in a scalable manner for ultra-high-dimensional data. Additionally, a companion postprocessing framework is introduced to provide posterior summaries of the discrete posterior distributions of both the consensus ranking and the variable selector. Simulation studies are performed to assess the performance of the method. The usefulness of the method is also shown in an application to signature discovery for cancer genomics, where RNA-seq bulk gene expression data obtained from breast cancer patients are clustered genome-wide.

stat.ME

Bayesian Semiparametric Mixture Cure (Frailty) Models

In recent years, mixture cure models have gained increasing popularity in survival analysis as an alternative to the Cox proportional hazards model, particularly in settings where a subset of patients is considered cured. The proportional hazards mixture cure model is especially advantageous when the presence of a cured fraction can be reasonably assumed, providing a more accurate representation of long-term survival dynamics. In this study, we propose a novel hierarchical Bayesian framework for the semiparametric mixture cure model, which accommodates both the inclusion and exclusion of a frailty component, allowing for greater flexibility in capturing unobserved heterogeneity among patients. Samples from the posterior distribution are obtained using a Markov chain Monte Carlo method, leveraging a hierarchical structure inspired by Bayesian Lasso. Comprehensive simulation studies are conducted across diverse scenarios to evaluate the performance and robustness of the proposed models. Bayesian model comparison and assessment are performed using various criteria. Finally, the proposed approaches are applied to two well-known datasets in the cure model literature: the E1690 melanoma trial and a colon cancer clinical trial.

stat.ME

Functional structural equation modeling with latent variables

Handling latent variables in Structural Equation Models (SEMs) in a case where both the latent variables and their corresponding indicators in the measurement error part of the model are random curves presents significant challenges, especially with sparse data. In this paper, we develop a novel family of Functional Structural Equation Models (FSEMs) that incorporate latent variables modeled as Gaussian Processes (GPs). The introduced FSEMs are built upon functional regression models having response variables modeled as underlying GPs. The model flexibly adapts to cases when the random curves' realizations are observed only over a sparse subset of the domain, and the inferential framework is based on a restricted maximum likelihood approach. The advantage of this framework lies in its ability and flexibility in handling various data scenarios, including regularly and irregularly spaced points and thus missing data. To extract smooth estimates for the functional parameters, we employ a penalized likelihood approach that selects the smoothing parameters using a cross-validation method. We evaluate the performance of the proposed model using simulation studies and a real data example, which suggests that our model performs well in practice. The uncertainty associated with the estimates of the functional coefficients is also assessed by constructing confidence regions for each estimate. The goodness of fit indices that are commonly used to evaluate the fit of SEMs are developed for the FSEMs introduced in this paper. Overall, the proposed method is a promising approach for modeling functional data in SEMs with functional latent variables.

stat.ME

A Weibull Mixture Cure Frailty Model for High-dimensional Covariates

A novel mixture cure frailty model is introduced for handling censored survival data. Mixture cure models are preferable when the existence of a cured fraction among patients can be assumed. However, such models are heavily underexplored: frailty structures within cure models remain largely undeveloped, and furthermore, most existing methods do not work for high-dimensional datasets, when the number of predictors is significantly larger than the number of observations. In this study, we introduce a novel extension of the Weibull mixture cure model that incorporates a frailty component, employed to model an underlying latent population heterogeneity with respect to the outcome risk. Additionally, high-dimensional covariates are integrated into both the cure rate and survival part of the model, providing a comprehensive approach to employ the model in the context of high-dimensional omics data. We also perform variable selection via an adaptive elastic-net penalization, and propose a novel approach to inference using the expectation-maximization (EM) algorithm. Extensive simulation studies are conducted across various scenarios to demonstrate the performance of the model, and results indicate that our proposed method outperforms competitor models. We apply the novel approach to analyze RNAseq gene expression data from bulk breast cancer patients included in The Cancer Genome Atlas (TCGA) database. A set of prognostic biomarkers is then derived from selected genes, and subsequently validated via both functional enrichment analysis and comparison to the existing biological literature. Finally, a prognostic risk score index based on the identified biomarkers is proposed and validated by exploring the patients' survival.

stat.ME

Rank-based Bayesian clustering via covariate-informed Mallows mixtures

Data in the form of rankings, ratings, pair comparisons or clicks are frequently collected in diverse fields, from marketing to politics, to understand assessors' individual preferences. Combining such preference data with features associated with the assessors can lead to a better understanding of the assessors' behaviors and choices. The Mallows model is a popular model for rankings, as it flexibly adapts to different types of preference data, and the previously proposed Bayesian Mallows Model (BMM) offers a computationally efficient framework for Bayesian inference, also allowing capturing the users' heterogeneity via a finite mixture. We develop a Bayesian Mallows-based finite mixture model that performs clustering while also accounting for assessor-related features, called the Bayesian Mallows model with covariates (BMMx). BMMx is based on a similarity function that a priori favours the aggregation of assessors into a cluster when their covariates are similar, using the Product Partition models (PPMx) proposal. We present two approaches to measure the covariate similarity: one based on a novel deterministic function measuring the covariates' goodness-of-fit to the cluster, and one based on an augmented model as in PPMx. We investigate the performance of BMMx in both simulation experiments and real-data examples, showing the method's potential for advancing the understanding of assessor preferences and behaviors in different applications.

stat.ME

Pseudo-Mallows for Efficient Probabilistic Preference Learning

We propose the Pseudo-Mallows distribution over the set of all permutations of $n$ items, to approximate the posterior distribution with a Mallows likelihood. The Mallows model has been proven to be useful for recommender systems where it can be used to learn personal preferences from highly incomplete data provided by the users. Inference based on MCMC is however slow, preventing its use in real time applications. The Pseudo-Mallows distribution is a product of univariate discrete Mallows-like distributions, constrained to remain in the space of permutations. The quality of the approximation depends on the order of the $n$ items used to determine the factorization sequence. In a variational setting, we optimise the variational order parameter by minimising a marginalized KL-divergence. We propose an approximate algorithm for this discrete optimization, and conjecture a certain form of the optimal variational order that depends on the data. Empirical evidence and some theory support our conjecture. Sampling from the Pseudo-Mallows distribution allows fast preference learning, compared to alternative MCMC based options, when the data exists in form of partial rankings of the items or of clicking on some items. Through simulations and a real life data case study, we demonstrate that the Pseudo-Mallows model learns personal preferences very well and makes recommendations much more efficiently, while maintaining similar accuracy compared to the exact Bayesian Mallows model.

stat.ME

Rank-based Bayesian variable selection for genome-wide transcriptomic analyses

Variable selection is crucial in high-dimensional omics-based analyses, since it is biologically reasonable to assume only a subset of non-noisy features contributes to the data structures. However, the task is particularly hard in an unsupervised setting, and a priori ad hoc variable selection is still a very frequent approach, despite the evident drawbacks and lack of reproducibility. We propose a Bayesian variable selection approach for rank-based unsupervised transcriptomic analysis. Making use of data rankings instead of the actual continuous measurements increases the robustness of conclusions when compared to classical statistical methods, and embedding variable selection into the inferential tasks allows complete reproducibility. Specifically, we develop a novel extension of the Bayesian Mallows model for variable selection that allows for a full probabilistic analysis, leading to coherent quantification of uncertainties. Simulation studies demonstrate the versatility and robustness of the proposed method in a variety of scenarios, as well as its superiority with respect to several competitors when varying the data dimension or data generating process. We use the novel approach to analyse genome-wide RNAseq gene expression data from ovarian cancer patients: several genes that affect cancer development are correctly detected in a completely unsupervised fashion, showing the usefulness of the method in the context of signature discovery for cancer genomics. Moreover, the possibility to also perform uncertainty quantification plays a key role in the subsequent biological investigation.

stat.ME

Latent function-on-scalar regression models for observed sequences of binary data: a restricted likelihood approach

In this paper, we study a functional regression setting where the random response curve is unobserved, and only its dichotomized version observed at a sequence of correlated binary data is available. We propose a practical computational framework for maximum likelihood analysis via the parameter expansion technique. Compared to existing methods, our proposal relies on the use of a complete data likelihood, with the advantage of being able to handle non-equally spaced and missing observations effectively. The proposed method is used in the Function-on-Scalar regression setting, with the latent response variable being a Gaussian random element taking values in a separable Hilbert space. Smooth estimations of functional regression coefficients and principal components are provided by introducing an adaptive MCEM algorithm that circumvents selecting the smoothing parameters. Finally, the performance of our novel method is demonstrated by various simulation studies and on a real case study. The proposed method is implemented in the R package dfrr.

stat.ME

A novel framework for joint sparse clustering and alignment of functional data

We propose a novel framework for sparse functional clustering that also embeds an alignment step. Sparse functional clustering means finding a grouping structure while jointly detecting the parts of the curves' domains where their grouping structure shows the most. Misalignment is a well-known issue in functional data analysis, that can heavily affect functional clustering results if not properly handled. Therefore, we develop a sparse functional clustering procedure that accounts for the possible curve misalignment: the coherence of the functional measure used in the clustering step to the class where the warping functions are chosen is ensured, and the well-posedness of the sparse clustering problem is proved. A possible implementing algorithm is also proposed, that jointly performs all these tasks: clustering, alignment, and domain selection. The method is tested on simulated data in various realistic situations, and its application to the Berkeley Growth Study data and to the AneuRisk65 data set is discussed.

stat.ME

BayesMallows: An R Package for the Bayesian Mallows Model

BayesMallows is an R package for analyzing data in the form of rankings or preferences with the Mallows rank model, and its finite mixture extension, in a Bayesian probabilistic framework. The Mallows model is a well-known model, grounded on the idea that the probability density of an observed ranking decreases exponentially fast as its distance to the location parameter increases. Despite the model being quite popular, this is the first Bayesian implementation that allows a wide choice of distances, and that works well with a large amount of items to be ranked. BayesMallows supports footrule, Spearman, Kendall, Cayley, Hamming and Ulam distances, allowing full use of the rich expressiveness of the Mallows model. This is possible thanks to the implementation of fast algorithms for approximating the partition function of the model under various distances. Although developed for being used in computing the posterior distribution of the model, these algorithms may be of interest in their own right. BayesMallows handles non-standard data: partial rankings and pairwise comparisons, even in cases including non-transitive preference patterns. The advantage of the Bayesian paradigm in this context comes from its ability to coherently quantify posterior uncertainties of estimates of any quantity of interest. These posteriors are fully available to the user, and the package comes with convienient tools for summarizing and visualizing the posterior distributions.

stat.CO

A Bayesian Mallows approach to non-transitive pair comparison data: how human are sounds?

We are interested in learning how listeners perceive sounds as having human origins. An experiment was performed with a series of electronically synthesized sounds, and listeners were asked to compare them in pairs. We propose a Bayesian probabilistic method to learn individual preferences from non-transitive pairwise comparison data, as happens when one (or more) individual preferences in the data contradicts what is implied by the others. We build a Bayesian Mallows model in order to handle non-transitive data, with a latent layer of uncertainty which captures the generation of preference misreporting. We then develop a mixture extension of the Mallows model, able to learn individual preferences in a heterogeneous population. The results of our analysis of the musicology experiment are of interest to electroacoustic composers and sound designers, and to the audio industry in general, whose aim is to understand how computer generated sounds can be produced in order to sound more human.

stat.AP

Probabilistic preference learning with the Mallows rank model

Ranking and comparing items is crucial for collecting information about preferences in many areas, from marketing to politics. The Mallows rank model is among the most successful approaches to analyse rank data, but its computational complexity has limited its use to a particular form based on Kendall distance. We develop new computationally tractable methods for Bayesian inference in Mallows models that work with any right-invariant distance. Our method performs inference on the consensus ranking of the items, also when based on partial rankings, such as top-k items or pairwise comparisons. We prove that items that none of the assessors has ranked do not influence the maximum a posteriori consensus ranking, and can therefore be ignored. When assessors are many or heterogeneous, we propose a mixture model for clustering them in homogeneous subgroups, with cluster-specific consensus rankings. We develop approximate stochastic algorithms that allow a fully probabilistic analysis, leading to coherent quantifications of uncertainties. We make probabilistic predictions on the class membership of assessors based on their ranking of just some items, and predict missing individual preferences, as needed in recommendation systems. We test our approach using several experimental and benchmark datasets.

stat.ME

Sparse Clustering of Functional Data

We consider the problem of clustering functional data while jointly selecting the most relevant features for classification. This problem has never been tackled before in the functional data context, and it requires a proper definition of the concept of sparsity for functional data. Functional sparse clustering is here analytically defined as a variational problem with a hard thresholding constraint ensuring the sparsity of the solution. First, a unique solution to sparse clustering with hard thresholding in finite dimensions is proved to exist. Then, the infinite dimensional generalization is given and proved to have a unique solution. Both the multivariate and the functional version of sparse clustering with hard thresholding exhibits improvements on other standard and sparse clustering strategies on simulated data. A real functional data application is also shown.

stat.ME