SearcharxivSearch

arXiv subjects

Silvia D'Angelo

Publications and source records attributed to Silvia D'Angelo.

6 recordsLinked to original sources

Data compression for fast dimension reduction and clustering of high-dimensional discrete data

High-dimensional discrete data are common in genomics, microbiomics, survey research, and digital behavioral analysis. Clustering such data is challenging because many existing methods are computationally expensive, sensitive to sparsity and discreteness, or designed for specific data types. We introduce a deterministic dimension-reduction framework for clustering high-dimensional discrete observations. The approach compresses observations into a low-dimensional continuous representation using weighted sums derived from a scaled positional encoding, yielding a numerically stable transformation applicable to both binary and count data. Several theoretical properties are established. The compression mapping is injective, ensuring that distinct observations remain distinguishable after transformation. Under mild regularity conditions, the compressed variables are approximately Gaussian, supporting the use of model-based clustering in the reduced space. We further show that separation between cluster centroids is preserved, indicating that location-based cluster structure remains identifiable following dimension reduction. Simulation studies demonstrate accurate cluster recovery across diverse settings, while achieving substantial computational savings compared with commonly used dimension-reduction techniques. Applications to microbiome data and United Nations rolling call voting data highlight the method's practical utility. Overall, the framework offers a scalable, efficient, and broadly applicable solution for clustering high-dimensional discrete data.

stat.ME

A Dynamic Latent Space Model for Healthcare Mobility Networks: the Italian National Health Service case

Healthcare mobility -- patients seeking treatment outside their territory of residence -- represents a major source of inequality and financial imbalance in decentralised health systems. In Italy, persistent north-south asymmetries in patient flows among Local Health Authorities (ASLs) have reinforced existing disparities within the National Health Service; yet the structural organisation and temporal dynamics of these flows remain poorly understood at the sub-regional level. We propose a Bayesian dynamic latent space model for directed weighted networks with a hurdle negative binomial likelihood, and apply it to administrative discharge records on mobility for hip replacement procedures among 109 Italian ASLs over 2018-2024. The model jointly addresses excess zeros, overdispersion and network dependence, while capturing directional heterogeneity through multiplicative sender and receiver effects and controlling for differences in territorial size via an appropriate exposure term. Applied to Italian mobility data, the model reveals the evolving geometry of the healthcare system, quantifies the disruption induced by the COVID-19 pandemic, and uncovers structural asymmetries in outward propensity and ASLs attractiveness. The framework provides a flexible tool for the statistical analysis of dynamic healthcare mobility networks with direct relevance to the monitoring and evaluation of territorial healthcare provision.

stat.ME

Model-based clustering for multidimensional social networks

Social network data are relational data recorded among a group of actors, interacting in different contexts. Often, the same set of actors can be characterized by multiple social relations, captured by a multidimensional network. A common situation is that of colleagues working in the same institution, whose social interactions can be defined on professional and personal levels. In addition, individuals in a network tend to interact more frequently with similar others, naturally creating communities. Latent space models for network data are useful to recover clustering of the actors, as they allow to represent similarities between them by their positions and relative distances in an interpretable low dimensional social space. We propose the infinite latent position cluster model for multidimensional network data, which enables model-based clustering of actors interacting across multiple social dimensions. The model is based on a Bayesian nonparametric framework, that allows to perform automatic inference on the clustering allocations, the number of clusters, and the latent social space. The method is tested on simulated data experiments, and it is employed to investigate the presence of communities in two multidimensional networks recording relationships of different types among colleagues.

stat.ME

Inferring food intake from multiple biomarkers using a latent variable model

Metabolomic based approaches have gained much attention in recent years due to their promising potential to deliver objective tools for assessment of food intake. In particular, multiple biomarkers have emerged for single foods. However, there is a lack of statistical tools available for combining multiple biomarkers to infer food intake. Furthermore, there is a paucity of approaches for estimating the uncertainty around biomarker based prediction of intake. Here, to facilitate inference on the relationship between multiple metabolomic biomarkers and food intake in an intervention study conducted under the A-DIET research programme, a latent variable model, multiMarker, is proposed. The proposed model draws on factor analytic and mixture of experts models, describing intake as a continuous latent variable whose value gives raise to the observed biomarker values. We employ a mixture of Gaussian distributions to flexibly model the latent variable. A Bayesian hierarchical modelling framework provides flexibility to adapt to different biomarker distributions and facilitates prediction of the latent intake along with its associated uncertainty. Simulation studies are conducted to assess the performance of the proposed multiMarker framework, prior to its application to the motivating application of quantifying apple intake.

stat.ME

Modelling heterogeneity in Latent Space Models for Multidimensional Networks

Multidimensional network data can have different levels of complexity, as nodes may be characterized by heterogeneous individual-specific features, which may vary across the networks. This paper introduces a class of models for multidimensional network data, where different levels of heterogeneity within and between networks can be considered. The proposed framework is developed in the family of latent space models, and it aims to distinguish symmetric relations between the nodes and node-specific features. Model parameters are estimated via a Markov Chain Monte Carlo algorithm. Simulated data and an application to a real example, on fruits import/export data, are used to illustrate and comment on the performance of the proposed models.

stat.ME

Latent Space Modeling of Multidimensional Networks with Application to the Exchange of Votes in Eurovision Song Contest

The Eurovision Song Contest is a popular TV singing competition held annually among country members of the European Broadcasting Union. In this competition, each member can be both contestant and jury, as it can participate with a song and/or vote for other countries' tunes. Throughout the years, the voting system has repeatedly been accused of being biased by the presence of tactical voting, according to which votes would represent strategic interests rather than actual musical preferences of the voting countries. In this work, we develop a latent space model to investigate the presence of a latent structure underlying the exchange of votes. Focusing on the period from 1998 to 2015, we represent the vote exchange as a multivariate network: each edition is a network, where countries are the nodes and two countries are linked by an edge if one voted for the other. The different networks are taken to be independent replicates of a common latent space capturing the overall relationships among the countries. Proximity denotes similarity, and countries close in the latent space are assumed to be more likely to exchange votes. Therefore, if the exchange of votes depends on the similarity between countries, the quality of the competing songs might not be a relevant factor in the determination of the voting preferences, and this would suggest the presence of bias. A Bayesian hierarchical modelling approach is employed to model the probability of a connection between any two countries as a function of their distance in the latent space, and of network-specific parameters and edge-specific covariates. The inferred latent space is found to be relevant in the determination of edge probabilities, however, the positions of the countries in such space only partially correspond to their actual geographical positions.

stat.AP