SearcharxivSearch

arXiv subjects

Michael Fop

Publications and source records attributed to Michael Fop.

At least 19 recordsLinked to original sources

Data compression for fast dimension reduction and clustering of high-dimensional discrete data

High-dimensional discrete data are common in genomics, microbiomics, survey research, and digital behavioral analysis. Clustering such data is challenging because many existing methods are computationally expensive, sensitive to sparsity and discreteness, or designed for specific data types. We introduce a deterministic dimension-reduction framework for clustering high-dimensional discrete observations. The approach compresses observations into a low-dimensional continuous representation using weighted sums derived from a scaled positional encoding, yielding a numerically stable transformation applicable to both binary and count data. Several theoretical properties are established. The compression mapping is injective, ensuring that distinct observations remain distinguishable after transformation. Under mild regularity conditions, the compressed variables are approximately Gaussian, supporting the use of model-based clustering in the reduced space. We further show that separation between cluster centroids is preserved, indicating that location-based cluster structure remains identifiable following dimension reduction. Simulation studies demonstrate accurate cluster recovery across diverse settings, while achieving substantial computational savings compared with commonly used dimension-reduction techniques. Applications to microbiome data and United Nations rolling call voting data highlight the method's practical utility. Overall, the framework offers a scalable, efficient, and broadly applicable solution for clustering high-dimensional discrete data.

stat.ME

A density-based framework for community detection in attributed networks

Community structure in social and collaborative networks often emerges from a complex interplay between structural mechanisms, such as degree heterogeneity and leader-driven attraction, and homophily on node attributes. Existing community detection methods typically focus on these dimensions in isolation, limiting their ability to recover interpretable communities in presence of such mechanisms. In this paper, we propose AttDeCoDe, an attribute-driven extension of a density-based community detection framework, developed to analyse networks where node characteristics play a central role in group formation. Instead of defining density purely from network topology, AttDeCoDe estimates node-wise density in the attribute space, allowing communities to form around attribute-based community representatives while preserving structural connectivity constraints. This approach naturally captures homophily-driven aggregation while remaining sensitive to leader influence. We evaluate the proposed method through a simulation study based on a novel generative model that extends the degree-corrected stochastic block model by incorporating attribute-driven leader attraction, reflecting key features of collaborative research networks. We perform an empirical application to research collaboration data from the Horizon programmes, where organisations are characterised by project-level thematic descriptors. Both results show that AttDeCoDe offers a flexible and interpretable framework for community detection in attributed networks achieving competitive performance relative to topology-based and attribute-assisted benchmarks.

cs.SI

A latent variable model for identifying and characterizing food adulteration

Recently, growing consumer awareness of food quality and sustainability has led to a rising demand for effective food authentication methods. Vibrational spectroscopy techniques have emerged as a promising tool for collecting large volumes of data to detect food adulteration. However, spectroscopic data pose significant challenges from a statistical viewpoint, highlighting the need for more sophisticated modeling strategies. To address these challenges, in this work we propose a latent variable model specifically tailored for food adulterant detection, while accommodating the features of spectral data. Our proposal offers greater granularity with respect to existing approaches, since it does not only identify adulterated samples but also estimates the level of adulteration, and detects the spectral regions most affected by the adulterant. Consequently, the methodology offers deeper insights, and could facilitate the development of portable and faster instruments for efficient data collection in food authenticity studies. The method is applied to both synthetic and real honey mid-infrared spectroscopy data, delivering precise estimates of the adulteration level and accurately identifying which portions of the spectra are most impacted by the adulterant.

stat.ME

Community-level core-periphery structures in collaboration networks

Uncovering structural patterns in collaboration networks is key for understanding how knowledge flows and innovation emerges. These networks often exhibit a rich interplay of meso-scale structures, such as communities, core-periphery organization, and influential hubs, which shape the complexity of scientific collaboration. The coexistence of such structures challenges traditional approaches, which typically isolate specific network patterns at the node level. We introduce a novel framework for detecting core-periphery structures at the community level. Given a reference grouping of the nodes, the method optimizes an objective function that assigns core or peripheral roles to communities by accounting for the density and strength of their inter-community connections. The node-level partition may correspond to either inferred communities or to a node-attribute classification, such as discipline or location, enabling direct interpretation of how different social or organizational groups occupy central positions in the network. The method is motivated by an application to a co-authorship network of Italian academics in three different disciplines, where it reveals a hierarchical core-periphery structure associated with institutional role, regional location, and research topics.

stat.ME

A Latent Position Co-Clustering Model for Multiplex Networks

Multiplex networks are increasingly common across diverse domains, motivating the development of clustering methods that uncover patterns at multiple levels. Existing approaches typically focus on clustering either entire networks or nodes within a single network. We address the lack of a unified latent space framework for simultaneous network- and node-level clustering by proposing a latent position co-clustering model (LaPCoM), based on a hierarchical mixture-of-mixtures formulation. LaPCoM enables co-clustering of networks and their constituent nodes, providing joint dimension reduction and two-level cluster detection. At the network level, it identifies global homogeneity in topological patterns by grouping networks that share similar latent representations. At the node level, it captures local connectivity and community patterns. The model adopts a Bayesian nonparametric framework using a mixture of finite mixtures, which places priors on the number of clusters at both levels and incorporates sparse priors to encourage parsimonious clustering. Inference is performed via Markov chain Monte Carlo with automatic selection of the number of clusters. LaPCoM accommodates both binary and count-valued multiplex data. Simulation studies and comparisons with existing methods demonstrate accurate recovery of latent structure and clusters. Applications to real-world social multiplexes reveal interpretable network-level clusters aligned with context-specific patterns, and node-level clusters reflecting social patterns and roles.

stat.ME

A Gaussian process approach for rapid evaluation of skin tension

Skin tension plays a pivotal role in clinical settings, it affects scarring, wound healing and skin necrosis. Despite its importance, there is no widely accepted method for assessing in vivo skin tension or its natural pre-stretch. This study aims to utilise modern machine learning (ML) methods to develop a model that uses non-invasive measurements of surface wave speed to predict clinically useful skin properties such as stress and natural pre-stretch. A large dataset consisting of simulated wave propagation experiments was created using a simplified two-dimensional finite element (FE) model. Using this dataset, a sensitivity analysis was performed, highlighting the effect of the material parameters and material model on the Rayleigh and supersonic shear wave speeds. Then, a Gaussian process regression model was trained to solve the ill-posed inverse problem of predicting stress and pre-stretch of skin using measurements of surface wave speed. This model had good predictive performance (R2 = 0.9570) and it was possible to interpolate simplified parametric equations to calculate the stress and pre-stretch. To demonstrate that wave speed measurements could be obtained cheaply and easily, a simple experiment was devised to obtain wave speed measurements from synthetic skin at different values of pre-stretch. These experimental wave speeds agree well with the FE simulations and a model trained solely on the FE data provided accurate predictions of synthetic skin stiffness. Both the simulated and experimental results provide further evidence that elastic wave measurements coupled with ML models are a viable non-invasive method to determine in vivo skin tension.

physics.med-ph

Analysis of in-vivo skin anisotropy using elastic wave measurements and Bayesian modelling

In vivo skin exhibits viscoelastic, hyper-elastic and non-linear characteristics. It is under a constant non-equibiaxial tension in its natural configuration and is reinforced with oriented collagen fibers, giving rise to anisotropic behaviour. Understanding the complex mechanical behaviour of skin has relevance across many sectors including pharmaceuticals, cosmetics and surgery. However, there is a dearth of quality data characterizing human skin anisotropy in vivo. The available data is usually confined to limited population groups and/or limited angular resolution. Here, we use elastic waves travelling through the skin to obtain measurements from 78 volunteers from 3 to 93 years old. Using a Bayesian framework, we analyse the effect that age, gender and level of skin tension have on the skin anisotropy and stiffness. First, we propose a new measurement of anisotropy based on the eccentricity of angular data and conclude that it is a more robust measurement compared to the classic ``anisotropic ratio". We then find that in vivo skin anisotropy increases logarithmically with age, while the skin stiffness increases linearly along the direction of Langer Lines. We also conclude that gender does not significantly affect the skin anisotropy level, but does affect the overall stiffness, with males having stiffer skin on average. Finally, we find that skin tension significantly affects both the anisotropy and stiffness measurements, indicating that elastic wave measurements have promising applications in determining in vivo skin tension. In contrast to earlier studies, these results represent a comprehensive assessment of the variation of skin anisotropy with age and gender using a sizeable dataset and robust modern statistical analysis. This data has implications for the planning of surgical procedures and the adoption of universal cosmetic surgery practices for young or elderly patients.

physics.bio-ph

Scalable Durational Event Models: Application to Physical and Digital Interactions

Durable interactions are increasingly observed in social network analysis with precise timestamps. Phone and video calls, for instance, are events to which a specific duration can be assigned. We term data encoding interactions with start and end times ``durational event data''. Recent advances in data collection have enabled the observation of such data over extended periods and across large populations of actors. Methodologically, we propose the Durational Event Model, an extension of Relational Event Models that decouples the modeling of event incidence from event duration. Computationally, we derive a fast, memory-efficient, and exact block-coordinate ascent algorithm to facilitate large-scale inference. Theoretical complexity analysis and numerical simulations demonstrate the computational superiority of this approach over state-of-the-art methods. We apply the model implemented in the R package redeem to physical and digital interactions among college students in Copenhagen. Our empirical findings reveal that past interactions drive physical interactions, whereas digital interactions are influenced predominantly by friendship ties and prior dyadic contact.

stat.ME

Multiplex Dirichlet stochastic block model for clustering multidimensional compositional networks

Network data often represent multiple types of relations, which can also denote exchanged quantities, and are typically encompassed in a weighted multiplex. Such data frequently exhibit clustering structures, however, traditional clustering methods are not well-suited for multiplex networks. Additionally, standard methods treat edge weights in their raw form, potentially biasing clustering towards a node's total weight capacity rather than reflecting cluster-related interaction patterns. To address this, we propose transforming edge weights into a compositional format, enabling the analysis of connection strengths in relative terms and removing the impact of nodes' total weights. We introduce a multiplex Dirichlet stochastic block model designed for multiplex networks with compositional layers. This model accounts for sparse compositional networks and enables joint clustering across different types of interactions. We validate the model through a simulation study and apply it to the international export data from the Food and Agriculture Organization of the United Nations.

stat.ME

A Dirichlet stochastic block model for composition-weighted networks

Network data are observed in various applications where the individual entities of the system interact with or are connected to each other, and often these interactions are defined by their associated strength or importance. Clustering is a common task in network analysis that involves finding groups of nodes displaying similarities in the way they interact with the rest of the network. However, most clustering methods use the strengths of connections between entities in their original form, ignoring the possible differences in the capacities of individual nodes to send or receive edges. This often leads to clustering solutions that are heavily influenced by the nodes' capacities. One way to overcome this is to analyse the strengths of connections in relative rather than absolute terms, expressing each edge weight as a proportion of the sending (or receiving) capacity of the respective node. This, however, induces additional modelling constraints that most existing clustering methods are not designed to handle. In this work we propose a stochastic block model for composition-weighted networks based on direct modelling of compositional weight vectors using a Dirichlet mixture, with the parameters determined by the cluster labels of the sender and the receiver nodes. Inference is implemented via an extension of the classification expectation-maximisation algorithm that uses a working independence assumption, expressing the complete data likelihood of each node of the network as a function of fixed cluster labels of the remaining nodes. A model selection criterion is derived to aid the choice of the number of clusters. The model is validated using simulation studies, and showcased on network data from the Erasmus exchange program and a bike sharing network for the city of London.

stat.ME

Hausdorff Distance-Based Record Linkage for Improved Matching of Households and Individuals in Different Databases

Matching households and individuals across different databases poses challenges due to the lack of unique identifiers, typographical errors, and changes in attributes over time. Record linkage tools play a crucial role in overcoming these difficulties. This paper presents a multi-step record linkage procedure that incorporates household information to enhance the entity-matching process across multiple databases. Our approach utilizes the Hausdorff distance to estimate the probability of a match between households in multiple files. Subsequently, probabilities of matching individuals within these households are computed using a logistic regression model based on attribute-level distances. These estimated probabilities are then employed in a linear programming optimization framework to infer one-to-one matches between individuals. To assess the efficacy of our method, we apply it to link data from the Italian Survey of Household Income and Wealth across different years. Through internal and external validation procedures, the proposed method is shown to provide a significant enhancement in the quality of the individual matching process, thanks to the incorporation of household information. A comparison with a standard record linkage approach based on direct matching of individuals, which neglects household information, underscores the advantages of accounting for such information.

stat.AP

A consensus-constrained parsimonious Gaussian mixture model for clustering hyperspectral images

The use of hyperspectral imaging to investigate food samples has grown due to the improved performance and lower cost of instrumentation. Food engineers use hyperspectral images to classify the type and quality of a food sample, typically using classification methods. In order to train these methods, every pixel in each training image needs to be labelled. Typically, computationally cheap threshold-based approaches are used to label the pixels, and classification methods are trained based on those labels. However, threshold-based approaches are subjective and cannot be generalized across hyperspectral images taken in different conditions and of different foods. Here a consensus-constrained parsimonious Gaussian mixture model (ccPGMM) is proposed to label pixels in hyperspectral images using a model-based clustering approach. The ccPGMM utilizes information that is available on some pixels and specifies constraints on those pixels belonging to the same or different clusters while clustering the rest of the pixels in the image. A latent variable model is used to represent the high-dimensional data in terms of a small number of underlying latent factors. To ensure computational feasibility, a consensus clustering approach is employed, where the data are divided into multiple randomly selected subsets of variables and constrained clustering is applied to each data subset; the clustering results are then consolidated across all data subsets to provide a consensus clustering solution. The ccPGMM approach is applied to simulated datasets and real hyperspectral images of three types of puffed cereal, corn, rice, and wheat. Improved clustering performance and computational efficiency are demonstrated when compared to other current state-of-the-art approaches.

stat.ME

Variational Inference for the Latent Shrinkage Position Model

The latent position model (LPM) is a popular method used in network data analysis where nodes are assumed to be positioned in a $p$-dimensional latent space. The latent shrinkage position model (LSPM) is an extension of the LPM which automatically determines the number of effective dimensions of the latent space via a Bayesian nonparametric shrinkage prior. However, the LSPM reliance on Markov chain Monte Carlo for inference, while rigorous, is computationally expensive, making it challenging to scale to networks with large numbers of nodes. We introduce a variational inference approach for the LSPM, aiming to reduce computational demands while retaining the model's ability to intrinsically determine the number of effective latent dimensions. The performance of the variational LSPM is illustrated through simulation studies and its application to real-world network data. To promote wider adoption and ease of implementation, we also provide open-source code.

stat.ME

Model-based Clustering for Network Data via a Latent Shrinkage Position Cluster Model

Low-dimensional representation and clustering of network data are tasks of great interest across various fields. Latent position models are routinely used for this purpose by assuming that each node has a location in a low-dimensional latent space, and by enabling node clustering. However, these models fall short through their inability to simultaneously determine the latent space dimension and number of clusters. Here we introduce the latent shrinkage position cluster model (LSPCM), which addresses this limitation. The LSPCM posits an infinite dimensional latent space and assumes a Bayesian nonparametric shrinkage prior on the latent positions' variance parameters resulting in higher dimensions having increasingly smaller variances, aiding the identification of dimensions with non-negligible variance. Further, the LSPCM assumes the latent positions follow a sparse finite Gaussian mixture model, allowing for automatic inference on the number of clusters related to non-empty mixture components. As a result, the LSPCM simultaneously infers the effective dimension of the latent space and the number of clusters, eliminating the need to fit and compare multiple models. The performance of the LSPCM is assessed via simulation studies and demonstrated through application to two real Twitter network datasets from sporting and political contexts. Open source software is available to facilitate widespread use of the LSPCM.

stat.ME

Sparse model-based clustering of three-way data via lasso-type penalties

Mixtures of matrix Gaussian distributions provide a probabilistic framework for clustering continuous matrix-variate data, which are becoming increasingly prevalent in various fields. Despite its widespread adoption and successful application, this approach suffers from over-parameterization issues, making it less suitable even for matrix-variate data of moderate size. To overcome this drawback, we introduce a sparse model-based clustering approach for three-way data. Our approach assumes that the matrix mixture parameters are sparse and have different degree of sparsity across clusters, allowing to induce parsimony in a flexible manner. Estimation of the model relies on the maximization of a penalized likelihood, with specifically tailored group and graphical lasso penalties. These penalties enable the selection of the most informative features for clustering three-way data where variables are recorded over multiple occasions and allow to capture cluster-specific association structures. The proposed methodology is tested extensively on synthetic data and its validity is demonstrated in application to time-dependent crime patterns in different US cities.

stat.CO

A Latent Shrinkage Position Model for Binary and Count Network Data

Interactions between actors are frequently represented using a network. The latent position model is widely used for analysing network data, whereby each actor is positioned in a latent space. Inferring the dimension of this space is challenging. Often, for simplicity, two dimensions are used or model selection criteria are employed to select the dimension, but this requires choosing a criterion and the computational expense of fitting multiple models. Here the latent shrinkage position model (LSPM) is proposed which intrinsically infers the effective dimension of the latent space. The LSPM employs a Bayesian nonparametric multiplicative truncated gamma process prior that ensures shrinkage of the variance of the latent positions across higher dimensions. Dimensions with non-negligible variance are deemed most useful to describe the observed network, inducing automatic inference on the latent space dimension. While the LSPM is applicable to many network types, logistic and Poisson LSPMs are developed here for binary and count networks respectively. Inference proceeds via a Markov chain Monte Carlo algorithm, where novel surrogate proposal distributions reduce the computational burden. The LSPM's properties are assessed through simulation studies, and its utility is illustrated through application to real network datasets. Open source software assists wider implementation of the LSPM.

stat.ME

Group-wise shrinkage estimation in penalized model-based clustering

Finite Gaussian mixture models provide a powerful and widely employed probabilistic approach for clustering multivariate continuous data. However, the practical usefulness of these models is jeopardized in high-dimensional spaces, where they tend to be over-parameterized. As a consequence, different solutions have been proposed, often relying on matrix decompositions or variable selection strategies. Recently, a methodological link between Gaussian graphical models and finite mixtures has been established, paving the way for penalized model-based clustering in the presence of large precision matrices. Notwithstanding, current methodologies implicitly assume similar levels of sparsity across the classes, not accounting for different degrees of association between the variables across groups. We overcome this limitation by deriving group-wise penalty factors, which automatically enforce under or over-connectivity in the estimated graphs. The approach is entirely data-driven and does not require additional hyper-parameter specification. Analyses on synthetic and real data showcase the validity of our proposal.

stat.ME

Unobserved classes and extra variables in high-dimensional discriminant analysis

In supervised classification problems, the test set may contain data points belonging to classes not observed in the learning phase. Moreover, the same units in the test data may be measured on a set of additional variables recorded at a subsequent stage with respect to when the learning sample was collected. In this situation, the classifier built in the learning phase needs to adapt to handle potential unknown classes and the extra dimensions. We introduce a model-based discriminant approach, Dimension-Adaptive Mixture Discriminant Analysis (D-AMDA), which can detect unobserved classes and adapt to the increasing dimensionality. Model estimation is carried out via a full inductive approach based on an EM algorithm. The method is then embedded in a more general framework for adaptive variable selection and classification suitable for data of large dimensions. A simulation study and an artificial experiment related to classification of adulterated honey samples are used to validate the ability of the proposed framework to deal with complex situations.

stat.ME