SearcharxivSearch

arXiv subjects

Nial Friel

Publications and source records attributed to Nial Friel.

At least 19 recordsLinked to original sources

Bayesian Plackett--Luce latent block models for ranked data

We introduce a Bayesian latent block model that jointly partitions assessors and items under a Plackett--Luce observation model. Assessors are assigned to $C$ clusters and items to $K$ blocks; items in a block share a common strength parameter within each assessor cluster, yielding a parsimonious $C\times K$ co-clustering representation. Independent Gnedin priors infer $C$ and $K$. Data augmentation gives conjugate Gibbs updates and a tractable MCMC sampler with split-merge moves. Simulations characterize recovery and posterior uncertainty as signal, ranking depth, and group balance vary. Applied to the cancer gene atlas (TCGA) pan-cancer top-500 gene-expression rankings, the model reveals tissue-driven sample structure while compressing gene-level heterogeneity into interpretable blocks. Rank-based GSEA of posterior gene scores supports biological interpretation.

stat.ME

Bayesian Conway-Maxwell-Poisson model with spike-and slab priors for dispersed count data with application to football scores

Statistical modeling for goals scored in football is typically achieved using the Poisson distribution and its variants. Here we propose a Bayesian framework for modeling under- and over-dispersion in count data by combining the Conway-Maxwell-Poisson (CMP) likelihood with a spikeand-slab (SAS) prior on unit-specific dispersion parameters. The proposed methodology generalizes Poisson-based count data models by treating equidispersion as an explicit baseline, and offering probabilistic quantification of departures from this regime, while simultaneously estimating their magnitude. Posterior inference is performed through a tailored Metropolis-within-Gibbs sampler that handles the doubly-intractable likelihood and provides efficient posterior exploration. The new method is examined using simulated data to confirm its ability to capture non-equidispersion, and applied to English Premier League (EPL) data. Dispersion is modeled at the team level and linked to goal-scoring behavior, and allows for thresholding mechanisms to distinguish teams based on their posterior probability of non-equidispersion. The results reveal heterogeneities in team-specific dispersion in the EPL, and demonstrate improvements in both model fit and predictive performance with respect to the standard Poisson model.

stat.ME

Ordering Stochastic Block Models via prior transitivity

In directed networks, nodes may form groups with similar interaction patterns, while these groups may themselves follow an ordered structure. Existing methods typically treat these features separately, either clustering nodes without enforcing a coherent block order, or ranking individual nodes without allowing for structurally equivalent groups. We introduce the Transitive Stochastic Block Model (TSBM), a Bayesian model for directed weighted networks that uses transitivity-inducing priors to infer ordered blocks. The model separates the total volume of interaction between two nodes from the direction of interaction conditional on interaction occurring, so that hierarchy is imposed on directional imbalance rather than interaction frequency. We consider two order-restricted specifications: a flexible weak-stochastic-transitivity version, which excludes cyclic dominance patterns while allowing heterogeneous block-pair strengths, and a Toeplitz strong-stochastic-transitivity version, in which directional advantage increases with rank separation. Posterior inference is performed through a Gibbs sampler using P\'olya-Gamma data augmentation. Since ordered block labels are not exchangeable, we introduce an age-ordered partition prior to infer the number of blocks jointly with node allocation. Simulation studies show that order-constrained priors improve prediction and partition recovery, especially in sparse networks. Across six empirical directed networks, the TSBM improves predictive performance in four cases and yields partitions with clearer ordered structure. The results also identify cases, such as nearly deterministic dominance networks or non-transitive citation networks, where imposing ordered blocks can harm prediction. The TSBM therefore provides a probabilistic framework for estimating ordered groups and assessing when a transitive block structure is supported by the data.

stat.ME

The Bradley-Terry Stochastic Block Model

The Bradley-Terry model is widely used for the analysis of pairwise comparison data and, in essence, produces a ranking of the items under comparison. We embed the Bradley-Terry model within a stochastic block model, allowing items to cluster. The resulting Bradley-Terry SBM (BT-SBM) ranks clusters so that items within a cluster share the same tied rank. We develop a fully Bayesian specification in which all quantities-the number of blocks, their strengths, and item assignments-are jointly learned via a fast Gibbs sampler derived through a Thurstonian data augmentation. Despite its efficiency, the sampler yields coherent and interpretable posterior summaries for all model components. Our motivating application analyzes men's tennis results from ATP tournaments over the seasons 2000-2022. We find that the top 100 players can be broadly partitioned into three or four tiers in most seasons. Moreover, the size of the strongest tier was small from the mid-2000s to 2018 and has increased since, providing evidence that men's tennis has become more competitive in recent years.

stat.ME

A Zero-Inflated Poisson Latent Position Cluster Model

The latent position network model (LPM) is a popular approach for the statistical analysis of network data. A central aspect of this model is that it assigns nodes to random positions in a latent space, such that the probability of an interaction between each pair of individuals or nodes is determined by their distance in this latent space. A key feature of this model is that it allows one to visualize nuanced structures via the latent space representation. The LPM can be further extended to the Latent Position Cluster Model (LPCM), to accommodate the clustering of nodes by assuming that the latent positions are distributed following a finite mixture distribution. In this paper, we extend the LPCM to accommodate missing network data and apply this to non-negative discrete weighted social networks. By treating missing data as ``unusual'' zero interactions, we propose a combination of the LPCM with the zero-inflated Poisson distribution. Statistical inference is based on a novel partially collapsed Markov chain Monte Carlo algorithm, where a Mixture-of-Finite-Mixtures (MFM) model is adopted to automatically determine the number of clusters and optimal group partitioning. Our algorithm features a truncated absorb-eject move, which is a novel adaptation of an idea commonly used in collapsed samplers, within the context of MFMs. Another aspect of our work is that we illustrate our results on 3-dimensional latent spaces, maintaining clear visualizations while achieving more flexibility than 2-dimensional models. The performance of this approach is illustrated via three carefully designed simulation studies, as well as four different publicly available real networks, where some interesting new perspectives are uncovered.

stat.ME

Modelling superspreading dynamics and circadian rhythms in online discussion boards using Hawkes processes

Online boards offer a platform for sharing and discussing content, where discussion emerges as a cascade of comments in response to a post. Branching point process models offer a practical approach to modelling these cascades; however, existing models do not account for apparent features of empirical data. We address this gap by illustrating the flexibility of Hawkes processes to model data arising from this context as well as outlining the computational tools needed to service this class of models. For example, the distribution of replies within discussions tends to have a heavy tail. As such, a small number of posts and comments may generate many replies, while most generate few or none, similar to `superspreading' in epidemics. Here, we propose a novel model for online discussion, motivated by a dataset arising from discussions on the r/ireland subreddit, that accommodates such phenomena and develop a framework for Bayesian inference that considers in- and out-of-sample tests for goodness-of-fit. This analysis shows that discussions within this community follow a circadian rhythm and are subject to moderate superspreading dynamics. For example, we estimate that the expected discussion size is approximately four for initial posts between 04:00 and 12:00 but approximately 2.5 from 15:00 to 02:00. We also estimate that 58% to 62% of posts fail to generate any discussion, with 95% posterior probability. Thus, we demonstrate that our framework offers a general approach to modelling discussion on online boards.

stat.AP

Zero-inflated stochastic block modeling of efficiency-security tradeoffs in weighted criminal networks

Criminal networks arise from the unique attempt to balance a need of establishing frequent ties among affiliates to facilitate the coordination of illegal activities, with the necessity to sparsify the overall connectivity architecture to hide from law enforcement. This efficiency-security tradeoff is also combined with the creation of groups of redundant criminals that exhibit similar connectivity patterns, thus guaranteeing resilient network architectures. State-of-the-art models for such data are not designed to infer these unique structures. In contrast to such solutions we develop a computationally-tractable Bayesian zero-inflated Poisson stochastic block model (ZIP-SBM), which identifies groups of redundant criminals with similar connectivity patterns, and infers both overt and covert block interactions within and across such groups. This is accomplished by modeling weighted ties (corresponding to counts of interactions among pairs of criminals) via zero-inflated Poisson distributions with block-specific parameters that quantify complex patterns in the excess of zero ties in each block (security) relative to the distribution of the observed weighted ties within that block (efficiency). The performance of ZIP-SBM is illustrated in simulations and in a study of summits co-attendances in a complex Mafia organization, where we unveil efficiency-security structures adopted by the criminal organization that were hidden to previous analyses.

stat.AP

Bayesian Strategies for Repulsive Spatial Point Processes

There is increasing interest to develop Bayesian inferential algorithms for point process models with intractable likelihoods. A purpose of this paper is to illustrate the utility of using simulation based strategies, including Approximate Bayesian Computation (ABC) and Markov Chain Monte Carlo (MCMC) methods for this task. Shirota and Gelfand (2017) proposed an extended version of an ABC approach for Repulsive Spatial Point Processes (RSPP), but their algorithm was not correctly detailed. In this paper, we correct their method and, based on this, we propose a new ABC-MCMC algorithm to which Markov property is introduced compared to a typical ABC method. Though it is generally impractical to use, Monte Carlo approximations can be leveraged for intractable terms. Another aspect of this paper is to explore the use of the exchange algorithm and the noisy Metropolis-Hastings algorithm (Alquier et al., 2016) on RSPP. Comparisons to ABC-MCMC methods are also provided. We find that the inferential approaches outlined above yield good performance for RSPP in both simulated and real data applications and should be considered as viable approaches for the analysis of these models.

stat.CO

Clustered Mallows Model

Rankings are a type of preference elicitation that arise in experiments where assessors arrange items, for example, in decreasing order of utility. Orderings of n items labelled {1,...,n} denoted are permutations that reflect strict preferences. For a number of reasons, strict preferences can be unrealistic assumptions for real data. For example, when items share common traits it may be reasonable to attribute them equal ranks. Also, there can be different importance attributions to decisions that form the ranking. In a situation with, for example, a large number of items, an assessor may wish to rank at top a certain number items; to rank other items at the bottom and to express indifference to all others. In addition, when aggregating opinions, a judging body might be decisive about some parts of the rank but ambiguous for others. In this paper we extend the well-known Mallows (Mallows, 1957) model (MM) to accommodate item indifference, a phenomenon that can be in place for a variety of reasons, such as those above mentioned.The underlying grouping of similar items motivates the proposed Clustered Mallows Model (CMM). The CMM can be interpreted as a Mallows distribution for tied ranks where ties are learned from the data. The CMM provides the flexibility to combine strict and indifferent relations, achieving a simpler and robust representation of rank collections in the form of ordered clusters. Bayesian inference for the CMM is in the class of doubly-intractable problems since the model's normalisation constant is not available in closed form. We overcome this challenge by sampling from the posterior with a version of the exchange algorithm \citep{murray2006}. Real data analysis of food preferences and results of Formula 1 races are presented, illustrating the CMM in practical situations.

stat.ME

Bayesian Testing of Scientific Expectations Under Exponential Random Graph Models

The exponential random graph (ERGM) model is a commonly used statistical framework for studying the determinants of tie formations from social network data. To test scientific theories under the ERGM framework, statistical inferential techniques are generally used based on traditional significance testing using p-values. This methodology has certain limitations, however, such as its inconsistent behavior when the null hypothesis is true, its inability to quantify evidence in favor of a null hypothesis, and its inability to test multiple hypotheses with competing equality and/or order constraints on the parameters of interest in a direct manner. To tackle these shortcomings, this paper presents Bayes factors and posterior probabilities for testing scientific expectations under a Bayesian framework. The methodology is implemented in the R package 'BFpack'. The applicability of the methodology is illustrated using empirical collaboration networks and policy networks.

stat.ME

Latent Position Network Models

In this chapter, we present a review of latent position models for networks. We review the recent literature in this area and illustrate the basic aspects and properties of this modeling framework. Through several illustrative examples we highlight how the latent position model is able to capture important features of observed networks. We emphasize how the canonical design of this model has made it popular thanks to its ability to provide interpretable visualizations of complex network interactions. We outline the main extensions that have been introduced to this model, illustrating its flexibility and applicability.

stat.ME

Assessing competitive balance in the English Premier League for over forty seasons using a stochastic block model

Competitive balance is the subject of much interest in the sports analytics literature and beyond. In this paper, we develop a statistical network model based on an extension of the stochastic block model to assess the balance between teams in a league. Here we represent the outcome of all matches in a football season as a dense network with nodes identified by teams and categorical edges representing the outcome of each game as a win, draw or a loss. The main focus and motivation for this paper is to provide a statistical framework to assess the issue of competitive balance in the context of the English First Division / Premier League over more than 40 seasons. The Premier League is arguably one of the most popular leagues in the world, in terms of its global reach and the revenue which it generates. Therefore it is of wide interest to assess its competitiveness. Our analysis provides evidence suggesting a structural change around the early 2000's from a reasonably balanced league to a two-tier league.

stat.AP

Multivariate Conway-Maxwell-Poisson Distribution: Sarmanov Method and Doubly-Intractable Bayesian Inference

In this paper, a multivariate count distribution with Conway-Maxwell (COM)-Poisson marginals is proposed. To do this, we develop a modification of the Sarmanov method for constructing multivariate distributions. Our multivariate COM-Poisson (MultCOMP) model has desirable features such as (i) it admits a flexible covariance matrix allowing for both negative and positive non-diagonal entries; (ii) it overcomes the limitation of the existing bivariate COM-Poisson distributions in the literature that do not have COM-Poisson marginals; (iii) it allows for the analysis of multivariate counts and is not just limited to bivariate counts. Inferential challenges are presented by the likelihood specification as it depends on a number of intractable normalizing constants involving the model parameters. These obstacles motivate us to propose a Bayesian inferential approach where the resulting doubly-intractable posterior is dealt with via the exchange algorithm and the Grouped Independence Metropolis-Hastings algorithm. Numerical experiments based on simulations are presented to illustrate the proposed Bayesian approach. We analyze the potential of the MultCOMP model through a real data application on the numbers of goals scored by the home and away teams in the Premier League from 2018 to 2021. Here, our interest is to assess the effect of a lack of crowds during the COVID-19 pandemic on the well-known home team advantage. A MultCOMP model fit shows that there is evidence of a decreased number of goals scored by the home team, not accompanied by a reduced score from the opponent. Hence, our analysis suggests a smaller home team advantage in the absence of crowds, which agrees with the opinion of several football experts.

stat.ME

Assessing epidemic curves for evidence of superspreading

The expected number of secondary infections arising from each index case, referred to as the reproduction or $R$ number, is a vital summary statistic for understanding and managing epidemic diseases. There are many methods for estimating $R$; however, few explicitly model heterogeneous disease reproduction, which gives rise to superspreading within the population. We propose a parsimonious discrete-time branching process model for epidemic curves that incorporates heterogeneous individual reproduction numbers. Our Bayesian approach to inference illustrates that this heterogeneity results in less certainty on estimates of the time-varying cohort reproduction number $R_t$. We apply these methods to a COVID-19 epidemic curve for the Republic of Ireland and find support for heterogeneous disease reproduction. Our analysis allows us to estimate the expected proportion of secondary infections attributable to the most infectious proportion of the population. For example, we estimate that the 20% most infectious index cases account for approximately 75-98% of the expected secondary infections with 95% posterior probability. In addition, we highlight that heterogeneity is a vital consideration when estimating $R_t$.

stat.AP

Calibrating COVID-19 SEIR models with time-varying effective contact rates

We describe the population-based SEIR (susceptible, exposed, infected, removed) model developed by the Irish Epidemiological Modelling Advisory Group (IEMAG), which advises the Irish government on COVID-19 responses. The model assumes a time-varying effective contact rate (equivalently, a time-varying reproduction number) to model the effect of non-pharmaceutical interventions. A crucial technical challenge in applying such models is their accurate calibration to observed data, e.g., to the daily number of confirmed new cases, as the past history of the disease strongly affects predictions of future scenarios. We demonstrate an approach based on inversion of the SEIR equations in conjunction with statistical modelling and spline-fitting of the data, to produce a robust methodology for calibration of a wide class of models of this type.

physics.soc-ph

A Bayesian latent allocation model for clustering compositional data with application to the Great Barrier Reef

Relative abundance is a common metric to estimate the composition of species in ecological surveys reflecting patterns of commonness and rarity of biological assemblages. Measurements of coral reef compositions formed by four communities along Australia's Great Barrier Reef (GBR) gathered between 2012 and 2017 are the focus of this paper. We undertake the task of finding clusters of transect locations with similar community composition and investigate changes in clustering dynamics over time. During these years, an unprecedented sequence of extreme weather events (cyclones and coral bleaching) impacted the 58 surveyed locations. The dependence between constituent parts of a composition presents a challenge for existing multivariate clustering approaches. In this paper, we introduce a finite mixture of Dirichlet distributions with group-specific parameters, where cluster memberships are dictated by unobserved latent variables. The inference is carried in a Bayesian framework, where MCMC strategies are outlined to sample from the posterior model. Simulation studies are presented to illustrate the performance of the model in a controlled setting. The application of the model to the 2012 coral reef data reveals that clusters were spatially distributed in similar ways across reefs which indicates a potential influence of wave exposure at the origin of coral reef community composition. The number of clusters estimated by the model decreased from four in 2012 to two from 2014 until 2017. Posterior probabilities of transect allocations to the same cluster substantially increase through time showing a potential homogenization of community composition across the whole GBR. The Bayesian model highlights the diversity of coral reef community composition within a coral reef and rapid changes across large spatial scales that may contribute to undermining the future of the GBR's biodiversity.

stat.AP

Statistical Network Analysis with Bergm

Recent advances in computational methods for intractable models have made network data increasingly amenable to statistical analysis. Exponential random graph models (ERGMs) emerged as one of the main families of models capable of capturing the complex dependence structure of network data in a wide range of applied contexts. The Bergm package for R has become a popular package to carry out Bayesian parameter inference, missing data imputation, model selection and goodness-of-fit diagnostics for ERGMs. Over the last few years, the package has been considerably improved in terms of efficiency by adopting some of the state-of-the-art Bayesian computational methods for doubly-intractable distributions. Recently, version 5 of the package has been made available on CRAN having undergone a substantial makeover, which has made it more accessible and easy to use for practitioners. New functions include data augmentation procedures based on the approximate exchange algorithm for dealing with missing data, adjusted pseudo-likelihood and pseudo-posterior procedures, which allow for fast approximate inference of the ERGM parameter posterior and model evidence for networks on several thousands nodes.

stat.CO

Bayesian inference, model selection and likelihood estimation using fast rejection sampling: the Conway-Maxwell-Poisson distribution

Bayesian inference for models with intractable likelihood functions represents a challenging suite of problems in modern statistics. In this work we analyse the Conway-Maxwell-Poisson (COM-Poisson) distribution, a two parameter generalisation of the Poisson distribution. COM-Poisson regression modelling allows the flexibility to model dispersed count data as part of a generalised linear model (GLM) with a COM-Poisson response, where exogenous covariates control the mean and dispersion level of the response. The major difficulty with COM-Poisson regression is that the likelihood function contains multiple intractable normalising constants and is not amenable to standard inference and MCMC techniques. Recent work by Chanialidis et al. (2017) has seen the development of a sampler to draw random variates from the COM-Poisson likelihood using a rejection sampling algorithm. We provide a new rejection sampler for the COM-Poisson distribution which significantly reduces the CPU time required to perform inference for COM-Poisson regression models. A novel extension of this work shows that for any intractable likelihood function with an associated rejection sampler it is possible to construct unbiased estimators of the intractable likelihood which proves useful for model selection or for use within pseudo-marginal MCMC algorithms (Andrieu and Roberts, 2009). We demonstrate all of these methods on a real-world dataset of takeover bids.

stat.CO