SearcharxivSearch

arXiv subjects

Leonardo Egidi

Publications and source records attributed to Leonardo Egidi.

15 recordsLinked to original sources

Stochastic Bayes factors: why, when, and how

The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capture uncertainty inherent in the data, leading to an analogy with frequentist p-values, and primarily reflect prior-predictive rather than posterior-predictive performance. We introduce the stochastic Bayes factor (SBF), a new framework that extends the BF by explicitly incorporating uncertainty via replicated data. Formally, the SBF is defined as a push-forward measure transferring the BF from the observed data space to that of replications. This approach generalizes previous calibration proposals, while emphasizing posterior-predictive replication as a robust alternative. We establish key theoretical properties, including model consistency, compatibility and dominance, ensuring that SBFs preserve desirable Bayesian guarantees. An algorithmic routine is then proposed to operationalize the SBF, guiding model discrimination in a principled way while naturally providing model calibration. Simulation studies and real applications confirm that the SBF offers improved robustness and predictive reliability compared to the classical BF, by providing a valuable tool for model comparison.

stat.ME

Mixture priors for replication studies

Replication of scientific studies is important for assessing the credibility of their results. However, there is no consensus on how to quantify the extent to which a replication study replicates an original result. We propose a novel Bayesian approach for replication studies based on mixture priors. The idea is to use a mixture of the posterior distribution based on the original study and a non-informative distribution as the prior for the analysis of the replication study. The mixture weight then determines the extent to which the original and replication data are pooled. Two distinct strategies are presented: one with fixed mixture weights, and one that introduces uncertainty by assigning a prior distribution to the mixture weight itself. Furthermore, it is shown how within this framework Bayes factors can be used for formal testing of relevant scientific hypotheses, such as tests on the presence or absence of an effect or whether the mixture weight equals zero (completely discounting the original data) or one (fully pooling with the original data). To showcase the practical application of the methodology, we analyze data from three replication studies. Our findings suggest that mixture priors are a valuable and intuitive alternative to other Bayesian methods for analyzing replication studies, such as hierarchical models and power priors. We provide the free and open source R package repmix that implements the proposed methodology.

stat.ME

Leicester's Tale: Another Perspective on the EPL 2015/16 Through Expected Goals (xG) Modelling

Probabilistic modeling is an effective tool for evaluating team performance and predicting outcomes in sports. However, an important question that hasn't been fully explored is whether these models can reliably reflect actual performance while assigning meaningful probabilities to rare results that differ greatly from expectations. In this study, we create an inference-based probabilistic framework built on expected goals (xG). This framework converts shot-level event data into season-level simulations of points, rankings, and outcome probabilities. Using the English Premier League 2015/16 season as a data, we demonstrate that the framework captures the overall structure of the league table. It correctly identifies the top-four contenders and relegation candidates while explaining a significant portion of the variance in final points and ranks. In a full-season evaluation, the model assigns a low probability to extreme outcomes, particularly Leicester City's historic title win, which stands out as a statistical anomaly. We then look at the ex ante inferential and early-diagnostic role of xG by only using mid-season information. With first-half data, we simulate the rest of the season and show that teams with stronger mid-season xG profiles tend to earn more points in the second half, even after considering their current league position. In this mid-season assessment, Leicester City ranks among the top teams by xG and is given a small but noteworthy chance of winning the league. This suggests that their ultimate success was unlikely but not entirely detached from their actual performance. Our analysis indicates that expected goals models work best as probabilistic baselines for analysis and early-warning diagnostics, rather than as certain predictors of rare season outcomes.

stat.ME

Bayesian weighted discrete-time dynamic models for association football prediction

In recent years, great emphasis has been placed on the prediction of association football. Due to this, several studies have proposed different types of statistical models to predict the outcome of a football match. However, most existing approaches usually assume that the offensive and defensive abilities of teams remain static over time. We introduce a Bayesian dynamic approach for football goal based models that uses period-specific commensurate priors to flexibly weight the evolution of attacking and defensive abilities. Our approach assigns separate, time varying precisions for each ability and period, controlled via spike and slab hyperpriors. This adaptive shrinkage borrows information about teams' strength when past and current performance aligns and allows rapid adjustments when teams experience substantial changes (e.g., transfer windows or coaching changes). We integrate this framework into six standard goal based models evaluating predictive performance using data from the last five seasons of the German Bundesliga, English Premier League, and Spanish La Liga. Compared with the other discrete time dynamic models, our adaptive approach yields better predictive performance. The proposed methodology has also been implemented in the free and open source R package footBayes.

stat.ME

Eliciting prior information from clinical trials via calibrated Bayes factor

In the Bayesian framework power prior distributions are increasingly adopted in clinical trials and similar studies to incorporate external and past information, typically to inform the parameter associated to a treatment effect. Their use is particularly effective in scenarios with small sample sizes and where robust prior information is actually available. A crucial component of this methodology is represented by its weight parameter, which controls the volume of historical information incorporated into the current analysis. This parameter can be considered as either fixed or random. Although various strategies exist for its determination, eliciting the prior distribution of the weight parameter according to a full Bayesian approach remains a challenge. In general, this parameter should be carefully selected to accurately reflect the available prior information without dominating the posterior inferential conclusions. To this aim, we propose a novel method for eliciting the prior distribution of the weight parameter through a simulation-based calibrated Bayes factor procedure. This approach allows for the prior distribution to be updated based on the strength of evidence provided by the data: The goal is to facilitate the integration of historical data when it aligns with current information and to limit it when discrepancies arise in terms, for instance, of prior-data conflicts. The performance of the proposed method is tested through simulation studies and applied to real data from clinical trials.

stat.ME

Alternative ranking measures to predict international football results

Over the last few years, there has been a growing interest in the prediction and modelling of competitive sports outcomes, with particular emphasis placed on this area by the Bayesian statistics and machine learning communities. In this paper, we have carried out a comparative evaluation of statistical and machine learning models to assess their predictive performance for the 2022 FIFA World Cup and for the 2023 CAF Africa Cup of Nations by evaluating alternative summaries of past performances related to the involved teams. More specifically, we consider the Bayesian Bradley-Terry-Davidson model, which is a widely used statistical framework for ranking items based on paired comparisons that have been applied successfully in various domains, including football. The analysis was performed including in some canonical goal-based models both the Bradley-Terry-Davidson derived ranking and the widely recognized Coca-Cola FIFA ranking commonly adopted by football fans and amateurs.

stat.AP

Assessing replication success via skeptical mixture priors

There is a growing interest in the analysis of replication studies of original findings across many disciplines. When testing a hypothesis for an effect size, two Bayesian approaches stand out for their principled use of the Bayes factor (BF), namely the replication BF and the skeptical BF. In particular, the latter BF is based on the skeptical prior, which represents the opinion of an investigator who is unconvinced by the original findings and wants to challenge them. We embrace the skeptical perspective, and elaborate a novel mixture prior which incorporates skepticism while at the same time controlling for prior-data conflict within the original data. Consistency properties of the resulting skeptical mixture BF are provided together with an extensive analysis of the main features of our proposal. Finally, we apply our methodology to data from the Social Sciences Replication Project. In particular we show that, for some case studies where prior-data conflict is an issue, our method uses a more realistic prior and leads to evidence-classification for replication success which differs from the standard skeptical approach.

stat.ME

pivmet: Pivotal Methods for Bayesian Relabelling and k-Means Clustering

The identification of groups' prototypes, i.e. elements of a dataset that represent different groups of data points, may be relevant to the tasks of clustering, classification and mixture modeling. The R package pivmet presented in this paper includes different methods for extracting pivotal units from a dataset. One of the main applications of pivotal methods is a Markov Chain Monte Carlo (MCMC) relabelling procedure to solve the label switching in Bayesian estimation of mixture models. Each method returns posterior estimates, and a set of graphical tools for visualizing the output. The package offers JAGS and Stan sampling procedures for Gaussian mixtures, and allows for user-defined priors' parameters. The package also provides functions to perform consensus clustering based on pivotal units, which may allow to improve classical techniques (e.g. k-means) by means of a careful seeding. The paper provides examples of applications to both real and simulated datasets.

stat.CO

A Bayesian Quest for Finding a Unified Model for Predicting Volleyball Games

Volleyball is a team sport with unique and specific characteristics. We introduce a new two level-hierarchical Bayesian model which accounts for theses volleyball specific characteristics. In the first level, we model the set outcome with a simple logistic regression model. Conditionally on the winner of the set, in the second level, we use a truncated negative binomial distribution for the points earned by the loosing team. An additional Poisson distributed inflation component is introduced to model the extra points played in the case that the two teams have point difference less than two points. The number of points of the winner within each set is deterministically specified by the winner of the set and the points of the inflation component. The team specific abilities and the home effect are used as covariates on all layers of the model (set, point, and extra inflated points). The implementation of the proposed model on the Italian Superlega 2017/2018 data shows an exceptional reproducibility of the final league table and a satisfactory predictive ability.

stat.AP

Mendelian Randomization with Incomplete Exposure Data: a Bayesian Approach

We expand Mendelian Randomization (MR) methodology to deal with randomly missing data on either the exposure or the outcome variable, and furthermore with data from nonindependent individuals (eg components of a family). Our method rests on the Bayesian MR framework proposed by Berzuini et al (2018), which we apply in a study of multiplex Multiple Sclerosis (MS) Sardinian families to characterise the role of certain plasma proteins in MS causation. The method is robust to presence of pleiotropic effects in an unknown number of instruments, and is able to incorporate inter-individual kinship information. Introduction of missing data allows us to overcome the bias introduced by the (reverse) effect of treatment (in MS cases) on level of protein. From a substantive point of view, our study results confirm recent suspicion that an increase in circulating IL12A and STAT4 protein levels does not cause an increase in MS risk, as originally believed, suggesting that these two proteins may not be suitable drug targets for MS.

stat.AP

Bayesian Mendelian Randomization identifies disease causing proteins via pedigree data, partially observed exposures and correlated instruments

Background In a study performed on multiplex Multiple Sclerosis (MS) Sardinian families to identify disease causing plasma proteins, application of Mendelian Randomization (MR) methods encounters difficulties due to relatedness of individuals, correlation between finely mapped genotype instrumental variables (IVs) and presence of missing exposures. Method We specialize the method of Berzuini et al (2018) to deal with these difficulties. The proposed method allows pedigree structure to enter the specification of the outcome distribution via kinship matrix, and treating missing exposures as additional parameters to be estimated from the data. It also acknowledges possible correlation between instruments by replacing the originally proposed independence prior for IV-specific pleiotropic effect with a g-prior. Based on correlated (r2< 0.2) IVs, we analysed the data of four candidate MS-causing proteins by using both the independence and the g-prior. Results 95% credible intervals for causal effect for proteins IL12A and STAT4 lay within the strictly negative real semiaxis, in both analyses, suggesting potential causality. Those instruments whose estimated pleiotropic effect exceeded 85% of total effect on outcome were found to act in trans. Analysis via frequentist MR gave inconsistent results. Replacing the independence with a g-prior led to smaller credible intervals for causal effect. Conclusions Bayesian MR may be a good way to study disease causation at a protein level based on family data and moderately correlated instruments.

stat.AP

Combining historical data and bookmakers'odds in modelling football scores

Modelling football outcomes has gained increasing attention, in large part due to the potential for making substantial profits. Despite the strong connection existing between football models and the bookmakers' betting odds, no authors have used the latter for improving the fit and the predictive accuracy of these models. We have developed a hierarchical Bayesian Poisson model in which the scoring rates of the teams are convex combinations of parameters estimated from historical data and the additional source of the betting odds. We apply our analysis to a nine-year dataset of the most popular European leagues in order to predict match outcomes for their tenth seasons. In this paper, we provide numerical and graphical checks for our model.

stat.AP

Mixture Data-Dependent Priors

We propose a two-component mixture of a noninformative (diffuse) and an informative prior distribution, weighted through the data in such a way to prefer the first component if a prior-data conflict arises. The data-driven approach for computing the mixture weights makes this class data-dependent. Although rarely used with any theoretical motivation, data-dependent priors are often used for different reasons, and their use has been a lot debated over the last decades. However, our approach is justified in terms of Bayesian inference as an approximation of a hierarchical model and as a conditioning on a data statistic. This class of priors turns out to provide less information than an informative prior, perhaps it represents a suitable option for not dominating the inference in presence of small samples. First evidences from simulation studies show that this class could also be a good proposal for reducing mean squared errors.

stat.ME

Maxima Units Search (MUS) algorithm: methodology and applications

An algorithm for extracting identity submatrices of small rank and pivotal units from large and sparse matrices is proposed. The procedure has already been satisfactorily applied for solving the label switching problem in Bayesian mixture models. Here we introduce it on its own and explore possible applications in different contexts.

stat.CO

Relabelling in Bayesian mixture models by pivotal units

In this paper a simple procedure to deal with label switching when exploring complex posterior distributions by MCMC algorithms is proposed. Although it cannot be generalized to any situation, it may be handy in many applications because of its simplicity and very low computational burden. A possible area where it proves to be useful is when deriving a sample for the posterior distribution arising from finite mixture models when no simple or rational ordering between the components is available.

stat.CO