SearcharxivSearch

arXiv subjects

Manuele Leonelli

Publications and source records attributed to Manuele Leonelli.

At least 19 recordsLinked to original sources

How exceptional was the Big Three era? Extremes and persistence in men's professional tennis

Three players won 66 of the 81 Grand Slam titles contested between 2003 and 2023, and their era is widely held to be the most dominant in the history of tennis. Assessing it means comparing players who never met, so that every comparison passes through the opponents each did face. We measure dominance by how far a player stands above the field of his own day, how many stand there at once, and how long they remain, the last through averages over fixed windows rather than through unbroken runs, which are unstable when the strengths beneath them are themselves estimated. Strengths come from a Bayesian dynamic Bradley-Terry state space model, and dominance is characterized through exceedances of a high threshold. Applied to 197,926 matches played between 1968 and 2025, only Djokovic separates from the field on peak strength, Federer and Nadal being indistinguishable from Borg, McEnroe and Lendl. What distinguishes the recent period is not how high its best players stood but how long three of them stood there together. At its height three players held a top-three position in at least four fifths of all weeks, where in the strongest comparable decade of the past those positions were shared among four. The sharper anomaly lies between: for twelve years from 1990 the upper tail was not merely unoccupied but compressed.

stat.AP

Tail Dependence in EU Carbon Markets: Graphical Models of Extremes for EUA Futures

Understanding how extreme price movements propagate across financial and energy markets is critical for risk management and regulatory design in the EU Emissions Trading System (EU ETS). We apply H\"{u}sler-Reiss graphical models of extremes to a system of 20 daily variables centred on EU allowances futures across Phases 3 and 4 of the EU ETS (2013--2025), with a Gaussian graphical model as the average-dependence baseline. The tail networks are structurally distinct from the average dependence network: substantially denser, organized around different central nodes, and governed by within-sector homophily that binds sector boundaries more tightly than at the average-dependence level. EU allowances futures are peripheral in the standard graphical model but achieve the highest centrality in the tail networks, while equity indices and major FX pairs follow the opposite trajectory. Exponential random graph models confirm equity and FX peripherality in tail networks across all sample periods and identify triadic closure during market downturns as a Phase~3 phenomenon that vanishes in Phase~4. The phase transition restructures the tail network without thinning it: average dependence contracts sharply while tail dependence persists, and crash contagion shifts from clustered to diffuse propagation. These findings have direct implications for hedge construction by compliance entities, stress-test calibration by regulators, and the design of systemic-risk monitoring tools for EU ETS markets.

stat.AP

Crashing Together, Rallying Apart: Dynamic Conditional Tail Dependence in Cryptocurrency Markets

Cryptocurrency markets are prone to violent, synchronised drawdowns, challenging the claim that a basket of crypto-assets offers genuine internal diversification. Because standard covariance-based metrics fail to capture asymptotic tail dependence, they systematically understate systemic risk and overstate diversification benefits precisely when markets crash. This study maps the conditional dependence structure of the cryptocurrency market directly in the joint tails, isolating direct extremal linkages from those mediated by the rest of the system. We analyse the daily returns of the thirteen largest cryptocurrencies over a sequence of 89 overlapping windows spanning late 2021 to 2025. We apply dynamic H\"usler-Reiss graphical models of extremes, estimated separately for joint crashes and rallies, and benchmark them against a Gaussian graphical model of ordinary co-movement. The results reveal a near-complete and stable lower-tail graph, an upper tail that thins over time to re-form sectoral structures, and the dissolution of ordinary token categories into a single block anchored by a Bitcoin-Ethereum core. These findings imply that intra-crypto diversification fails on the downside, standard risk models underestimate market-wide crash probabilities by roughly eight-fold, and dynamic extremal graphs offer a superior tool for systemic risk monitoring.

q-fin.ST

Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation

Large language models are rapidly becoming infrastructural components in high-stakes institutional settings, including public administration, legal reasoning, and healthcare, where opacity is not merely inconvenient but institutionally and legally untenable. Existing approaches to explainability are predominantly post-hoc, offering unstable, non-contestable accounts that have no formal relationship to the reasoning process that produced the output. We argue that the problem is not the absence of explanation but the absence of structured reasoning in the first place. This paper makes the case for a fundamentally different architecture, which we call the Glassbox Framework, in which Bayesian networks serve as transparent, ante-hoc mediation layers for generative models. Bayesian networks encode domain knowledge, causal assumptions, and probabilistic dependencies before inference occurs, enabling auditable reasoning traces, uncertainty quantification, and contestable outputs. We characterise the architecture of this framework and ground it in a benefit eligibility scenario, identifying the foundational challenges spanning semantic alignment, dynamic model construction, probabilistic grounding, and human governance that must be solved to realise it at scale. By shifting from post-hoc explanation to ante-hoc probabilistic mediation, this work outlines a principled path toward AI systems that are not only powerful but fundamentally accountable.

cs.AI

Estimating Staged Event Tree Models via Hierarchical Clustering on the Simplex

Staged tree models enhance Bayesian networks by incorporating context-specific dependencies through a stage-based structure. In this study, we present a new framework for estimating staged trees using hierarchical clustering on the probability simplex, utilizing simplex basesd divergences. We conduct a thorough evaluation of several distance and divergence metrics including Total Variation, Hellinger, Fisher, and Kaniadakis; alongside various linkage methods such as Ward.D2, average, complete, and McQuitty. We conducted the simulation experiments that reveals Total Variation, especially when combined with Ward.D2 linkage, consistently produces staged trees with better model fit, structure recovery, and computational efficiency. We assess performance by utilizing relative Bayesian Information Criterion (BIC), and Hamming distance. Our findings indicate that although Backward Hill Climbing (BHC) delivers competitive outcomes, it incurs a significantly higher computational cost. On the other, Total Variation divergence with Ward.D2 linkage, achieves similar performance while providing significantly better computational efficiency, making it a more viable option for large-scale or time sensitive tasks.

stat.ML

Bayesian Causal Effect Estimation for Categorical Data using Staged Tree Models

We propose a fully Bayesian approach for causal inference with multivariate categorical data based on staged tree models, a class of probabilistic graphical models capable of representing asymmetric and context-specific dependencies. To account for uncertainty in both structure and parameters, we introduce a flexible family of prior distributions over staged trees. These include product partition models to encourage parsimony, a novel distance-based prior to promote interpretable dependence patterns, and an extension that incorporates continuous covariates into the learning process. Posterior inference is achieved via a tailored Markov Chain Monte Carlo algorithm with split-and-merge moves, yielding posterior samples of staged trees from which average treatment effects and uncertainty measures are derived. Posterior summaries and uncertainty measures are obtained via techniques from the Bayesian nonparametrics literature. Two case studies on electronic fetal monitoring and cesarean delivery and on anthracycline therapy and cardiac dysfunction in breast cancer illustrate the methods.

stat.ME

Staged Event Trees for Transparent Treatment Effect Estimation

Average and conditional treatment effects are fundamental causal quantities used to evaluate the effectiveness of treatments in various critical applications, including clinical settings and policy-making. Beyond the gold-standard estimators from randomized trials, numerous methods have been proposed to estimate treatment effects using observational data. In this paper, we provide a novel characterization of widely used causal inference techniques within the framework of staged event trees, demonstrating their capacity to enhance treatment effect estimation. These models offer a distinct advantage due to their interpretability, making them particularly valuable for practical applications. We implement classical estimators within the framework of staged event trees and illustrate their capabilities through both simulation studies and real-world applications. Furthermore, we showcase how staged event trees explicitly and visually describe when standard causal assumptions, such as positivity, hold, further enhancing their practical utility.

stat.ME

Modeling Psychological Profiles in Volleyball via Mixed-Type Bayesian Networks

Psychological attributes rarely operate in isolation: coaches reason about networks of related traits. We analyze a new dataset of 164 female volleyball players from Italy's C and D leagues that combines standardized psychological profiling with background information. To learn directed relationships among mixed-type variables (ordinal questionnaire scores, categorical demographics, continuous indicators), we introduce latent MMHC, a hybrid structure learner that couples a latent Gaussian copula and a constraint-based skeleton with a constrained score-based refinement to return a single DAG. We also study a bootstrap-aggregated variant for stability. In simulations spanning sample size, sparsity, and dimension, latent Max-Min Hill-Climbing (MMHC) attains lower structural Hamming distance and higher edge recall than recent copula-based learners while maintaining high specificity. Applied to volleyball, the learned network organizes mental skills around goal setting and self-confidence, with emotional arousal linking motivation and anxiety, and locates Big-Five traits (notably neuroticism and extraversion) upstream of skill clusters. Scenario analyses quantify how improvements in specific skills propagate through the network to shift preparation, confidence, and self-esteem. The approach provides an interpretable, data-driven framework for profiling psychological traits in sport and for decision support in athlete development.

cs.LG

Will AI Take My Job? Evolving Perceptions of Automation and Labor Risk in Latin America

As artificial intelligence and robotics increasingly reshape the global labor market, understanding public perceptions of these technologies becomes critical. We examine how these perceptions have evolved across Latin America, using survey data from the 2017, 2018, 2020, and 2023 waves of the Latinobarómetro. Drawing on responses from over 48,000 individuals across 16 countries, we analyze fear of job loss due to artificial intelligence and robotics. Using statistical modeling and latent class analysis, we identify key structural and ideological predictors of concern, with education level and political orientation emerging as the most consistent drivers. Our findings reveal substantial temporal and cross-country variation, with a notable peak in fear during 2018 and distinct attitudinal profiles emerging from latent segmentation. These results offer new insights into the social and structural dimensions of AI anxiety in emerging economies and contribute to a broader understanding of public attitudes toward automation beyond the Global North.

cs.CY

Understanding support for AI regulation: A Bayesian network perspective

As artificial intelligence (AI) becomes increasingly embedded in public and private life, understanding how citizens perceive its risks, benefits, and regulatory needs is essential. To inform ongoing regulatory efforts such as the European Union's proposed AI Act, this study models public attitudes using Bayesian networks learned from the nationally representative 2023 German survey Current Questions on AI. The survey includes variables on AI interest, exposure, perceived threats and opportunities, awareness of EU regulation, and support for legal restrictions, along with key demographic and political indicators. We estimate probabilistic models that reveal how personal engagement and techno-optimism shape public perceptions, and how political orientation and age influence regulatory attitudes. Sobol indices and conditional inference identify belief patterns and scenario-specific responses across population profiles. We show that awareness of regulation is driven by information-seeking behavior, while support for legal requirements depends strongly on perceived policy adequacy and political alignment. Our approach offers a transparent, data-driven framework for identifying which public segments are most responsive to AI policy initiatives, providing insights to inform risk communication and governance strategies. We illustrate this through a focused analysis of support for AI regulation, quantifying the influence of political ideology, perceived risks, and regulatory awareness under different scenarios.

cs.CY

Disentangling Spatial and Structural Drivers of Housing Prices through Bayesian Networks: A Case Study of Madrid, Barcelona, and Valencia

Understanding how housing prices respond to spatial accessibility, structural attributes, and typological distinctions is central to contemporary urban research and policy. In cities marked by affordability stress and market segmentation, models that offer both predictive capability and interpretive clarity are increasingly needed. This study applies discrete Bayesian networks to model residential price formation across Madrid, Barcelona, and Valencia using over 180,000 geo-referenced housing listings. The resulting probabilistic structures reveal distinct city-specific logics. Madrid exhibits amenity-driven stratification, Barcelona emphasizes typology and classification, while Valencia is shaped by spatial and structural fundamentals. By enabling joint inference, scenario simulation, and sensitivity analysis within a transparent framework, the approach advances housing analytics toward models that are not only accurate but actionable, interpretable, and aligned with the demands of equitable urban governance.

stat.AP

Uncovering Drivers of EU Carbon Futures with Bayesian Networks

The European Union Emissions Trading System (EU ETS) is a key policy tool for reducing greenhouse gas emissions and advancing toward a net-zero economy. Under this scheme, tradeable carbon credits, European Union Allowances (EUAs), are issued to large emitters, who can buy and sell them on regulated markets. We investigate the influence of financial, economic, and energy-related factors on EUA futures prices using discrete and dynamic Bayesian networks to model both contemporaneous and time-lagged dependencies. The analysis is based on daily data spanning the third and fourth ETS trading phases (2013-2025), incorporating a wide range of indicators including energy commodities, equity indices, exchange rates, and bond markets. Results reveal that EUA pricing is most influenced by energy commodities, especially coal and oil futures, and by the performance of the European energy sector. Broader market sentiment, captured through stock indices and volatility measures, affects EUA prices indirectly via changes in energy demand. The dynamic model confirms a modest next-day predictive influence from oil markets, while most other effects remain contemporaneous. These insights offer regulators, institutional investors, and firms subject to ETS compliance a clearer understanding of the interconnected forces shaping the carbon market, supporting more effective hedging, investment strategies, and policy design.

stat.AP

Predicting and understanding shooting performance in professional biathlon: A Bayesian approach

Biathlon is a unique winter sport that combines precision rifle marksmanship with the endurance demands of cross-country skiing. We develop a Bayesian hierarchical model to predict and understand shooting performance using data from the 2021/22 Women's World Cup season. The model captures athlete-specific, position-specific, race-type, and stage-dependent effects, providing a comprehensive view of shooting accuracy variability. By incorporating dynamic components, we reveal how performance evolves over the season, with model validation showing strong predictive ability at both overall and individual levels. Our findings highlight substantial athlete-specific differences and underscore the value of personalized performance analysis for optimizing coaching strategies. This work demonstrates the potential of advanced Bayesian modeling in sports analytics, paving the way for future research in biathlon and similar sports requiring the integration of technical and endurance skills.

stat.AP

bnRep: A repository of Bayesian networks from the academic literature

Bayesian networks (BNs) are widely used for modeling complex systems with uncertainty, yet repositories of pre-built BNs remain limited. This paper introduces bnRep, an open-source R package offering a comprehensive collection of documented BNs, facilitating benchmarking, replicability, and education. With over 200 networks from academic publications, bnRep integrates seamlessly with bnlearn and other R packages, providing users with interactive tools for network exploration.

cs.AI

The diameter of a stochastic matrix: A new measure for sensitivity analysis in Bayesian networks

Bayesian networks are one of the most widely used classes of probabilistic models for risk management and decision support because of their interpretability and flexibility in including heterogeneous pieces of information. In any applied modelling, it is critical to assess how robust the inferences on certain target variables are to changes in the model. In Bayesian networks, these analyses fall under the umbrella of sensitivity analysis, which is most commonly carried out by quantifying dissimilarities using Kullback-Leibler information measures. In this paper, we argue that robustness methods based instead on the familiar total variation distance provide simple and more valuable bounds on robustness to misspecification, which are both formally justifiable and transparent. We introduce a novel measure of dependence in conditional probability tables called the diameter to derive such bounds. This measure quantifies the strength of dependence between a variable and its parents. We demonstrate how such formal robustness considerations can be embedded in building a Bayesian network.

stat.ME

Global Sensitivity Analysis of Uncertain Parameters in Bayesian Networks

Traditionally, the sensitivity analysis of a Bayesian network studies the impact of individually modifying the entries of its conditional probability tables in a one-at-a-time (OAT) fashion. However, this approach fails to give a comprehensive account of each inputs' relevance, since simultaneous perturbations in two or more parameters often entail higher-order effects that cannot be captured by an OAT analysis. We propose to conduct global variance-based sensitivity analysis instead, whereby $n$ parameters are viewed as uncertain at once and their importance is assessed jointly. Our method works by encoding the uncertainties as $n$ additional variables of the network. To prevent the curse of dimensionality while adding these dimensions, we use low-rank tensor decomposition to break down the new potentials into smaller factors. Last, we apply the method of Sobol to the resulting network to obtain $n$ global sensitivity indices. Using a benchmark array of both expert-elicited and learned Bayesian networks, we demonstrate that the Sobol indices can significantly differ from the OAT indices, thus revealing the true influence of uncertain parameters and their interactions.

cs.AI

Context-Specific Refinements of Bayesian Network Classifiers

Supervised classification is one of the most ubiquitous tasks in machine learning. Generative classifiers based on Bayesian networks are often used because of their interpretability and competitive accuracy. The widely used naive and TAN classifiers are specific instances of Bayesian network classifiers with a constrained underlying graph. This paper introduces novel classes of generative classifiers extending TAN and other famous types of Bayesian network classifiers. Our approach is based on staged tree models, which extend Bayesian networks by allowing for complex, context-specific patterns of dependence. We formally study the relationship between our novel classes of classifiers and Bayesian networks. We introduce and implement data-driven learning routines for our models and investigate their accuracy in an extensive computational study. The study demonstrates that models embedding asymmetric information can enhance classification accuracy.

stat.ML

Learning Staged Trees from Incomplete Data

Staged trees are probabilistic graphical models capable of representing any class of non-symmetric independence via a coloring of its vertices. Several structural learning routines have been defined and implemented to learn staged trees from data, under the frequentist or Bayesian paradigm. They assume a data set has been observed fully and, in practice, observations with missing entries are either dropped or imputed before learning the model. Here, we introduce the first algorithms for staged trees that handle missingness within the learning of the model. To this end, we characterize the likelihood of staged tree models in the presence of missing data and discuss pseudo-likelihoods that approximate it. A structural expectation-maximization algorithm estimating the model directly from the full likelihood is also implemented and evaluated. A computational experiment showcases the performance of the novel learning algorithms, demonstrating that it is feasible to account for different missingness patterns when learning staged trees.

stat.ML