SearcharxivSearch

arXiv subjects

Fulvia Mecatti

Publications and source records attributed to Fulvia Mecatti.

6 recordsLinked to original sources

Enhancing Gender Equality Assessment through Object-Oriented Bayesian networks: the European Gender Equality Index Case

A novel data-driven framework is introduced to assess gender equality by complementing and empowering a widely used European gender composite indicator, the Gender Equality Index (GEI). The GEI synthetizes the latent construct of gender equality into a single score and is extensively employed for cross-country comparison and monitoring. While effective for communication and benchmarking, this practice is affected by conceptual and methodological limitations, including marginal analysis that leaves interactions and conditional (in)dependencies unmeasured, and a lack of predictive capability. To address these limitations, this paper proposes the use of Object-Oriented Bayesian Networks (OOBNs) to model the GEI. By preserving the hierarchical structure of the index, OOBNs extend Bayesian Networks and enable a multivariate and probabilistic representation of interdependencies among the components of gender equality. This approach advances intersectional gender statistics by shifting the focus from computing a single composite score to modelling the underlying mechanisms that shape gender inequalities. The proposed methodology enhances the assessment and monitoring of gender equality and adds a predictive dimension through scenario-based evaluation, thereby supporting Gender Impact Assessment and policy decision-making. An application to Italian official statistics illustrates the practical relevance of the framework and its applicability to other national contexts and policy needs.

stat.AP

Bayesian Network Propensity Score to Evaluate Treatment Effects in Observational Studies

This paper focuses on the Bayesian Network Propensity Score (BNPS), a novel approach for estimating treatment effects in observational studies characterized by unknown (and likely unbalanced) designs and complex dependency structures among covariates. Traditional methods, such as logistic regression, often impose rigid parametric assumptions that may lead to misspecification errors, compromising causal inference. Recent classical and machine learning alternatives, such as boosted CART, random forests, and Stable Balancing Weights, seem to be attractive in a predictive perspective, but they typically lack asymptotic properties, such as consistency, efficiency, and valid variance estimation. In contrast, the recently proposed BNPS to estimate propensity scores uses Bayesian Networks to flexibly model conditional dependencies while preserving essential statistical properties such as consistency, asymptotic normality and asymptotic efficiency. Combined with the Hájek estimator, BNPS enables robust estimation of the Average Treatment Effect (ATE) in scenarios with strong covariate interactions and unknown data-generating mechanisms. Through extensive simulations across fifteen realistic scenarios and varying sample sizes, BNPS consistently outperforms benchmark methods in both empirical rejection rates and coverage accuracy. Finally, an application to a real-world dataset of 7,162 prostate cancer patients from San Raffaele Hospital (Milan, Italy) demonstrates BNPS's practical value in assessing the impact of pelvic lymph node dissection on hospitalization duration and biochemical recurrence. The findings support BNPS as a statistically robust, interpretable and transparent alternative for causal inference in complex observational settings, enhancing the reliability of evidence from real-world biomedical data.

stat.ME

Statistical Challenges in Analyzing Migrant Backgrounds Among University Students: a Case Study from Italy

The methodological issues and statistical complexities of analyzing university students with migrant backgrounds is explored, focusing on Italian data from the University of Milano-Bicocca. With the increasing size of migrant populations and the growth of the second and middle generations, the need has risen for deeper knowledge of the various strata of this population, including university students with migrant backgrounds. This presents challenges due to inconsistent recording in university datasets. By leveraging both administrative records and an original targeted survey we propose a methodology to fully identify the study population of students with migrant histories, and to distinguish relevant subpopulations within it such as second-generation born in Italy. Traditional logistic regression and machine learning random forest models are used and compared to predict migrant status. The primary contribution lies in creating an expanded administrative dataset enriched with indicators of students' migrant backgrounds and status. The expanded dataset provides a critical foundation for analyzing the characteristics of students with migration histories across all variables routinely registered in the administrative data set. Additionally, findings highlight the presence of selection bias in the targeted survey data, underscoring the need of further research.

stat.AP

Testing for causal effect for binary data when propensity scores are estimated through Bayesian Networks

This paper proposes a new statistical approach for assessing treatment effect using Bayesian Networks (BNs). The goal is to draw causal inferences from observational data with a binary outcome and discrete covariates. The BNs are here used to estimate the propensity score, which enables flexible modeling and ensures maximum likelihood properties, including asymptotic efficiency. %As a result, other available approaches cannot perform better. When the propensity score is estimated by BNs, two point estimators are considered - Hájek and Horvitz-Thompson - based on inverse probability weighting, and their main distributional properties are derived for constructing confidence intervals and testing hypotheses about the absence of the treatment effect. Empirical evidence is presented to show the goodness of the proposed methodology on a simulation study mimicking the characteristics of a real dataset of prostate cancer patients from Milan San Raffaele Hospital.

stat.ME

Sequential adaptive strategy for population-based sampling of a rare and clustered disease

An innovative sampling strategy is proposed, which applies to large-scale population-based surveys targeting a rare trait that is unevenly spread over a geographical area of interest. Our proposal is characterised by the ability to tailor the data collection to specific features and challenges of the survey at hand. It is based on integrating an adaptive component into a sequential selection, which aims to both intensify detection of positive cases, upon exploiting the spatial clusterisation, and provide a flexible framework for managing logistical and budget constraints. To account for the selection bias, a ready-to-implement weighting system is provided to release unbiased and accurate estimates. Empirical evidence is illustrated from tuberculosis prevalence surveys, which are recommended in many countries and supported by the WHO as an emblematic example of the need for an improved sampling design. Simulation results are also given to illustrate strengths and weaknesses of the proposed sampling strategy with respect to traditional cross-sectional sampling.

stat.ME

A unified principled framework for resampling based on pseudo-populations: asymptotic theory

In this paper, a class of resampling techniques for finite populations under complex sampling design is introduced. The basic idea on which it rests is a two-step procedure consisting in : (i) constructing a pseudo-population on the basis of sample data; (ii) drawing a sample from the predicted population according to an appropriate resampling design. From a logical point of view, this approach is essentially based on the plug-in principle by Efron, at the "sampling design level". Theoretical justifications based on large sample theory are provided. New approaches to construct pseudo-populations based on various forms of calibrations are proposed. Finally, a simulation study is performed.

stat.ME