SearcharxivSearch

arXiv subjects

Lea Multerer

Publications and source records attributed to Lea Multerer.

3 recordsLinked to original sources

Counterfactual Methods for Detecting Unfairness in Anti-Money Laundering Algorithms

The application of machine learning-based predictive algorithms to Anti-Money Laundering (AML) has grown rapidly, driven by the vast volume of financial transaction data available to banks. These algorithms are typically trained not only on transactional data but also on sensitive client information, which may raise fairness concerns. Despite this, AML detection systems remain largely underexplored from a fairness perspective, even though deeper analytical methods based on counterfactuals are now available. Such techniques enable the decomposition of the direct and indirect effects of potentially sensitive features on model predictions, thereby supporting the evaluation of whether their influence is acceptable from a fairness perspective. Closing this gap, we consider the synthetic IBM AMLSim transaction dataset and construct additional features of the country of an account and its average behaviour. This improves the predictive performance of diverse machine learning models, ranging from baseline decision trees to state-of-the-art graph neural networks. We assess the potential unfairness associated with these features through a counterfactual, path-specific effect analysis. This reveals that fairness violations tend to be more pronounced for models whose predictive performance benefits the most from the extended features. Such a finding highlights a concrete instance of the trade-off between predictive accuracy and fairness in AML applications, thus underscoring the urgency of a systematic fairness analysis in such critical domains.

cs.LG

Learning Dynamical Systems from Multiple Sparse Datasets: A Hierarchical Bayesian Modeling Approach

Estimating parameters of dynamical systems from sparse, noisy, and irregularly sampled data is often severely ill-conditioned. When multiple related datasets are available, they provide additional information if the shared structure and variability are properly modeled. We propose a hierarchical Bayesian framework for probabilistic meta-learning in dynamical systems, modeling dataset-specific parameters as draws from a shared population distribution. A numerical ODE solver is embedded within gradient-based MCMC to enable efficient posterior inference of the shared population and dataset-specific parameter distribution. Experiments show improved predictive performance over unpooled methods, highlighting the potential for data-efficient system identification in settings with sparse data.

cs.LG

A reduced model for the long-term effects of physical activity on type 2 diabetes

Type 2 diabetes progresses slowly and may be reversed through lifestyle changes, but quantifying the long-term impact of regular physical activity remains challenging due to sparse longitudinal data. Mechanistic models offer a powerful tool by simulating metabolic processes over extended timescales. However, multi-scale formulations that capture both the short-term effects of exercise sessions and the slow evolution of disease tend to be computationally demanding, limiting their practical use in personalized decision support. To address this limitation, we derived a reduced version of a two-scale model that captures the short- and long-term effects of physical activity on blood glucose regulation. By analytically averaging the short-term effects induced by exercise, we developed a homogenized formulation that transmits the average contribution of physical activity to the slower glucose-insulin dynamics. This reduction preserves the key model dynamics while decreasing computational complexity by almost a factor 2000. We prove that the approximation error remains bounded and confirm the model's accuracy through a parameter-based simulation study. The resulting model provides a mathematically grounded reduction that retains key physiological mechanisms while enabling fast long-term simulations. This substantial computational gain makes it suitable for integration into medical decision support systems, where it can be used to design and evaluate personalized physical activity plans aimed at reducing the risk of type 2 diabetes.

q-bio.QM