Searcharxiv⌕ Search

arXiv subjects

Jean-Michel Loubes

Publications and source records attributed to Jean-Michel Loubes.

At least 19 recordsLinked to original sources

On the Comparison of Optimizers for Imbalanced Learning

Data imbalance is pervasive in machine learning, from rare words and anomalies to underrepresented patterns in heterogeneous or cross-tabulated data. We study idealized optimizers geometries in continuous time to model small-step training in deep learning. We assume that the source of imbalance is unobserved: the optimizer has only access to the aggregate training loss ignoring the exact contributions of the majority and minority groups. In this setting, we characterize a region where majority losses are optimized regardless of the admissible minority structure. We derive explicit equations of this zone and bounds on the time needed to leave it. These bounds exhibit a milder dependence on minority amplitude for sign, spectral, and Newton descent than for Euclidean gradient descent. Experiments with AdamW and Muon suggest similar advantages over SGD across language, tabular, and image tasks.

stat.ML↗

Minimax Additive Regression under Unknown Dependent Designs

We study additive regression under an unknown and potentially non product design distribution, allowing the number of covariates to grow with the sample size. We consider a coupled class that separately controls the smoothness of the marginal densities and of each additive component multiplied by the corresponding marginal density. Under joint-density bounds that hold uniformly in the dimension and suitable dimension-growth conditions, we establish matching minimax bounds for prediction. With known marginal densities, the classical additive rate is attainable. When the marginals are unknown, this rate is preserved if the densities are at least as smooth as the weighted components. Otherwise, marginal-density smoothness determines the minimax rate over the coupled class. Finally, we recover all additive components with total squared error of the same order as the prediction error.

stat.ML↗

Generalized Functional ANOVA: A Complete Theoretical Framework

The functional ANOVA provides a fundamental representation of square-integrable multivariate functions into main effects and higher-order interactions. For independent inputs, the components belong to mutually orthogonal Hilbert subspaces and admit an explicit representation. For dependent inputs, however, the components are only hierarchically orthogonal: although existence and uniqueness results are available, the Hilbert subspaces underlying the generalized decomposition have remained implicit. We resolve this representation problem for continuous inputs supported on a bounded hyperrectangle whose joint density is bounded above and away from zero. We introduce a distribution-adapted family of functions and prove that it forms a Riesz basis of the $L^2$ space, thereby guaranteeing a unique, stable, and unconditionally convergent representation. We then show that, for every coalition of variables, the corresponding block of this basis exactly characterizes the Hilbert subspace containing the functional ANOVA component. Our construction recovers the classical orthogonal decomposition under input independence. As a direct consequence, computing the generalized functional ANOVA reduces to estimating coefficients in an explicit, distribution-adapted basis. Finally, as a \emph{proof of concept}, we introduce an elementary, fast and model-agnostic estimator based on our theoretical results. Experiments on synthetic and real-world datasets illustrate its connections with established tabular explanation methods and show that low order components often capture most of the signal in the model output.

stat.ML↗

When majority rules, minority loses: bias amplification of gradient descent

Despite growing empirical evidence of bias amplification in machine learning, its theoretical foundations remain poorly understood. We develop a formal framework for majority-minority learning tasks, showing how standard training can favor majority groups and produce stereotypical predictors that neglect minority-specific features. Assuming population and variance imbalance, our analysis reveals three key findings: (i) the close proximity between ``full-data'' and stereotypical predictors, (ii) the dominance of a region where training the entire model tends to merely learn the majority traits, and (iii) a lower bound on the additional training required. Our results are illustrated through experiments in deep learning for tabular and image classification tasks.

cs.LG↗

An Explainable GNN Framework for Component-Level Anomaly Diagnosis

Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in such systems is critical for reliability and safety, yet understanding their origin is equally important. Existing Graph Neural Network (GNN)based methods for anomaly detection primarily focus on sensor-level deviations and either attribute anomalies directly to the deviating sensors. When diagnosis is attempted, generally, the most deviated sensor is identified as a root cause of a system fault. However, in many industrial systems, anomalies do not arise from faulty sensors but from disruptions in the influences governing the system dynamics. We propose an explainable GNN-based anomaly detection framework that shifts the perspective from sensor-level anomalies to component-level diagnosis, hypothesizing that anomalous measurements are symptoms of altered inter-sensor influences. Experiments show that the method effectively identifies and prioritizes the true faulty components, providing interpretable insights into system failures.

cs.AI↗

Probing Cultural Signals in Large Language Models through Author Profiling

Large language models (LLMs) are increasingly deployed in applications with societal impact, raising concerns about the cultural biases they encode. We probe these representations by evaluating whether LLMs can perform author profiling from song lyrics in a zero-shot setting, inferring singers' gender and ethnicity without task-specific fine-tuning. Across several open-source models evaluated on more than 10,000 lyrics, we find that LLMs achieve non-trivial profiling performance but demonstrate systematic cultural alignment: most models default toward North American ethnicity, while DeepSeek-1.5B aligns more strongly with Asian ethnicity. This finding emerges from both the models' prediction distributions and an analysis of their generated rationales. To quantify these disparities, we introduce two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), and show that Ministral-8B displays the strongest ethnicity bias among the evaluated models, whereas Gemma-12B shows the most balanced behavior. Our code is available on [GitHub](https://github.com/ValentinLafargue/CulturalProbingLLM) and results on [HuggingFace](https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults).

cs.CL↗

Token-Efficient Change Detection in LLM APIs

Remote change detection in LLMs is a difficult problem. Existing methods are either too expensive for deployment at scale, or require initial white-box access to model weights or grey-box access to log probabilities. We aim to achieve both low cost and strict black-box operation, observing only output tokens. Our approach hinges on specific inputs we call Border Inputs, for which there exists more than one output top token. From a statistical perspective, optimal change detection depends on the model's Jacobian and the Fisher information of the output distribution. Analyzing these quantities in low-temperature regimes shows that border inputs enable powerful change detection tests. Building on this insight, we propose the Black-Box Border Input Tracking (B3IT) scheme. Extensive in-vivo and in-vitro experiments show that border inputs are easily found for non-reasoning tested endpoints, and achieve performance on par with the best available grey-box approaches. B3IT reduces costs by $30\times$ compared to existing methods, while operating in a strict black-box setting.

cs.LG↗

OT-FairBoost: Optimal Transport-Guided Gradient Boosting for Fairness Regularization on Tabular Data

Although neural-based machine learning models have received a lot of attention recently, tree-based models such as gradient boosting are competitive for tabular data and therefore remain widely used in various applications of AI. As when using other machine learning predictive models, they can however yield discriminative predictions across demographic groups, due to so-called algorithmic biases. These undesirable phenomena have motivated the emergence of new regulatory frameworks and various AI fairness strategies. While several pre-and post-processing methodologies exist to mitigate such bias on gradient boosting models, only a few in-processing methods have been proposed. To bridge this gap, we introduce OT-FairBoost, a novel in-processing framework that incorporates a Wasserstein-2 distance penalty directly into the objective function of gradient-boosted trees. This OT-based mitigation strategy has been shown to efficiently optimize group fairness criteria such as Demographic Parity and Equalized Odds on neural-based predictions. To adapt this approach for gradient boosting, we extend the sample-wise gradient estimation of the Wasserstein-2 distance between group predictions to discrete distributions and hessian diagonals. We then integrate our approach into the LightGBM training procedure and evaluate it across binary classification, regression, and multi-group sensitive attribute settings. Experimental results in each of these settings demonstrate that OT-FairBoost achieves best accuracy-fairness trade-offs against alternatives.

math.ST↗

Distributional Limit Theory for Optimal Transport

Optimal Transport (OT) is a resource allocation problem with applications in biology, data science, economics and statistics, among others. In some of the applications, practitioners have access to samples which approximate the continuous measure. Hence the quantities of interest derived from OT -- plans, maps and costs -- are only available in their empirical versions. Statistical inference on OT aims at finding confidence intervals of the population plans, maps and costs. In recent years this topic gained an increasing interest in the statistical community. In this paper we provide a comprehensive review of the most influential results on this research field, underlying the some of the applications. Finally, we provide a list of open problems.

math.ST↗

Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation

3D Morphable Models (3DMMs) remain the standard parametric shape priors for many state-of-the-art 3D face reconstruction algorithms. However, as these models are derived from a finite number of 3D face samples, they inherit the morphological biases of their training data, potentially limiting their generalizability across diverse global populations. In this paper, we propose a novel framework to analyze 3DMM reconstructions through the lens of surface curvature, with the objective to discover, quantify and visualize biases. While standard evaluation metrics often rely on Euclidean distances, our reconstruction error captures subtle surface nuances such as local topology or undulations. To do so, we leverage the Laplace-Beltrami Operator (LBO) to generate high-resolution curvature error maps, providing a localized and geometrically meaningful visualization of discrepancies between ground truth faces and reconstructed meshes. We derive from it an error metric that we validated through a user study, observing a significantly higher correlation to human perception compared to traditional methods. Furthermore, we conduct extensive experiments across several 3DMM bases and fitting algorithms, uncovering systematic age-related biases and providing preliminary evidence of biases associated with gender and ethnicity. Our findings highlight the necessity of adopting curvature-aware evaluation protocols to ensure demographic fairness and geometric precision in future 3D face reconstruction research.

cs.CV↗

Exposing the Illusion of Fairness: Auditing Vulnerabilities to Distributional Manipulation Attacks

The rapid deployment of AI systems in high-stakes domains, including those classified as high-risk under the The EU AI Act (Regulation (EU) 2024/1689), has intensified the need for reliable compliance auditing. For binary classifiers, regulatory risk assessment often relies on global fairness metrics such as the Disparate Impact ratio, widely used to evaluate potential discrimination. In typical auditing settings, the auditee provides a subset of its dataset to an auditor, while a supervisory authority may verify whether this subset is representative of the full underlying distribution. In this work, we investigate to what extent a malicious auditee can construct a fairness-compliant yet representative-looking sample from a non-compliant original distribution, thereby creating an illusion of fairness. We formalize this problem as a constrained distributional projection task and introduce mathematically grounded manipulation strategies based on entropic and optimal transport projections. These constructions characterize the minimal distributional shift required to satisfy fairness constraints. To counter such attacks, we formalize representativeness through distributional distance based statistical tests and systematically evaluate their ability to detect manipulated samples. Our analysis highlights the conditions under which fairness manipulation can remain statistically undetected and provides practical guidelines for strengthening supervisory verification. We validate our theoretical findings through experiments on standard tabular datasets for bias detection. Code is publicly available at https://github.com/ValentinLafargue/Inspection.

cs.LG↗

Exact Functional ANOVA Decomposition for Categorical Inputs Models

Functional ANOVA offers a principled framework for interpretability by decomposing a model's prediction into main effects and higher-order interactions. For independent features, this decomposition is well-defined, strongly linked with SHAP values, and serves as a cornerstone of additive explainability. However, the lack of an explicit closed-form expression for general dependent distributions has forced practitioners to rely on costly sampling-based approximations. We completely resolve this limitation for categorical inputs. By bridging functional analysis with the extension of discrete Fourier analysis, we derive a closed-form decomposition without any assumption. Our formulation is computationally very efficient. It seamlessly recovers the classical independent case and extends to arbitrary dependence structures, including distributions with non-rectangular support. Furthermore, leveraging the intrinsic link between SHAP and ANOVA under independence, our framework yields a natural generalization of SHAP values for the general categorical setting.

stat.ML↗

Minimax Private Estimation of Smooth Optimal-Transport Maps

We study the problem of estimating smooth optimal transport (OT) maps between two probability distributions under differential privacy (DP) constraints. Leveraging wavelet-based density estimators and recent stability bounds for smooth OT maps, we propose differentially private estimators that apply to both central and local DP models. Our main estimator achieves near-minimax optimal rates in dimension $d \geq 2$, and we complement it with a quantile-based estimator that attains minimax optimal rates in dimension $d = 1$ under central DP. We further establish matching minimax lower bounds, confirming the near-optimality of our approach. To the best of our knowledge, this constitutes the first differentially private procedure for OT map estimation with minimax optimality guarantees.

math.ST↗

Buzz, Choose, Forget: A Meta-Bandit Framework for Bee-Like Decision Making

This work introduces MAYA, a sequential imitation learning model based on multi-armed bandits, designed to reproduce and predict individual bees' decisions in contextualized foraging tasks. The model accounts for bees' limited memory through a temporal window $τ$, whose optimal value is around 7 trials, with a slight dependence on weather conditions. Experimental results on real, simulated, and complementary (mice) datasets show that MAYA (particularly with the Wasserstein distance) outperforms imitation baselines and classical statistical models, while providing interpretability of individual learning strategies and enabling the inference of realistic trajectories for prospective ecological applications.

cs.LG↗

An improved central limit theorem for the empirical sliced Wasserstein distance

Wasserstein distances are widely used in modern data analysis but pose significant computational and statistical challenges in high dimensions. The sliced Wasserstein distance alleviates these challenges by leveraging one-dimensional projections. Building on the Efron-Stein inequality-a technique proven effective in related problems-and a non-trivial control of the optimal transport potentials across directions, we establish a central limit theorem for the p-sliced Wasserstein distance, for p>1, centered at the expected empirical cost. Unlike for the general Wasserstein distance, the centering can be replaced by the population cost, enabling valid statistical inference. This generalizes and refines existing one-dimensional results, providing the first asymptotically valid inference framework for the sliced Wasserstein distance between possibly non-compact measures. Finally, we address other practical aspects crucial for inference, including Monte Carlo approximation of the slicing integral and consistent variance estimation.

math.ST↗

Evaluating Black-Box Vulnerabilities with Wasserstein-Constrained Data Perturbations

The growing use of Machine Learning (ML) tools comes with critical challenges, such as limited model explainability. We propose a global explainability framework that leverages Optimal Transport and Distributionally Robust Optimization to analyze how ML algorithms respond to constrained data perturbations. Our approach enforces constraints on feature-level statistics (e.g., brightness, age distribution), generating realistic perturbations that preserve semantic structure. We provide a model-agnostic diagnostic bench that applies to both tabular and image domains with solid theoretical guarantees. We validate the approach on real-world datasets providing interpretable robustness diagnostics that complement standard evaluation and fairness auditing tools.

cs.LG↗

Wasserstein Spatial Depth

Modeling observations as random distributions embedded within Wasserstein spaces is becoming increasingly popular across scientific fields, as it captures the variability and geometric structure of the data more effectively. However, the distinct geometry and unique properties of Wasserstein spaces pose challenges to the application of conventional statistical tools, which are primarily designed for Euclidean spaces. Consequently, adapting and developing new methodologies for analysis within Wasserstein spaces has become essential. The space of distributions on $\mathbb{R}^d$ with $d>1$ is not linear, and "mimic" the geometry of a Riemannian manifold. In this paper, we extend the concept of statistical depth to distribution-valued data, introducing the notion of Wasserstein spatial depth. This new measure provides a way to rank and order distributions, enabling the development of order-based clustering techniques and inferential tools. We show that Wasserstein spatial depth (WSD) preserves critical properties of conventional statistical depths, notably, ranging within $[0,1]$, transformation and geodesic invariance, vanishing at infinity, reaching a maximum at the geometric median, and continuity. Regarding robustness, we characterize the breakdown points of the empirical depth regions and the influence function of the WSD. Additionally, the population WSD has a straightforward plug-in estimator based on sampled empirical distributions. We establish the estimator's consistency and asymptotic normality. We also provide a two-sample test for populations of distributions based on the WSD. Finally, extensive simulations and a real-data application showcase the practical efficacy of the WSD.

math.ST↗

Fourier Analysis on the Boolean Hypercube via Hoeffding Functional Decomposition

Fourier analysis on the Boolean hypercube is fundamentally defined as the orthogonal decomposition of the space of pseudo-Boolean functions with respect to the uniform probability measure. In this work, we propose an ANOVA-based generalization of the Fourier decomposition on the Boolean hypercube endowed with any arbitrary probability measure. We provide an \emph{explicit} decomposition basis which generalizes the Walsh-Hadamard (or parity functions) basis under any \emph{arbitrary} probability measure on the Boolean hypercube. We formulate the computation of the entire functional decomposition as a least squares problem and also provide a method to address the classical \emph{curse of dimensionality} challenge. We provide a comprehensive generalization of Fourier analysis on the Boolean hypercube, enabling the handling of non-uniform configuration spaces inherent to real-world machine learning tasks, \textit{e.g.} when dealing with \emph{one-hot encoded} features. Finally, we demonstrate its practical impact in the field of explainable AI, by conducting comparative studies with feature attribution methods such as SHAP or TreeHFD.

stat.ML↗