Searcharxiv⌕ Search

arXiv subjects

Valérie Chavez-Demoulin

Publications and source records attributed to Valérie Chavez-Demoulin.

At least 19 recordsLinked to original sources

Spatiotemporally Consistent Multivariate Bias Correction for Climate Projections via Nested Vine Copulas

Climate models are essential for understanding large-scale climate dynamics and long-term climate change, yet they exhibit systematic biases when compared with historical observations. Existing multivariate bias correction (MBC) approaches do not explicitly handle spatiotemporal dependence. However, preserving both spatiotemporal and inter-variable consistency is essential for realistic climate dynamics and reliable regional impact assessments. To address this gap, we propose a novel MBC method called GN-VBC that uses generalized additive models (GAMs) to disentangle spatiotemporal deterministic effects from stochastic residuals. To model joint distributions and dependencies across variables and locations, we introduce nested vine copulas (NVCs), a hierarchical vine merging strategy. NVC in the context of MBC combines two dependence levels: (i) spatial dependence across locations, modeled separately for each variable, and (ii) inter-variable dependence modeled at a selected reference location, which links the spatial models into a coherent multivariate and spatial structure. An application to Switzerland shows improvements in preserving inter-variable, spatial and temporal dependence across a wide range of evaluation metrics.

stat.ME↗

Identifiability of causal graphs under nonadditive conditionally parametric causal models

Existing approaches to causal discovery often rely on restrictive modeling assumptions that limit their applicability in real-world settings, particularly when data are heavy-tailed or contain a mixture of discrete and continuous variables. Identifiability of causal graphs has been established under several structural models, including linear non-Gaussian models, post-nonlinear models, and location-scale models. However, these frameworks may not capture the diversity of distributions observed in practice. To address this, we introduce Conditionally Parametric Causal Models (CPCM), a flexible class of models where the conditional distribution of the effect, given its cause, belongs to a known parametric family such as Gaussian, Poisson, Gamma, or Pareto. These models are adaptable to a wide range of practical situations, where the cause influences not only the mean but also the variance or tail behavior of the effect. We demonstrate the identifiability of CPCM by leveraging the concept of sufficient statistics. Furthermore, we propose an algorithm for estimating the causal structure from random samples drawn from CPCM. We evaluate the empirical properties of our methodology on various datasets, demonstrating state-of-the-art performance across multiple benchmarks.

stat.ME↗

SEMF: Supervised Expectation-Maximization Framework for Predicting Intervals

This work introduces the Supervised Expectation-Maximization Framework (SEMF), a versatile and model-agnostic approach for generating prediction intervals with any ML model. SEMF extends the Expectation-Maximization algorithm, traditionally used in unsupervised learning, to a supervised context, leveraging latent variable modeling for uncertainty estimation. Through extensive empirical evaluation of diverse simulated distributions and 11 real-world tabular datasets, SEMF consistently produces narrower prediction intervals while maintaining the desired coverage probability, outperforming traditional quantile regression methods. Furthermore, without using the quantile (pinball) loss, SEMF allows point predictors, including gradient-boosted trees and neural networks, to be calibrated with conformal quantile regression. The results indicate that SEMF enhances uncertainty quantification under diverse data distributions and is particularly effective for models that otherwise struggle with inherent uncertainty representation.

stat.ML↗

Structural restrictions in local causal discovery: identifying direct causes of a target variable

We consider the problem of learning a set of direct causes of a target variable from an observational joint distribution. Learning directed acyclic graphs (DAGs) that represent the causal structure is a fundamental problem in science. Several results are known when the full DAG is identifiable from the distribution, such as assuming a nonlinear Gaussian data-generating process. Here, we are only interested in identifying the direct causes of one target variable (local causal structure), not the full DAG. This allows us to relax the identifiability assumptions and develop possibly faster and more robust algorithms. In contrast to the Invariance Causal Prediction framework, we only assume that we observe one environment without any interventions. We discuss different assumptions for the data-generating process of the target variable under which the set of direct causes is identifiable from the distribution. While doing so, we put essentially no assumptions on the variables other than the target variable. In addition to the novel identifiability results, we provide two practical algorithms for estimating the direct causes from a finite random sample and demonstrate their effectiveness on several benchmark and real datasets.

stat.ME↗

Tail asymptotics and precise large deviations for some Poisson cluster processes

We study the tail asymptotics of two functionals (the maximum and the sum of the marks) of a generic cluster in two sub-models of the marked Poisson cluster process, namely the renewal Poisson cluster process and the Hawkes process. Under the hypothesis that the governing components of the processes are regularly varying, we extend results due to [18] and [5] notably, relying on Karamata's Tauberian Theorem to do so. We use these asymptotics to derive precise large deviation results in the fashion of [30] for the above-mentioned processes.

math.PR↗

Statistical Inference on the Miss Distance Compared to Collision Probability for Conjunction Analysis

Satellite conjunctions involving near misses of space objects are increasingly common, especially with the growth of satellite constellations and space debris. Accurate risk analysis for these events is essential to prevent collisions and manage space traffic. Traditional methods for assessing collision risk, such as calculating the so-called collision probability, are widely used but have limitations, including counterintuitive interpretations when uncertainty in the state vector is large. To address these limitations, we build on an alternative approach proposed by Elkantassi and Davison (2022) that uses a statistical model allowing inference on the miss distance between two objects in the presence of nuisance parameters. This model provides significance probabilities for a null hypothesis that assumes a small miss distance and allows the construction of confidence intervals, leading to another interpretation of collision risk. In this study, we compare this approach with the traditional use of pc across a large, NASA-provided dataset of real conjunctions, in order to evaluate its reliability and to refine the statistical framework to improve its suitability for operational decision-making. We also discuss constraints that could limit the practical use of such alternative approaches.

stat.AP↗

Causal Discovery in Multivariate Extremes with a Hydrological Analysis of Swiss River Discharges

Causal asymmetry is based on the principle that an event is a cause only if its absence would not have been a cause. From there, uncovering causal effects becomes a matter of comparing a well-defined score in both directions. Motivated by studying causal effects at extreme levels of a multivariate random vector, we propose to construct a model-agnostic causal score relying solely on the assumption of the existence of a max-domain of attraction. Based on a representation of a Generalized Pareto random vector, we construct the causal score as the Wasserstein distance between the margins and a well-specified random variable. The proposed methodology is illustrated on a hydrologically simulated dataset of different characteristics of catchments in Switzerland: discharge, precipitation, and snowmelt.

stat.ME↗

Causality and extremes

In this work, we summarize the state-of-the-art methods in causal inference for extremes. In a non-exhaustive way, we start by describing an extremal approach to quantile treatment effect where the treatment has an impact on the tail of the outcome. Then, we delve into two primary causal structures for extremes, offering in-depth insights into their identifiability. Additionally, we discuss causal structure learning in relation to these two models as well as in a model-agnostic framework. To illustrate the practicality of the approaches, we apply and compare these different methods using a Seine network dataset. This work concludes with a summary and outlines potential directions for future research.

stat.ME↗

Causal Modelling of Heavy-Tailed Variables and Confounders with Application to River Flow

Confounding variables are a recurrent challenge for causal discovery and inference. In many situations, complex causal mechanisms only manifest themselves in extreme events, or take simpler forms in the extremes. Stimulated by data on extreme river flows and precipitation, we introduce a new causal discovery methodology for heavy-tailed variables that allows the effect of a known potential confounder to be almost entirely removed when the variables have comparable tails, and also decreases it sufficiently to enable correct causal inference when the confounder has a heavier tail. We also introduce a new parametric estimator for the existing causal tail coefficient and a permutation test. Simulations show that the methods work well and the ideas are applied to the motivating dataset.

stat.ME↗

Detecting causal covariates for extreme dependence structures

Determining the causes of extreme events is a fundamental question in many scientific fields. An important aspect when modelling multivariate extremes is the tail dependence. In application, the extreme dependence structure may significantly depend on covariates. As for the general case of modelling including covariates, only some of the covariates are causal. In this paper, we propose a methodology to discover the causal covariates explaining the tail dependence structure between two variables. The proposed methodology for discovering causal variables is based on comparing observations from different environments or perturbations. It is a desired methodology for predicting extremal behaviour in a new, unobserved environment. The methodology is applied to a dataset of $\text{NO}_2$ concentration in the UK. Extreme $\text{NO}_2$ levels can cause severe health problems, and understanding the behaviour of concurrent severe levels is an important question. We focus on revealing causal predictors for the dependence between extreme $\text{NO}_2$ observations at different sites.

stat.ME↗

Linear prediction of point process times and marks

In this paper, we are interested in linear prediction of a particular kind of stochastic process, namely a marked temporal point process. The observations are event times recorded on the real line, with marks attached to each event. We show that in this case, linear prediction extends straightforwardly from the theory of prediction for stationary stochastic processes. Following classical lines, we derive a Wiener-Hopf-type integral equation to characterise the linear predictor, extending the "model independent origin" of the Hawkes process (Jaisson, 2015) as a corollary. We propose two recursive methods to solve the linear prediction problem and show that these are computationally efficient in known cases. The first solves the Wiener-Hopf equation via a set of differential equations. It is particularly well-adapted to autoregressive processes. In the second method, we develop an innovations algorithm tailored for moving-average processes. A small simulation study on two typical examples shows the application of numerical schemes for estimation of a Hawkes process intensity.

stat.ME↗

A competing risks interpretation of Hawkes processes

We give a construction of the Hawkes process as a piecewise competing risks model. We argue that the most natural interpretation of the self-excitation kernel is the hazard function of a defective random variable. This establishes a link between desired qualitative features of the process and a parametric form for the kernel, which we illustrate using examples from the literature. Two families of cure rate models taken from the survival analysis literature are proposed as new models for the self-excitation kernel. Finally, we show that the competing risks viewpoint leads to a general simulation algorithm which avoids inverting the compensator of the point process or performing an accept-reject step, and is therefore fast and quite general.

stat.ME↗

Distinguishing Cause from Effect Using Quantiles: Bivariate Quantile Causal Discovery

Causal inference using observational data is challenging, especially in the bivariate case. Through the minimum description length principle, we link the postulate of independence between the generating mechanisms of the cause and of the effect given the cause to quantile regression. Based on this theory, we develop Bivariate Quantile Causal Discovery (bQCD), a new method to distinguish cause from effect assuming no confounding, selection bias or feedback. Because it uses multiple quantile levels instead of the conditional mean only, bQCD is adaptive not only to additive, but also to multiplicative or even location-scale generating mechanisms. To illustrate the effectiveness of our approach, we perform an extensive empirical comparison on both synthetic and real datasets. This study shows that bQCD is robust across different implementations of the method (i.e., the quantile regression), computationally efficient, and compares favorably to state-of-the-art methods.

stat.ML↗

Modelling the Extremes of Seasonal Viruses and Hospital Congestion: The Example of Flu in a Swiss Hospital

Viruses causing flu or milder coronavirus colds are often referred to as "seasonal viruses" as they tend to subside in warmer months. In other words, meteorological conditions tend to impact the activity of viruses, and this information can be exploited for the operational management of hospitals. In this study, we use three years of daily data from one of the biggest hospitals in Switzerland and focus on modelling the extremes of hospital visits from patients showing flu-like symptoms and the number of positive cases of flu. We propose employing a discrete Generalized Pareto distribution for the number of positive and negative cases, and a Generalized Pareto distribution for the odds of positive cases. Our modelling framework allows for the parameters of these distributions to be linked to covariate effects, and for outlying observations to be dealt with via a robust estimation approach. Because meteorological conditions may vary over time, we use meteorological and not calendar variations to explain hospital charge extremes, and our empirical findings highlight their significance. We propose a measure of hospital congestion and a related tool to estimate the resulting CaRe (Charge-at-Risk-estimation) under different meteorological conditions. The relevant numerical computations can be easily carried out using the freely available GJRM R package. The introduced approach could be applied to several types of seasonal disease data such as those derived from the new virus SARS-CoV-2 and its COVID-19 disease which is at the moment wreaking havoc worldwide. The empirical effectiveness of the proposed method is assessed through a simulation study.

stat.ME↗

Causal mechanism of extreme river discharges in the upper Danube basin network

Extreme hydrological events in the Danube river basin may severely impact human populations, aquatic organisms, and economic activity. One often characterizes the joint structure of the extreme events using the theory of multivariate and spatial extremes and its asymptotically justified models. There is interest however in cascading extreme events and whether one event causes another. In this paper, we argue that an improved understanding of the mechanism underlying severe events is achieved by combining extreme value modelling and causal discovery. We construct a causal inference method relying on the notion of the Kolmogorov complexity of extreme conditional quantiles. Tail quantities are derived using multivariate extreme value models and causal-induced asymmetries in the data are explored through the minimum description length principle. Our CausEV, for Causality for Extreme Values, approach uncovers causal relations between summer extreme river discharges in the upper Danube basin and finds significant causal links between the Danube and its Alpine tributary Lech.

stat.AP↗

Intraday Retail Sales Forecast: An Efficient Algorithm for Quantile Additive Modeling

With the ever increasing prominence of data in retail operations, sales forecasting has become an essential pillar in the efficient management of inventories. When facing high demand, the use of backroom storage and intraday shelf replenishment is necessary to avoid stock-out. In that context, the mandatory input for any successful replenishment policy to be implemented is access to reliable forecasts for the sales at an intraday granularity. To that end, we use quantile regression to adapt different patterns from one product to the other, and we develop a stable and efficient quantile additive model algorithm to compute sales forecasts in an intradaily context. Our algorithm is computationally fast and is therefore suitable for use in real-time dynamic shelf replenishment. As an illustration, we examine the case of a highly frequented store, where the demand for various alimentary products is accurately estimated over the day with the help of the proposed algorithm.

stat.AP↗

Analyzing privacy-aware mobility behavior using the evolution of spatio-temporal entropy

Analyzing mobility behavior of users is extremely useful to create or improve existing services. Several research works have been done in order to study mobility behavior of users that mainly use users' significant locations. However, these existing analysis are extremely intrusive because they require the knowledge of the frequently visited places of users, which thus makes it fairly easy to identify them. Consequently, in this paper, we present a privacy-aware methodology to analyze mobility behavior of users. We firstly propose a new metric based on the well-known Shannon entropy, called spatio-temporal entropy, to quantify the mobility level of a user during a time window. Then, we compute a sequence of spatio-temporal entropy from the location history of the user that expresses user's movements as rhythms. We secondly present how to study the effects of several groups of additional variables on the evolution of the spatio-temporal entropy of a user, such as spatio-temporal, demographic and mean of transportation variables. For this, we use Generalized Additive Models (GAMs). The results firstly show that the spatio-temporal entropy and GAMs are an ideal combination to understand mobility behavior of an individual user or a group of users. We also evaluate the prediction accuracy of a global GAM compared to individual GAMs and individual AutoRegressive Integrated Moving Average (ARIMA) models. These last results highlighted that the global GAM gives more accurate predictions of spatio-temporal entropy by checking the Mean Absolute Error (MAE). In addition, this research work opens various threads, such as the prediction of demographic data of users or the creation of personalized mobility prediction models by using movement rhythm characteristics of a user.

cs.LG↗

Discovering demographic data of users from the evolution of their spatio-temporal entropy

Inferring information related to users enables to highly improve the quality of many mobile services. For example, knowing the demographic characteristics of a user allows a service to display more accurate information. According to the literature, various works present models to detect them but, to the best of our knowledge, no one is based on the use of the spatio-temporal entropy and introduces Generalized Additive models (GAMs) in this context to reach this goal. In this preliminary work, we present a new approach including these two key elements. The spatio-temporal entropy enables to capture the regularity of the mobility behavior of a user, while GAMs help to predict her demographic data based on several co-variables including the spatio-temporal entropy. The preliminary results are very encouraging to do future work since we obtain a prediction accuracy of 87% about the prediction of the working profile of users.

cs.SI↗