SearcharxivSearch

arXiv subjects

Ernst C. Wit

Publications and source records attributed to Ernst C. Wit.

At least 19 recordsLinked to original sources

The Role of Uncertainty in Assessing the Fairness of Machine Learning Models

Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their outputs are biased against disadvantaged groups or individuals is crucial to ensuring they are fair and allowing their use in such settings. A rigorous risk assessment of possible fairness violations requires quantifying the uncertainty associated with selecting and estimating such models. Yet, this is rarely done in the literature, which focuses on identifying a single model with a suitable trade-off between predictive accuracy and fairness. In this paper, we move beyond point estimation and discuss frequentist and Bayesian approaches to uncertainty quantification for fair machine learning, with practical examples and implications for simulated and real data.

stat.ML

Introduction to Relational Event Modelling

Interactions and time shape many aspects of life. Everyday activities -- like conversations, emails, money transfers, citations, and even acts of violence -- are relational events: interactions between a sender and a receiver at a specific moment. At the intersection of event-history analysis and network modelling, relational event models (REMs) offer a powerful framework for studying when and why these events occur. Recent advances have made it possible to express REMs as generalized additive models, allowing researchers to capture complex, non-linear patterns over time. A hands-on guide to REMs is still missing. This tutorial fills that gap. It provides a practical introduction to REMs, incorporating the latest developments in the field. It demonstrates how to simulate synthetic relational-event data and walks through several empirical applications, comparing different modeling and inference strategies. By bringing together theory, simulation, and application, this tutorial lowers the barrier to entry and makes REMs a more accessible and practical tool.

stat.ME

Loglinear modelling of huge contingency tables

Contingency tables are the canonical representation of multivariate categorical data. As the size of the contingency table grows exponentially with the number of variables, even a moderate number of variables, each with a moderate number of levels, results in a huge number of cells, the majority of which remains empty even with a significant amount of data. We propose efficient methods for inferring higher-order loglinear models by performing subsampling on the set of the empty cells. First, we derive the likelihood under a zero-deflated Poisson sampling scheme. This is maximized via an efficient iteratively re-weighted least squares algorithm, leading to consistent and close to efficient estimators. This method works well for moderately sized contingency tables, but runs into computational instability when the number of dimensions grows. By sacrificing some efficiency, we show that nested case-control multinomial sampling combined with a degenerate logistic regression approach is also consistent and can be applied to arbitrarily large contingency tables. We illustrate the method with an analysis of data from the General Social Survey, which consists of $15014$ observations in a $69$-dimensional contingency table with a total of $6.6\times 10^{38}$ cells.

stat.ME

Inference of large scale relational state processes

Relational states refer to concepts such as friendship or collaboration, in which a relationship persists over a certain amount of time. Study of relational states often involves figuring out what factors contribute to the creation or dissolution of these relationships. However, most methods available now restrict their attention to binary states, i.e., ties that are either present or absent, even though many real-world systems evolve through multiple relational states (e.g., acquaintance, friendship, close friendship). We propose a continuous-time framework for modelling and inferring relational state networks in which each edge evolves by transitioning between two or more states. In our model, transition intensities are driven by state-dependent covariates that might be decomposed into anchoring (current-state) and pulling (target-state) mechanisms, with both linear and smooth non-linear effects. We address two common sampling regimes. With full event histories, a Cox-type partial likelihood with nested case-control sampling enables efficient estimation of both parametric and smooth effects. Instead, for panel data we derive a general ODE formulation for the likelihood, which leads to a particularly efficient inference procedure for binary state model. Simulation studies confirm accurate recovery of model parameters, and an empirical application to adolescent friendship data reproduces the substantive conclusions of established modelling techniques while offering substantial computational gains. The framework preserves the interpretability of classical network effects, generalizes them to multi-state ties, and scales to larger, more complex designs under both full-history and panel sampling designs.

stat.ME

Beyond Linearity and Time-Homogeneity: Relational Hyper Event Models with Time-Varying Non-Linear Effects

Recent technological advances have made it easier to collect large and complex networks of time-stamped relational events connecting two or more entities. Relational hyper-event models (RHEMs) aim to explain the dynamics of these events by modeling the event rate as a function of statistics based on past history and external information. However, despite the complexity of the data, most current RHEM approaches still rely on a linearity assumption to model this relationship. In this work, we address this limitation by introducing a more flexible model that allows the effects of statistics to vary non-linearly and over time. While time-varying and non-linear effects have been used in relational event modeling, we take this further by modeling joint time-varying and non-linear effects using tensor product smooths. We validate our methodology on both synthetic and empirical data. In particular, we use RHEMs to study how patterns of scientific collaboration and impact evolve over time. Our approach provides deeper insights into the dynamic factors driving relational hyper-events, allowing us to evaluate potential non-monotonic patterns that cannot be identified using linear models.

stat.ME

Some statistical aspects of the Covid-19 response

This paper discusses some statistical aspects of the U.K. Covid-19 pandemic response, focussing particularly on cases where we believe that a statistically questionable approach or presentation has had a substantial impact on public perception, or government policy, or both. We discuss the presentation of statistics relating to Covid risk, and the risk of the response measures, arguing that biases tended to operate in opposite directions, overplaying Covid risk and underplaying the response risks. We also discuss some issues around presentation of life loss data, excess deaths and the use of case data. The consequences of neglect of most individual variability from epidemic models, alongside the consequences of some other statistically important omissions are also covered. Finally the evidence for full stay at home lockdowns having been necessary to reverse waves of infection is examined, with new analyses provided for a number of European countries.

stat.AP

Hyperevent network modelling of partially observed gossip data

Gossiping is a widespread social phenomenon that shapes relationships and information flow in communities. From a network theoretic point of view, gossiping can be seen as a higher-order interaction, as it involves at least two persons talking about a non-present third. The mechanism of gossiping is complex: it is most likely dynamic, as its intensity changes over time, and possibly viral, if a gossiping event induces future gossiping, such as a repetition or retaliation. We define covariates of interest for these effects and propose a relational hyperevent model to study and quantify these complex dynamics. We consider survey data collected yearly from 44 secondary schools in Hungary. No information is available about the exact timing of the events nor about the aggregate number of events within the yearly time interval. What is measured is whether at least one gossiping event has occurred in a given time interval. We extend inference for relational hyperevent models to the case of rightcensored interval-time data and show how flexible and efficient generalized additive models can be used for estimation of effects of interest. Our analysis on the school data illustrates how a model that accounts for linear, smooth and random effects can identify the social drivers of gossiping, while revealing complex temporal dynamics.

stat.ME

Functional worst risk minimization

The aim of this paper is to extend worst risk minimization, also called worst average loss minimization, to the functional realm. This means finding a functional regression representation that will be robust to future distribution shifts on the basis of data from two environments. In the classical non-functional realm, structural equations are based on a transfer matrix $B$. In section~\ref{sec:sfr}, we generalize this to consider a linear operator $\mathcal{T}$ on square integrable processes that plays the the part of $B$. By requiring that $(I-\mathcal{T})^{-1}$ is bounded -- as opposed to $\mathcal{T}$ -- this will allow for a large class of unbounded operators to be considered. Section~\ref{sec:worstrisk} considers two separate cases that both lead to the same worst-risk decomposition. Remarkably, this decomposition has the same structure as in the non-functional case. We consider any operator $\mathcal{T}$ that makes $(I-\mathcal{T})^{-1}$ bounded and define the future shift set in terms of the covariance functions of the shifts. In section~\ref{sec:minimizer}, we prove a necessary and sufficient condition for existence of a minimizer to this worst risk in the space of square integrable kernels. Previously, such minimizers were expressed in terms of the unknown eigenfunctions of the target and covariate integral operators (see for instance \cite{HeMullerWang} and \cite{YaoAOS}). This means that in order to estimate the minimizer, one must first estimate these unknown eigenfunctions. In contrast, the solution provided here will be expressed in any arbitrary ON-basis. This completely removes any necessity of estimating eigenfunctions. This pays dividends in section~\ref{sec:estimation}, where we provide a family of estimators, that are consistent with a large sample bound. Proofs of all the results are provided in the appendix.

math.ST

Causal generalized linear models via Pearson risk invariance

Prediction invariance of causal models under heterogeneous settings has been exploited by a number of recent methods for causal discovery, typically focussing on recovering the causal parents of a target variable of interest. Existing methods require observational data from a number of sufficiently different environments, which is rarely available. In this paper, we consider a structural equation model where the target variable is described by a generalized linear model conditional on its parents. Besides having finite moments, no modelling assumptions are made on the conditional distributions of the other variables in the system, and nonlinear effects on the target variable can naturally be accommodated by a generalized additive structure. Under this setting, we characterize the causal model uniquely by means of two key properties: the Pearson risk invariant under the causal model and, conditional on the causal parents, the causal parameters maximize the expected likelihood. These two properties form the basis of a computational strategy for searching the causal model among all possible models. A stepwise greedy search is proposed for systems with a large number of variables. Crucially, for generalized linear models with a known dispersion parameter, such as Poisson and logistic regression, the causal model can be identified from a single data environment. The method is implemented in the R package causalreg.

stat.ME

Functional structural equation models with out-of-sample guarantees

Statistical learning methods typically assume that the training and test data originate from the same distribution, enabling effective risk minimization. However, real-world applications frequently involve distributional shifts, leading to poor model generalization. To address this, recent advances in causal inference and robust learning have introduced strategies such as invariant causal prediction and anchor regression. While these approaches have been explored for traditional structural equation models (SEMs), their extension to functional systems remains limited. This paper develops a risk minimization framework for functional SEMs using linear, potentially unbounded operators. We introduce a functional worst-risk minimization approach, ensuring robust predictive performance across shifted environments. Our key contribution is a novel worst-risk decomposition theorem, which expresses the maximum out-of-sample risk in terms of observed environments. We establish conditions for the existence and uniqueness of the worst-risk minimizer and provide consistent estimation procedures. Empirical results on functional systems illustrate the advantages of our method in mitigating distributional shifts. These findings contribute to the growing literature on robust functional regression and causal learning, offering practical guarantees for out-of-sample generalization in dynamic environments.

math.ST

Causal drivers of dynamic networks

Dynamic networks models describe temporal interactions between social actors, and as such have been used to describe financial fraudulent transactions, dispersion of destructive invasive species across the globe, and the spread of fake news. An important question in all of these examples is what are the causal drivers underlying these processes. Current network models are exclusively descriptive and based on correlative structures. In this paper we propose a causal extension of dynamic network modelling. In particular, we prove that the causal model satisfies a set of population conditions that uniquely identifies the causal drivers. The empirical analogue of these conditions provide a consistent causal discovery algorithm, which distinguishes it from other inferential approaches. Crucially, data from a single environment is sufficient. We apply the method in an analysis of bike sharing data in Washington D.C. in July 2023.

stat.ME

Relational event models with global covariates

Bike sharing is an increasingly popular mobility choice as it is a sustainable, healthy and economically viable transportation mode. By interpreting rides between bike stations over time as temporal events connecting two bike stations, relational event models can provide important insights into this phenomenon. The focus of relational event models, as a typical event history model, is normally on dyadic or node-specific covariates, as global covariates are considered nuisance parameters in a partial likelihood approach. As full likelihood approaches are infeasible given the sheer size of the relational process, we propose an innovative sampling approach of temporally shifted non-events to recover important global drivers of the relational process. The method combines nested case-control sampling on a time-shifted version of the event process. This leads to a partial likelihood of the relational event process that is identical to that of a degenerate logistic additive model, enabling efficient estimation of both global and non-global covariate effects. The computational effectiveness of the method is demonstrated through a simulation study. The analysis of around 350,000 bike rides in the Washington D.C. area reveals significant influences of weather and time of day on bike sharing dynamics, besides a number of traditional node-specific and dyadic covariates.

stat.ME

Inferring the dynamics of quasi-reaction systems via nonlinear local mean-field approximations

In the modelling of stochastic phenomena, such as quasi-reaction systems, parameter estimation of kinetic rates can be challenging, particularly when the time gap between consecutive measurements is large. Local linear approximation approaches account for the stochasticity in the system but fail to capture the nonlinear nature of the underlying process. At the mean level, the dynamics of the system can be described by a system of ODEs, which have an explicit solution only for simple unitary systems. An analytical solution for generic quasi-reaction systems is proposed via a first order Taylor approximation of the hazard rate. This allows a nonlinear forward prediction of the future dynamics given the current state of the system. Predictions and corresponding observations are embedded in a nonlinear least-squares approach for parameter estimation. The performance of the algorithm is compared to existing SDE and ODE-based methods via a simulation study. Besides the increased computational efficiency of the approach, the results show an improvement in the kinetic rate estimation, particularly for data observed at large time intervals. Additionally, the availability of an explicit solution makes the method robust to stiffness, which is often present in biological systems. An illustration on Rhesus Macaque data shows the applicability of the approach to the study of cell differentiation.

stat.ME

Bayesian Dynamic Generalized Additive Model for Mortality during COVID-19 Pandemic

While COVID-19 has resulted in a significant increase in global mortality rates, the impact of the pandemic on mortality from other causes remains uncertain. To gain insight into the broader effects of COVID-19 on various causes of death, we analyze an Italian dataset that includes monthly mortality counts for different causes from January 2015 to December 2020. Our approach involves a generalized additive model enhanced with correlated random effects. The generalized additive model component effectively captures non-linear relationships between various covariates and mortality rates, while the random effects are multivariate time series observations recorded in various locations, and they embody information on the dependence structure present among geographical locations and different causes of mortality. Adopting a Bayesian framework, we impose suitable priors on the model parameters. For efficient posterior computation, we employ variational inference, specifically for fixed effect coefficients and random effects, Gaussian variational approximation is assumed, which streamlines the analysis process. The optimisation is performed using a coordinate ascent variational inference algorithm and several computational strategies are implemented along the way to address the issues arising from the high dimensional nature of the data, providing accelerated and stabilised parameter estimation and statistical inference.

stat.AP

Worst-risk minimization in generalized structural equation models

We consider rather general structural equation models (SEMs) between a target and its covariates in several shifted environments. Given $k\in\mathbb{N}$ shifts we consider the set of shifts that are at most $γ$-times as strong as a given weighted linear combination of these $k$ shifts and the worst (quadratic) risk over this entire space. This worst risk has a nice decomposition which we refer to as the "worst risk decomposition". Then we find an explicit arg-min solution that minimizes the worst risk and consider its corresponding plug-in estimator which is the main object of this paper. This plug-in estimator is (almost surely) consistent and we first prove a concentration in measure result for it. The solution to the worst risk minimizer is rather reminiscent of the corresponding ordinary least squares solution in that it is product of a vector and an inverse of a Grammian matrix. Due to this, the central moments of the plug-in estimator is not well-defined in general, but we instead consider these moments conditioned on the Grammian inverse being bounded by some given constant. We also study conditional variance of the estimator with respect to a natural filtration for the incoming data. Similarly we consider the conditional covariance matrix with respect to this filtration and prove a bound for the determinant of this matrix. This SEM model generalizes the linear models that have been studied previously for instance in the setting of casual inference or anchor regression but the concentration in measure result and the moment bounds are new even in the linear setting.

math.ST

Constructive and consistent estimation of quadratic minimax

We consider $k$ square integrable random variables $Y_1,...,Y_k$ and $k$ random (row) vectors of length $p$, $X_1,...,X_k$ such that $X_i(l)$ is square integrable for $1\le i\le k$ and $1\le l\le p$. No assumptions whatsoever are made of any relationship between the $X_i$:s and $Y_i$:s. We shall refer to each pairing of $X_i$ and $Y_i$ as an environment. We form the square risk functions $R_i(β)=\mathbb{E}\left[(Y_i-βX_i)^2\right]$ for every environment and consider $m$ affine combinations of these $k$ risk functions. Next, we define a parameter space $Θ$ where we associate each point with a subset of the unique elements of the covariance matrix of $(X_i,Y_i)$ for an environment. Then we study estimation of the $\arg\min$-solution set of the maximum of a the $m$ affine combinations the of quadratic risk functions. We provide a constructive method for estimating the entire $\arg\min$-solution set which is consistent almost surely outside a zero set in $Θ^k$. This method is computationally expensive, since it involves solving polynomials of general degree. To overcome this, we define another approximate estimator that also provides a consistent estimation of the solution set based on the bisection method, which is computationally much more efficient. We apply the method to worst risk minimization in the setting of structural equation models.

math.ST

Nodal heterogeneity can induce ghost triadic effects in relational event models

Temporal network data is often encoded as time-stamped interaction events between senders and receivers, such as co-authoring scientific articles or communication via email. A number of relational event frameworks have been proposed to address specific issues raised by complex temporal dependencies. These models attempt to quantify how individual behaviour, endogenous and exogenous factors, as well as interactions with other individuals modify the network dynamics over time. It is often of interest to determine whether changes in the network can be attributed to endogenous mechanisms reflecting natural relational tendencies, such as reciprocity or triadic effects. The propensity to form or receive ties can also, at least partially, be related to actor attributes. Nodal heterogeneity in the network is often modelled by including actor-specific or dyadic covariates. However, comprehensively capturing all personality traits is difficult in practice, if not impossible. A failure to account for heterogeneity may confound the substantive effect of key variables of interest. This work shows that failing to account for node level sender and receiver effects can induce ghost triadic effects. We propose a random-effect extension of the relational event model to deal with these problems. We show that it is often effective over more traditional approaches, such as in-degree and out-degree statistics. These results that the violation of the hierarchy principle due to insufficient information about nodal heterogeneity can be resolved by including random effects in the relational event model as a standard.

stat.AP

Navigating Market Turbulence: Insights from Causal Network Contagion Value at Risk

Accurately defining, measuring and mitigating risk is a cornerstone of financial risk management, especially in the presence of financial contagion. Traditional correlation-based risk assessment methods often struggle under volatile market conditions, particularly in the face of external shocks, highlighting the need for a more robust and invariant predictive approach. This paper introduces the Causal Network Contagion Value at Risk (Causal-NECO VaR), a novel methodology that significantly advances causal inference in financial risk analysis. Embracing a causal network framework, this method adeptly captures and analyses volatility and spillover effects, effectively setting it apart from conventional contagion-based VaR models. Causal-NECO VaR's key innovation lies in its ability to derive directional influences among assets from observational data, thereby offering robust risk predictions that remain invariant to market shocks and systemic changes. A comprehensive simulation study and the application to the Forex market show the robustness of the method. Causal-NECO VaR not only demonstrates predictive accuracy, but also maintains its reliability in unstable financial environments, offering clearer risk assessments even amidst unforeseen market disturbances. This research makes a significant contribution to the field of risk management and financial stability, presenting a causal approach to the computation of VaR. It emphasises the model's superior resilience and invariant predictive power, essential for navigating the complexities of today's ever-evolving financial markets.

q-fin.RM