SearcharxivSearch

arXiv subjects

Rouven Michels

Publications and source records attributed to Rouven Michels.

9 recordsLinked to original sources

Predicting Qualification Thresholds in the UEFA Champions League and Europa League under the New League Phase Format

For the 2024/25 season, the Union of European Football Associations (UEFA) introduced an incomplete round-robin format in the Champions League and Europa League, replacing the traditional group stage with a single league table of all 36 teams. Under this structure, the top eight teams advance directly to the Round of 16, while teams ranked 9th-24th qualify for a play-off round. Simulation-based analyses, such as those by commercial data analyst Opta, provided indicative point thresholds for qualification but reveal deviations when compared with actual outcomes in the first season. To address these discrepancies, we employ a bivariate Dixon--Coles model that accounts for the lower frequency of draws observed in the 2024/25 and 2025/26 seasons, particularly in the Champions League, potentially driven by reduced incentives for teams to play for a draw. We proxy team strengths by Elo ratings and fit the model to different settings. This enables us to simulate match outcomes and to estimate qualification thresholds for both direct advancement and play-off participation. Our results provide scientific guidance for clubs and managers, supporting strategic decision-making under uncertainty regarding their progression prospects in the new format of UEFA club competitions.

econ.GN

Integrating Unsupervised and Supervised Learning for the Prediction of Defensive Schemes in American football

Anticipating defensive coverage schemes is a crucial yet challenging task for offenses in American football. Because defenders' assignments are intentionally disguised before the snap, they remain difficult to recognize in real time. To address this challenge, we develop a statistical framework that integrates supervised and unsupervised learning using player tracking data. Our goal is to forecast the defensive coverage scheme -- man or zone -- through elastic net logistic regression and gradient-boosted decision trees with incrementally derived features. We first use features from the pre-motion situation, then incorporate players' trajectories during motion in a naive way, and finally include features derived from a hidden Markov model (HMM). Based on player movements, the non-homogeneous HMM infers latent defensive assignments between offensive and defensive players during motion and transforms decoded state sequences into informative features for the supervised models. These HMM-based features enhance predictive performance and are significantly associated with coverage outcomes. Moreover, estimated random effects offer interpretable insights into how different defenses and positions adjust their coverage responsibilities.

stat.AP

Inference on the state process of periodically inhomogeneous hidden Markov models for animal behavior

Over the last decade, hidden Markov models (HMMs) have become increasingly popular in statistical ecology, where they constitute natural tools for studying animal behavior based on complex sensor data. Corresponding analyses sometimes explicitly focus on - and in any case need to take into account - periodic variation, for example by quantifying the activity distribution over the daily cycle or seasonal variation such as migratory behavior. For HMMs including periodic components, we establish important mathematical properties that allow for comprehensive statistical inference related to periodic variation, thereby also providing guidance for model building and model checking. Specifically, we derive the periodically varying unconditional state distribution as well as the time-varying and overall state dwell-time distributions - all of which are of key interest when the inferential focus lies on the dynamics of the state process. We use the associated novel inference and model-checking tools to investigate changes in the diel activity patterns of fruit flies in response to changing light conditions.

stat.ME

PEP: a tackle value measuring the prevention of expected points

Traditional assessments of tackling in American Football often only consider the number of tackles made, without adequately accounting for their context and importance for the game. Aiming for improvement, we develop a metric that quantifies the value of a tackle in terms of the prevented expected points (PEP). Specifically, we compare the real end-of-play yard line of tackles with the predicted yard line given the hypothetical situation that the tackle had been missed. For this, we use high-resolution tracking data, that capture the position and velocity of players, and a random forest to account for uncertainty and multi-modality in yard-line prediction. Moreover, we acknowledge the difference in the importance of tackles by assigning an expected points value to each individual tree prediction of the random forest. Finally, to relate the value of tackles to a player's ability to tackle, we fit a suitable mixed-effect model to the PEP values. Our approach contributes to a deeper understanding of defensive performances in American football and offers valuable insights for coaches and analysts.

stat.AP

Modelling handball outcomes using univariate and bivariate approaches

Handball has received growing interest during the last years, including academic research for many different aspects of the sport. On the other hand modelling the outcome of the game has attracted less interest mainly because of the additional challenges that occur. Data analysis has revealed that the number of goals scored by each team are under-dispersed relative to a Poisson distribution and hence new models are needed for this purpose. Here we propose to circumvent the problem by modelling the score difference. This removes the need for special models since typical models for integer data like the Skellam distribution can provide sufficient fit and thus reveal some of the characteristics of the game. In the present paper we propose some models starting from a Skellam regression model and also considering zero inflated versions as well as other discrete distributions in $\mathbb Z$. Furthermore, we develop some bivariate models using copulas to model the two halves of the game and thus providing insights on the game. Data from German Bundesliga are used to show the potential of the new models.

stat.ME

Extending the Dixon and Coles model: an application to women's football data

The prevalent model by Dixon and Coles (1997) extends the double Poisson model where two independent Poisson distributions model the number of goals scored by each team by moving probabilities between the scores 0-0, 0-1, 1-0, and 1-1. We show that this is a special case of a multiplicative model known as the Sarmanov family. Based on this family, we create more suitable models by moving probabilities between scores and employing other discrete distributions. We apply the new models to women's football scores, which exhibit some characteristics different than that of men's football.

stat.ME

Nonparametric estimation of multivariate hidden Markov models using tensor-product B-splines

For multivariate time series driven by underlying states, hidden Markov models (HMMs) constitute a powerful framework which can be flexibly tailored to the situation at hand. However, in practice it can be challenging to choose an adequate emission distribution for multivariate observation vectors. For example, the marginal data distribution may not immediately reveal the within-state distributional form, and also the different data streams may operate on different supports, rendering the common approach of using a multivariate normal distribution inadequate. Here we explore a nonparametric estimation of the emission distributions within a multivariate HMM based on tensor-product B-splines. In two simulation studies, we show the feasibility of our modelling approach and demonstrate potential pitfalls of inappropriate choices of parametric distributions. To illustrate the practical applicability, we present a case study where we use an HMM to model the bivariate time series comprising the lengths and angles of goalkeeper passes during the UEFA EURO 2020, investigating the effect of match dynamics on the teams' tactics.

stat.ME

Bettors' reaction to match dynamics -- Evidence from in-game betting

It is still largely unclear to what extent bettors update their prior assumptions about the strength and form of competing teams considering the dynamics during the match. This is of interest not only from the psychological perspective, but also as the pricing of live odds ideally should be driven both by the (objective) outcome probabilities and also the bettors' behaviour. Using state-space models (SSMs) to account for the dynamically evolving latent sentiment of the betting market, we analyse a unique high-frequency data set on stakes placed during the match. We find that stakes in the live-betting market are driven both by perceived pre-game strength and by in-game strength, the latter as measured by the Valuing Actions by Estimating Probabilities (VAEP) approach. Both effects vary over the course of the match.

stat.AP

The reaction to news in live betting

Sports betting markets have grown very rapidly recently, with the total European gambling market worth 98.6 billion euro in 2019. Considering a high-resolution (1 Hz) data set provided by a large European bookmaker, we investigate the effect of news on the dynamics of live betting. In particular, we consider stakes placed in a live betting market during football matches. Accounting for the general market activity level within a state-space modelling framework, we focus on the market's response to events such as goals (i.e. major news), but also to the general situation within a match such as the uncertainty about the outcome. Our results indicate that markets might overreact to recent news, confirming cognitive biases known from psychology and behavioural economics.

physics.soc-ph