SearcharxivSearch

arXiv subjects

Lawrence Clegg

Publications and source records attributed to Lawrence Clegg.

3 recordsLinked to original sources

A market-calibrated accelerated failure time model for in-play football forecasting

In-play football forecasting models have struggled to match the accuracy of betting exchange prices, which aggregate information from many market participants. We close this gap by combining two extensions to a Weibull accelerated failure time model: calibrating team strength parameters to Betfair Exchange prices at kick-off to capture pre-match market information, and including post-shot expected goals as a time-varying covariate to capture in-play information. The calibration approach, where we jointly fit team-strength parameters to 1X2 and over/under betting markets via squared-error minimisation, applies to any intensity-based goal arrival model and enables stronger in-play forecasting. Evaluated across 140 English Premier League matches at minute intervals, the calibrated model almost matches Betfair's classification accuracy (70.2% versus 70.6%) while retaining interpretable team-level parameters and covariate effects. A comparison with two alternative continuous-time scoring models, both calibrated to the same pre-match odds, confirms that market calibration is the dominant driver of predictive accuracy. A betting simulation against Betfair in-play odds yields a 4.5% return on investment (Sharpe ratio 5.94) over 17,458 bets, suggesting an inefficiency within in-play football markets.

stat.AP

Capturing Intransitive Dominance in Tennis Forecasting: A Graph Neural Network Approach

Intransitive player dominance, where player A beats B, B beats C, but C beats A, is common in competitive tennis. Yet, there are few known attempts to incorporate it within forecasting methods. We address this problem with a graph neural network approach that explicitly models these intransitive relationships through temporal directed graphs, with players as nodes and their historical match outcomes as directed edges. Our model (65.7% accuracy, 0.214 Brier score) forecasts competitively with established rating systems such as Weighted Elo. Although it does not improve on the baseline in unconditional accuracy, a forecast-encompassing test shows that it carries complementary information. A combined forecast significantly outperforms Weighted Elo, and there is some indication that the gain grows more strongly on the intransitive matchups our model targets. A graph-based representation of player interactions thus captures a forecasting signal that transitive rating systems discard, even between players who share no common opponents.

cs.LG

Not feeling the buzz: Correction study of mispricing and inefficiency in online sportsbooks

We present a replication and correction of a recent article (Ramirez, P., Reade, J.J., Singleton, C., Betting on a buzz: Mispricing and inefficiency in online sportsbooks, International Journal of Forecasting, 39:3, 2023, pp. 1413-1423, doi: 10.1016/j.ijforecast.2022.07.011). RRS measure profile page views on Wikipedia to generate a "buzz factor" metric for tennis players and show that it can be used to form a profitable gambling strategy by predicting bookmaker mispricing. Here, we use the same dataset as RRS to reproduce their results exactly, thus confirming the robustness of their mispricing claim. However, we discover that the published betting results are significantly affected by a single bet (the "Hercog" bet), which returns substantial outlier profits based on erroneously long odds. When this data quality issue is resolved, the majority of reported profits disappear and only one strategy, which bets on "competitive" matches, remains significantly profitable in the original out-of-sample period. While one profitable strategy offers weaker support than the original study, it still provides an indication that market inefficiencies may exist, as originally claimed by RRS. As an extension, we continue backtesting after 2020 on a cleaned dataset. Results show that (a) the "competitive" strategy generates no further profits, potentially suggesting markets have become more efficient, and (b) model coefficients estimated over this more recent period are no longer reliable predictors of bookmaker mispricing. We present this work as a case study demonstrating the importance of replication studies in sports forecasting, and the necessity to clean data. We open-source release comprehensive datasets and code.

stat.AP