SearcharxivSearch

arXiv subjects

Yannik Pitcan

Publications and source records attributed to Yannik Pitcan.

4 recordsLinked to original sources

Does a Structural Model Add Anything to the Closing Price? Calibrated forecasting, incremental information, and match leverage in the Italian Serie A

Studies of association-football forecasting routinely report three-way accuracy in the low fifties and present it as competitive with the betting market. Accuracy against a uniform benchmark answers the wrong question; the question worth asking is whether a model carries information a margin-free closing price has not already absorbed. We formalise that test as the fitted weight in a logarithmic opinion pool and apply it to nineteen complete Serie A seasons (7,220 matches). The answer is negative and stable. A Dixon-Coles model with tuned exponential decay attains 53.4% accuracy and a Ranked Probability Score of 0.1972 against the market's 0.1905; the paired difference is +0.0067 (95% CI [0.0046, 0.0088]) and the market wins in all seven test seasons. The fitted pooling weight on the structural model is 0.000, and the log-loss profile is monotone increasing in that weight on validation and test alike, so this is a boundary solution, not an optimisation artefact. Refitting the same machinery to shots on target yields a variant earning weight 0.35 against the goals model -- it carries information the goals model lacks -- and 0.000 against the market. Two structural signals, each informative about the other, both priced. The structural model is better calibrated than the market on the home-win margin (slope 0.995 versus 1.103) while clearly less sharp: the market's advantage is discrimination rather than honesty, which accuracy alone cannot distinguish. Value lies not in a better forecast but in what is built on a calibrated one. We define match leverage, the change in a club's probability of achieving a season objective between winning and losing a fixture, and compute it for ACF Fiorentina: an away fixture against a relegation rival carried 2.25x the leverage of hosting the eventual champions. The paper also documents and corrects errors in an earlier study of our own.

stat.AP

A Least-Squares Approach to Sample-Based Prior Elicitation

An expert who supplies examples of a quantity often also signals how plausible each one is; when is that signal worth using? We study eliciting a Bayesian prior from an expert who provides example points together with their approximate likelihoods. We propose fitting the prior by least squares--minimizing the squared discrepancy between a parametric density and the elicited likelihoods--which defines an M-estimator that remains well posed even for families whose moments do not exist. We establish consistency and asymptotic normality, and prove--under explicit regularity conditions, comprising a well-separation and a uniform-concentration requirement that we verify for the families considered--a non-asymptotic O(1/sqrt(n)) Berry--Esseen bound on its sampling distribution, uniform and nonuniform, by extending a result of Pinelis for maximum-likelihood estimators to the M-estimation setting. We then relax the assumptions that most limit the method in practice. Experts need not report on the density's own scale: an unknown reporting scale can be profiled out in closed form and estimated jointly, at no asymptotic cost for location families. The theory extends to multivariate parameters, where a directional Berry--Esseen bound follows from the multivariate delta method applied to a smooth implicit proxy for the estimator. An additive error floor in the noise model removes a degeneracy in the optimal design, making optimal designs interior. Simulations for normal and beta families confirm the predicted n^-1 error rate and the accuracy of the normal approximation at moderate sample sizes. Finally, we compare the estimator with the sample-only maximum-likelihood baseline, derive an explicit threshold on the expert's reporting noise below which the elicited likelihoods provably reduce estimation error, and calibrate that threshold against eleven datasets of human frequency judgments.

stat.ME

Sliced Multi-Marginal Optimal Transport

Multi-marginal optimal transport enables one to compare multiple probability measures, which increasingly finds application in multi-task learning problems. One practical limitation of multi-marginal transport is computational scalability in the number of measures, samples and dimensionality. In this work, we propose a multi-marginal optimal transport paradigm based on random one-dimensional projections, whose (generalized) distance we term the sliced multi-marginal Wasserstein distance. To construct this distance, we introduce a characterization of the one-dimensional multi-marginal Kantorovich problem and use it to highlight a number of properties of the sliced multi-marginal Wasserstein distance. In particular, we show that (i) the sliced multi-marginal Wasserstein distance is a (generalized) metric that induces the same topology as the standard Wasserstein distance, (ii) it admits a dimension-free sample complexity, (iii) it is tightly connected with the problem of barycentric averaging under the sliced-Wasserstein metric. We conclude by illustrating the sliced multi-marginal Wasserstein on multi-task density estimation and multi-dynamics reinforcement learning problems.

stat.ML

Concentration Inequalities for U-Statistics: A Survey with Explicit Constants

This survey gives a self-contained treatment of concentration inequalities for U-statistics. Using Hoeffding's blocking argument, which reduces a U-statistic to an average over sums of independent random variables, we derive Hoeffding-, Bennett-, and Bernstein-type tail bounds with explicit constants, together with their extensions to unbounded sub-Gaussian and sub-exponential kernels and to two-sample and incomplete U-statistics. The reduction is stated once, as a convex-domination lemma, from which all the tail bounds follow as corollaries. While these results are classical -- the blocking argument goes back to Hoeffding (1963) and Bernstein-type bounds appear in Arcones (1995) -- complete elementary derivations with explicit constants are scattered or omitted in the literature, and collecting them is the purpose of this survey. We close with an overview of sharper bounds available under degeneracy assumptions and of robust median-of-means alternatives for heavy-tailed kernels, and with a numerical illustration that quantifies how conservative the explicit bounds are and decomposes the observed gap into interpretable factors.

math.ST