Searcharxiv⌕ Search

arXiv subjects

Jacek P. Dmochowski

Publications and source records attributed to Jacek P. Dmochowski.

4 recordsLinked to original sources

The profit-bias identity in sports betting: bookmaker profit as the public's prediction error

Sports betting moves money continuously from a large public to a small number of firms. The most influential account of that flow, due to Levitt (2004), holds that books price away from the market-clearing point to exploit predictable public biases. It computes profit from two numbers (the probability that a side wins the proposition and the fraction of handle it attracts), treating the share on a side as independent of the outcome. Here we relax that assumption and derive a profit-bias identity: profit is affine and increasing in the expected share of handle on the losing side, with Levitt's expression as the special case of independence. The identity resolves the book's margin into exactly three channels: its hold, the product of its price shading and the public's lean, and the covariance between bet share and outcome. A lean is thus worthless without shading, and shading is worthless without a lean. Under a public-belief model, the profit driver is the public's Bayes error: the probability of a representative bettor selecting the losing side. We provide necessary and sufficient conditions for a "Goldilocks Zone": prices at which book and bettor both profit. Testing these predictions on 1,139 Major League Baseball games, we find that the apparent dependence between bet share and outcome is a Simpson's paradox: present when games are pooled, but absent once they are separated by which side the book favored. The public leans heavily toward favorites, but we detect no matching shading, and the realized margin is indistinguishable from the hold.

stat.AP↗

Learning latent causal relationships in multiple time series

Identifying the causal structure of systems with multiple dynamic elements is critical to several scientific disciplines. The conventional approach is to conduct statistical tests of causality, for example with Granger Causality, between observed signals that are selected a priori. Here it is posited that, in many systems, the causal relations are embedded in a latent space that is expressed in the observed data as a linear mixture. A technique for blindly identifying the latent sources is presented: the observations are projected into pairs of components -- driving and driven -- to maximize the strength of causality between the pairs. This leads to an optimization problem with closed form expressions for the objective function and gradient that can be solved with off-the-shelf techniques. After demonstrating proof-of-concept on synthetic data with known latent structure, the technique is applied to recordings from the human brain and historical cryptocurrency prices. In both cases, the approach recovers multiple strong causal relationships that are not evident in the observed data. The proposed technique is unsupervised and can be readily applied to any multiple time series to shed light on the causal relationships underlying the data.

stat.ML↗

Correlated Components Analysis - Extracting Reliable Dimensions in Multivariate Data

How does one find dimensions in multivariate data that are reliably expressed across repetitions? For example, in a brain imaging study one may want to identify combinations of neural signals that are reliably expressed across multiple trials or subjects. For a behavioral assessment with multiple ratings, one may want to identify an aggregate score that is reliably reproduced across raters. Correlated Components Analysis (CorrCA) addresses this problem by identifying components that are maximally correlated between repetitions (e.g. trials, subjects, raters). Here we formalize this as the maximization of the ratio of between-repetition to within-repetition covariance. We show that this criterion maximizes repeat-reliability, defined as mean over variance across repeats, and that it leads to CorrCA or to multi-set Canonical Correlation Analysis, depending on the constraints. Surprisingly, we also find that CorrCA is equivalent to Linear Discriminant Analysis for zero-mean signals, which provides an unexpected link between classic concepts of multivariate analysis. We present an exact parametric test of statistical significance based on the F-statistic for normally distributed independent samples, and present and validate shuffle statistics for the case of dependent samples. Regularization and extension to non-linear mappings using kernels are also presented. The algorithms are demonstrated on a series of data analysis applications, and we provide all code and data required to reproduce the results.

stat.ML↗

Maximally reliable spatial filtering of steady state visual evoked potentials

Due to their high signal-to-noise ratio (SNR) and robustness to artifacts, steady state visual evoked potentials (SSVEPs) are a popular technique for studying neural processing in the human visual system. SSVEPs are conventionally analyzed at individual electrodes or linear combinations of electrodes which maximize some variant of the SNR. Here we exploit the fundamental assumption of evoked responses -- reproducibility across trials -- to develop a technique that extracts a small number of high SNR, maximally reliable SSVEP components. This novel spatial filtering method operates on an array of Fourier coefficients and projects the data into a low-dimensional space in which the trial-to-trial spectral covariance is maximized. When applied to two sample data sets, the resulting technique recovers physiologically plausible components (i.e., the recovered topographies match the lead fields of the underlying sources) while drastically reducing the dimensionality of the data (i.e., more than 90% of the trial-to-trial reliability is captured in the first four components). Moreover, the proposed technique achieves a higher SNR than that of the single-best electrode or the Principal Components. We provide a freely-available MATLAB implementation of the proposed technique, herein termed "Reliable Components Analysis".

q-bio.QM↗