SearcharxivSearch

arXiv subjects

Vladimir Vovk

Publications and source records attributed to Vladimir Vovk.

At least 19 recordsLinked to original sources

The universal measure of probabilistically nonrandom objects

A finite object is called probabilistically random if it is an algorithmically random element of a finite set whose Kolmogorov complexity is relatively small. This note studies the amount of probabilistically nonrandom objects (those that are not probabilistically random) as gauged by the universal measure and without assuming any structure on the objects. The main results imply that the universal measure of probabilistically nonrandom objects is vanishingly small and establish the dependence of their universal measure on the required degree of probabilistic randomness.

math.LO

Inductive Venn-Abers and related regressors

Venn-Abers predictors are probabilistic predictors that enjoy appealing properties of validity, but their major limitation is that they have been applicable only to binary classification, apart from a recent extension to bounded regression. We generalize them to the case of unbounded regression, which requires adding an element of conformal prediction. In our simulation and empirical studies we investigate the predictive efficiency of point regressors derived from Venn-Abers regressors and argue that they somewhat improve the predictive efficiency of standard regressors for larger training sets.

cs.LG

The universal measure of nonstochastic objects

The usual definitions of stochasticity are very different from Kolmogorov's original definition as given in Shen's notes from November 1981. This note simplifies Kolmogorov's definition and shows that the universal machine generates nonstochastic objects with surprisingly high probability.

math.LO

The non-mathematical section of Kolmogorov's Grundbegriffe

Kolmogorov's "Grundbegriffe der Wahrscheinlichkeitsrechnung" introduced the standard measure-theoretic formalization of probability but included only a two-page section about the relation of the mathematical theory of probability to the world of experience. This note discusses this brief non-mathematical section concentrating on what I regard as its two most controversial features: having two bridges between the mathematical theory and the world of experience instead of the standard one, and the possibility of observing a prespecified event of probability zero in a sequence of trials.

math.PR

Exchangeability and randomness for infinite and finite sequences

Randomness (in the sense of being generated in an IID fashion) and exchangeability are standard assumptions in nonparametric statistics and machine learning, and relations between them have been a popular topic of research. This short paper draws the reader's attention to the fact that, while for infinite sequences of observations the two assumptions are almost indistinguishable, the difference between them becomes very significant for finite sequences of a given length.

math.ST

Aggregation in conformal e-classification

Aggregating conformal predictors is a standard way of balancing their predictive and computational efficiency while retaining their validity, at least approximately. An important advantage of conformal e-predictors is that they are easier to aggregate without sacrificing their validity. This paper studies experimentally cross-conformal e-prediction, which is an existing method of aggregating conformal e-predictors, and its modifications that are conceptually simpler and more flexible.

cs.LG

Confidence intervals for causal effects in sequential decision making

We derive confidence intervals and confidence sequences for causal effects in situations where the back-door criterion is applicable. Our tightest confidence intervals hold in the standard setting where the training data consists of IID observations over a system described by a given causal diagram. When interventions are allowed to depend on the past data, our confidence intervals become wider and involve a term coming from the law of the iterated logarithm, even where the number of observations is known in advance. In the sequential setting where the number of observations is not given, our confidence intervals, arranged into a confidence sequence for causal effects, involve more iterated logarithm terms and become even wider.

math.ST

A law of large numbers for predicting several steps ahead

This note proves a law of large numbers for predicting several steps ahead, which, in the case of uniformly bounded random variables, generalizes the standard law of large numbers for martingales; the standard law of large numbers corresponds to predicting one step ahead. Its main result shows that the law of large numbers holds for predicting $N$ uniformly bounded random variables $o(N)$ steps ahead, but it is much more precise and in some respects optimal. This law of large numbers is applied to a problem of decision making with a bounded loss function limiting the impact of each decision to $o(N)$ steps.

math.PR

Conformal e-prediction in the presence of confounding

This note extends conformal e-prediction to cover the case where there is observed confounding between the random object $X$ and its label $Y$. We consider both the case where the observed data is IID and a case where some dependence between observations is permitted.

math.ST

Inductive randomness predictors: beyond conformal

This paper introduces inductive randomness predictors, which form a proper superset of inductive conformal predictors but have the same principal property of validity under the assumption of randomness (i.e., of IID data). It turns out that every non-trivial inductive conformal predictor is strictly dominated by an inductive randomness predictor, although the improvement is not great, at most a factor of $\mathrm{e}\approx2.72$ in the case of e-prediction. The dominating inductive randomness predictors are more complicated and more difficult to compute; besides, an improvement by a factor of $\mathrm{e}$ is rare. Therefore, this paper does not suggest replacing inductive conformal predictors by inductive randomness predictors and only calls for a more detailed study of the latter.

cs.LG

Randomness, exchangeability, and conformal prediction

This paper argues for a wider use of the functional theory of randomness, a modification of the algorithmic theory of randomness getting rid of unspecified additive constants. Both theories are useful for understanding relationships between the assumptions of IID data and data exchangeability. While the assumption of IID data is standard in machine learning, conformal prediction relies on data exchangeability. Nouretdinov, V'yugin, and Gammerman showed, using the language of the algorithmic theory of randomness, that conformal prediction is a universal method under the assumption of IID data. In this paper (written for the Alex Gammerman Festschrift) I will selectively review connections between exchangeability and the property of being IID, early history of conformal prediction, my encounters and collaboration with Alex and other interesting people, and a translation of Nouretdinov et al.'s results into the language of the functional theory of randomness, which moves it closer to practice. Namely, the translation says that every confidence predictor that is valid for IID data can be transformed to a conformal predictor without losing much in predictive efficiency.

cs.LG

Universality of conformal prediction under the assumption of randomness

Conformal predictors provide set or functional predictions that are valid under the assumption of randomness, i.e., under the assumption of independent and identically distributed data. The question asked in this paper is whether there are predictors that are valid in the same sense under the assumption of randomness and that are more efficient than conformal predictors. The answer is that the class of conformal predictors is universal in that only limited gains in predictive efficiency are possible. The previous work in this area has relied on the algorithmic theory of randomness and so involved unspecified constants, whereas this paper's results are much more practical. They are also shown to be optimal in some respects.

cs.LG

Conformal e-prediction

This paper discusses a counterpart of conformal prediction for e-values, conformal e-prediction. Conformal e-prediction is conceptually simpler and had been developed in the 1990s as a precursor of conformal prediction. When conformal prediction emerged as result of replacing e-values by p-values, it seemed to have important advantages over conformal e-prediction without obvious disadvantages. This paper re-examines relations between conformal prediction and conformal e-prediction systematically from a modern perspective. Conformal e-prediction has advantages of its own, such as the ease of designing conditional conformal e-predictors and the guaranteed validity of cross-conformal e-predictors (whereas for cross-conformal predictors validity is only an empirical fact and can be broken with excessive randomization). Even where conformal prediction has clear advantages, conformal e-prediction can often emulate those advantages, more or less successfully.

cs.LG

Conditionality principle under unconstrained randomness

A very simple example demonstrates that Fisher's application of the conditionality principle to regression ("fixed-$x$ regression"), endorsed by Sprott and many other followers, makes prediction impossible in the context of statistical learning theory. On the other hand, relaxing the requirement of conditionality makes it possible via, e.g., conformal prediction.

math.ST

Conformal e-testing

There is a useful counterpart of conformal prediction for e-values, called conformal e-prediction. Conformal prediction can serve as basis for testing the assumption of exchangeability, leading to conformal testing. Similarly, conformal e-prediction can also serve as basis for testing. The resulting conformal e-testing looks very different from but inherits some strengths of conformal testing; it even has some advantages over conformal testing. In this paper we discuss systematically both strengths and limitations of conformal e-testing.

math.ST

True and false discoveries with independent and sequential e-values

In this paper we use e-values in the context of multiple hypothesis testing assuming that the base tests produce independent, or sequential, e-values. Our simulation and empirical studies and theoretical considerations suggest that, under this assumption, our new algorithms are superior to the known algorithms using independent p-values and to our recent algorithms designed for e-values without the assumption of independence.

stat.ME

Asymptotic uniqueness in long-term prediction

This paper establishes the asymptotic uniqueness of long-term probability forecasts in the following form. Consider two forecasters who repeatedly issue probability forecasts for the infinite future. The main result of the paper says that either at least one of the two forecasters will be discredited or their forecasts will converge in total variation. This can be regarded as a game-theoretic version of the classical Blackwell-Dubins result getting rid of some of its limitations. This result is further strengthened along the lines of Richard Jeffrey's radical probabilism.

math.ST