SearcharxivSearch

arXiv subjects

Nassim Nicholas Taleb

Publications and source records attributed to Nassim Nicholas Taleb.

At least 19 recordsLinked to original sources

Statistical Consequences of Fat Tails: Real World Preasymptotics, Epistemology, and Applications

(The third edition corrects minor typos and adds 3 chapters synthesized from published papers plus an appendix on maximum entropy distributions.) The monograph investigates the misapplication of conventional statistical techniques to fat tailed distributions and looks for remedies, when possible. Switching from thin tailed to fat tailed distributions requires more than "changing the color of the dress". Traditional asymptotics deal mainly with either n=1 or $n=\infty$, and the real world is in between, under of the "laws of the medium numbers" --which vary widely across specific distributions. Both the law of large numbers and the generalized central limit mechanisms operate in highly idiosyncratic ways outside the standard Gaussian or Levy-Stable basins of convergence. A few examples: + The sample mean is rarely in line with the population mean, with effect on "naive empiricism", but can be sometimes be estimated via parametric methods. + The "empirical distribution" is rarely empirical. + Parameter uncertainty has compounding effects on statistical metrics. + Dimension reduction (principal components) fails. + Inequality estimators (GINI or quantile contributions) are not additive and produce wrong results. + Many "biases" found in psychology become entirely rational under more sophisticated probability distributions + Most of the failures of financial economics, econometrics, and behavioral economics can be attributed to using the wrong distributions. This book, the first volume of the Technical Incerto, weaves a narrative around published journal articles.

stat.OT

Informational Rescaling of PCA Maps with Application to Genetic Distance

We discuss the inadequacy of covariances/correlations and other measures in L2 as relative distance metrics under some conditions. We propose a computationally simple heuristic to transform a map based on standard principal component analysis (PCA) (when the variables are asymptotically Gaussian) into an entropy-based map where distances are based on mutual information (MI). Rescaling Principal Component based distances using MI allows a representation of relative statistical associations when, as in genetics, it is applied on bit measurements between individuals' genomic mutual information. This entropy rescaled PCA, while preserving order relationships (along a dimension), changes the relative distances to make them linear to information. We show the effect on the entire world population and some subsamples, which leads to significant differences with the results of current research.

cs.IT

Tail Option Pricing Under Power Laws

We build a methodology that takes a given option price in the tails with strike $K$ and extends (for calls, all strikes > $K$, for puts all strikes $< K$) assuming the continuation falls into what we define as "Karamata Constant" over which the strong Pareto law holds. The heuristic produces relative prices for options, with for sole parameter the tail index $α$, under some mild arbitrage constraints. Usual restrictions such as finiteness of variance are not required. The methodology allows us to scrutinize the volatility surface and test various theories of relative tail option overpricing (usually built on thin tailed models and minor modifications/fudging of the Black-Scholes formula).

q-fin.PR

The Probability Conflation: A Reply

We respond to Tetlock et al. (2022) showing 1) how expert judgment fails to reflect tail risk, 2) the lack of compatibility between forecasting tournaments and tail risk assessment methods (such as extreme value theory). More importantly, we communicate a new result showing a greater gap between the properties of tail expectation and those of the corresponding probability.

q-fin.RM

Working With Convex Responses: Antifragility From Finance to Oncology

We extend techniques and learnings about the stochastic properties of nonlinear responses from finance to medicine, particularly oncology where it can inform dosing and intervention. We define antifragility. We propose uses of risk analysis to medical problems, through the properties of nonlinear responses (convex or concave). We 1) link the convexity/concavity of the dose-response function to the statistical properties of the results; 2) define "antifragility" as a mathematical property for local beneficial convex responses and the generalization of "fragility" as its opposite, locally concave in the tails of the statistical distribution; 3) propose mathematically tractable relations between dosage, severity of conditions, and iatrogenics. In short we propose a framework to integrate the necessary consequences of nonlinearities in evidence-based oncology and more general clinical risk management.

q-bio.QM

Bitcoin, Currencies, and Fragility

This discussion applies quantitative finance methods and economic arguments to cryptocurrencies in general and bitcoin in particular -- as there are about $10,000$ cryptocurrencies, we focus (unless otherwise specified) on the most discussed crypto of those that claim to hew to the original protocol (Nakamoto 2009) and the one with, by far, the largest market capitalization. In its current version, in spite of the hype, bitcoin failed to satisfy the notion of "currency without government" (it proved to not even be a currency at all), can be neither a short nor long term store of value (its expected value is no higher than $0$), cannot operate as a reliable inflation hedge, and, worst of all, does not constitute, not even remotely, a safe haven for one's investments, a shield against government tyranny, or a tail protection vehicle for catastrophic episodes. Furthermore, bitcoin promoters appear to conflate the success of a payment mechanism (as a decentralized mode of exchange), which so far has failed, with the speculative variations in the price of a zero-sum maximally fragile asset with massive negative externalities. Going through monetary history, we show how a true numeraire must be one of minimum variance with respect to an arbitrary basket of goods and services, how gold and silver lost their inflation hedge status during the Hunt brothers squeeze in the late 1970s and what would be required from a true inflation hedged store of value.

econ.GN

Unmasking the mask studies: why the effectiveness of surgical masks in preventing respiratory infections has been underestimated

Face masks have been widely used as a protective measure against COVID-19. However, pre-pandemic empirical studies have produced mixed statistical results on the effectiveness of masks against respiratory viruses. The implications of the studies' recognized limitations have not been quantitatively and statistically analyzed, leading to confusion regarding the effectiveness of masks. Such confusion may have contributed to organizations such as the WHO and CDC initially not recommending that the general public wear masks. Here we show that when the adherence to mask-usage guidelines is taken into account, the empirical evidence indicates that masks prevent disease transmission: all studies we analyzed that did not find surgical masks to be effective were under-powered to such an extent that even if masks were 100% effective, the studies in question would still have been unlikely to find a statistically significant effect. We also provide a framework for understanding the effect of masks on the probability of infection for single and repeated exposures. The framework demonstrates that more frequently wearing a mask provides super-linearly compounding protection, as does both the susceptible and infected individual wearing a mask. This work shows (1) that both theoretical and empirical evidence is consistent with masks protecting against respiratory infections and (2) that nonlinear effects and statistical considerations regarding the percentage of exposures for which masks are worn must be taken into account when designing empirical studies and interpreting their results.

q-bio.QM

On Single Point Forecasts for Fat-Tailed Variables

We discuss common errors and fallacies when using naive "evidence based" empiricism and point forecasts for fat-tailed variables, as well as the insufficiency of using naive first-order scientific methods for tail risk management. We use the COVID-19 pandemic as the background for the discussion and as an example of a phenomenon characterized by a multiplicative nature, and what mitigating policies must result from the statistical properties and associated risks. In doing so, we also respond to the points raised by Ioannidis et al. (2020).

physics.soc-ph

Tail Risk of Contagious Diseases

Applying a modification of Extreme value Theory (thanks to a dual distribution technique by the authors on data over the past 2,500 years, we show that pandemics are extremely fat-tailed in terms of fatalities, with a marked potentially existential risk for humanity. Such a macro property should invite the use of Extreme Value Theory (EVT) rather than naive interpolations and expected averages for risk management purposes. An implication is that potential tail risk overrides conclusions on decisions derived from compartmental epidemiological models and similar approaches.

physics.soc-ph

What You See and What You Don't See: The Hidden Moments of a Probability Distribution

Empirical distributions have their in-sample maxima as natural censoring. We look at the "hidden tail", that is, the part of the distribution in excess of the maximum for a sample size of $n$. Using extreme value theory, we examine the properties of the hidden tail and calculate its moments of order $p$. The method is useful in showing how large a bias one can expect, for a given $n$, between the visible in-sample mean and the true statistical mean (or higher moments), which is considerable for $α$ close to 1. Among other properties, we note that the "hidden" moment of order $0$, that is, the exceedance probability for power law distributions, follows an exponential distribution and has for expectation $\frac{1}{n}$ regardless of the parametrization of the scale and tail index.

q-fin.ST

On the Statistical Differences between Binary Forecasts and Real World Payoffs

What do binary (or probabilistic) forecasting abilities have to do with overall performance? We map the difference between (univariate) binary predictions, bets and "beliefs" (expressed as a specific "event" will happen/will not happen) and real-world continuous payoffs (numerical benefits or harm from an event) and show the effect of their conflation and mischaracterization in the decision-science literature. We also examine the differences under thin and fat tails. The effects are: A- Spuriousness of many psychological results particularly those documenting that humans overestimate tail probabilities and rare events, or that they overreact to fears of market crashes, ecological calamities, etc. Many perceived "biases" are just mischaracterizations by psychologists. There is also a misuse of Hayekian arguments in promoting prediction markets. We quantify such conflations with a metric for "pseudo-overestimation". B- Being a "good forecaster" in binary space doesn't lead to having a good actual performance}, and vice versa, especially under nonlinearities. A binary forecasting record is likely to be a reverse indicator under some classes of distributions. Deeper uncertainty or more complicated and realistic probability distribution worsen the conflation . C- Machine Learning: Some nonlinear payoff functions, while not lending themselves to verbalistic expressions and "forecasts", are well captured by ML or expressed in option contracts. D- Fattailedness: The difference is exacerbated in the power law classes of probability distributions.

q-fin.GN

Branching epistemic uncertainty and thickness of tails

This is an epistemological approach to errors in both inference and risk management, leading to necessary structural properties for the probability distribution. Many mechanisms have been used to show the emergence of fat tails. Here we follow an alternative route, the epistemological one, using counterfactual analysis, and show how nested uncertainty, that is, errors on the error in estimation of parameters, lead to fattailedness of the distribution. The results have relevant implications for forecasting, dealing with model risk and generally all statistical analyses. The more interesting results are as follows: + The forecasting paradox: The future is fatter tailed than the past. Further, out of sample results should be fatter tailed than in-sample ones. + Errors on errors can be explosive or implosive with different consequences. Infinite recursions can be easily dealt with, pending on the structure of the errors. We also present a method to perform counterfactual analysis without the explosion of branching counterfactuals.

stat.ME

Election Predictions as Martingales: An Arbitrage Approach

We consider the estimation of binary election outcomes as martingales and propose an arbitrage pricing when one continuously updates estimates. We argue that the estimator needs to be priced as a binary option as the arbitrage valuation minimizes the conventionally used Brier score for tracking the accuracy of probability assessors. We create a dual martingale process $Y$, in $[L,H]$ from the standard arithmetic Brownian motion, $X$ in $(-\infty, \infty)$ and price elections accordingly. The dual process $Y$ can represent the numerical votes needed for success. We show the relationship between the volatility of the estimator in relation to that of the underlying variable. When there is a high uncertainty about the final outcome, 1) the arbitrage value of the binary gets closer to 50\%, 2) the estimate should not undergo large changes even if polls or other bases show significant variations. There are arbitrage relationships between 1) the binary value, 2) the estimation of $Y$, 3) the volatility of the estimation of $Y$ over the remaining time to expiration. We note that these arbitrage relationships were often violated by the various forecasting groups in the U.S. presidential elections of 2016, as well as the notion that all intermediate assessments of the success of a candidate need to be considered, not just the final one.

q-fin.PR

How Much Data Do You Need? An Operational, Pre-Asymptotic Metric for Fat-tailedness

This note presents an operational measure of fat-tailedness for univariate probability distributions, in $[0,1]$ where 0 is maximally thin-tailed (Gaussian) and 1 is maximally fat-tailed. Among others,1) it helps assess the sample size needed to establish a comparative $n$ needed for statistical significance, 2) allows practical comparisons across classes of fat-tailed distributions, 3) helps understand some inconsistent attributes of the lognormal, pending on the parametrization of its scale parameter. The literature is rich for what concerns asymptotic behavior, but there is a large void for finite values of $n$, those needed for operational purposes. Conventional measures of fat-tailedness, namely 1) the tail index for the power law class, and 2) Kurtosis for finite moment distributions fail to apply to some distributions, and do not allow comparisons across classes and parametrization, that is between power laws outside the Levy-Stable basin, or power laws to distributions in other classes, or power laws for different number of summands. How can one compare a sum of 100 Student T distributed random variables with 3 degrees of freedom to one in a Levy-Stable or a Lognormal class? How can one compare a sum of 100 Student T with 3 degrees of freedom to a single Student T with 2 degrees of freedom? We propose an operational and heuristic measure that allow us to compare $n$-summed independent variables under all distributions with finite first moment. The method is based on the rate of convergence of the Law of Large numbers for finite sums, $n$-summands specifically. We get either explicit expressions or simulation results and bounds for the lognormal, exponential, Pareto, and the Student T distributions in their various calibrations --in addition to the general Pearson classes.

stat.ME

(Anti)Fragility and Convex Responses in Medicine

This paper applies risk analysis to medical problems, through the properties of nonlinear responses (convex or concave). It shows 1) necessary relations between the nonlinearity of dose-response and the statistical properties of the outcomes, particularly the effect of the variance (i.e., the expected frequency of the various results and other properties such as their average and variations); 2) The description of "antifragility" as a mathematical property for local convex response and its generalization and the designation "fragility" as its opposite, locally concave; 3) necessary relations between dosage, severity of conditions, and iatrogenics. Iatrogenics seen as the tail risk from a given intervention can be analyzed in a probabilistic decision-theoretic way, linking probability to nonlinearity of response. There is a necessary two-way mathematical relation between nonlinear response and the tail risk of a given intervention. In short we propose a framework to integrate the necessary consequences of nonlinearities in evidence-based medicine and medical risk management. Keywords: evidence based medicine, risk management, nonlinear responses

q-bio.QM

A Short Note on P-Value Hacking

We present the expected values from p-value hacking as a choice of the minimum p-value among $m$ independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the meta-distribution. We first present an exact probability distribution (meta-distribution) for p-values across ensembles of statistically identical phenomena. We derive the distribution for small samples $2<n \leq n^*\approx 30$ as well as the limiting one as the sample size $n$ becomes large. We also look at the properties of the "power" of a test through the distribution of its inverse for a given p-value and parametrization. The formulas allow the investigation of the stability of the reproduction of results and "p-hacking" and other aspects of meta-analysis. P-values are shown to be extremely skewed and volatile, regardless of the sample size $n$, and vary greatly across repetitions of exactly same protocols under identical stochastic copies of the phenomenon; such volatility makes the minimum $p$ value diverge significantly from the "true" one. Setting the power is shown to offer little remedy unless sample size is increased markedly or the p-value is lowered by at least one order of magnitude.

stat.AP

Gini estimation under infinite variance

We study the problems related to the estimation of the Gini index in presence of a fat-tailed data generating process, i.e. one in the stable distribution class with finite mean but infinite variance (i.e. with tail index $α\in(1,2)$). We show that, in such a case, the Gini coefficient cannot be reliably estimated using conventional nonparametric methods, because of a downward bias that emerges under fat tails. This has important implications for the ongoing discussion about economic inequality. We start by discussing how the nonparametric estimator of the Gini index undergoes a phase transition in the symmetry structure of its asymptotic distribution, as the data distribution shifts from the domain of attraction of a light-tailed distribution to that of a fat-tailed one, especially in the case of infinite variance. We also show how the nonparametric Gini bias increases with lower values of $α$. We then prove that maximum likelihood estimation outperforms nonparametric methods, requiring a much smaller sample size to reach efficiency. Finally, for fat-tailed data, we provide a simple correction mechanism to the small sample bias of the nonparametric estimator based on the distance between the mode and the mean of its asymptotic distribution.

stat.ME

Stochastic Tail Exponent For Asymmetric Power Laws

We examine random variables in the power law/regularly varying class with stochastic tail exponent, the exponent $α$ having its own distribution. We show the effect of stochasticity of $α$ on the expectation and higher moments of the random variable. For instance, the moments of a right-tailed or right-asymmetric variable, when finite, increase with the variance of $α$; those of a left-asymmetric one decreases. The same applies to conditional shortfall (CVar), or mean-excess functions. We prove the general case and examine the specific situation of lognormally distributed $α\in [b,\infty), b>1$. The stochasticity of the exponent induces a significant bias in the estimation of the mean and higher moments in the presence of data uncertainty. This has consequences on sampling error as uncertainty about $α$ translates into a higher expected mean. The bias is conserved under summation, even upon large enough a number of summands to warrant convergence to the stable distribution. We establish inequalities related to the asymmetry. We also consider the situation of capped power laws (i.e. with compact support), and apply it to the study of violence by Cirillo and Taleb (2016). We show that uncertainty concerning the historical data increases the true mean.

q-fin.ST