SearcharxivSearch

arXiv subjects

Spencer Wheatley

Publications and source records attributed to Spencer Wheatley.

11 recordsLinked to original sources

Revisiting the predictability of the Haicheng and Tangshan earthquakes

We analyse the compiled set of precursory data that were reported to be available in real time before the Ms 7.5 Haicheng earthquake in Feb. 1975 and the Ms 7.6-7.8 Tangshan earthquake in July 1976. We propose a robust and simple coarse-graining method consisting in aggregating and counting how all the anomalies together (geodesy, levelling, geomagnetism, soil resistivity, Earth currents, gravity, Earth stress, well water radon, well water level) develop as a function of time. We demonstrate a strong evidence for the existence of an acceleration of the number of anomalies leading up to the major Haicheng and Tangshan earthquakes. In particular for the Tangshan earthquake, the frequency of occurrence of anomalies is found to be well described by the log-periodic power law singularity (LPPLS) model, previously proposed for the prediction of engineering failures and later adapted to the prediction of financial crashes. Based on a mock real-time prediction experiment, and simulation study, we show the potential for an early warning system with lead-time of a few days, based on this methodology of monitoring accelerated rates of anomalies.

physics.geo-ph

Statistical witchhunts: Science, justice & the p-value crisis

We provide accessible insight into the current 'replication crisis' in 'statistical science', by revisiting the old metaphor of 'court trial as hypothesis test'. Inter alia, we define and diagnose harmful statistical witch-hunting both in justice and science, which extends to the replication crisis itself, where a hunt on p-values is currently underway.

stat.OT

Data breaches in the catastrophe framework & beyond

Development of sustainable insurance for cyber risks, with associated benefits, inter alia requires reduction of ambiguity of the risk. Considering cyber risk, and data breaches in particular, as a man-made catastrophe clarifies the actuarial need for multiple levels of analysis - going beyond claims-driven loss statistics alone to include exposure, hazard, breach size, and so on - and necessitating specific advances in scope, quality, and standards of both data and models. The prominent human element, as well as dynamic, networked, and multi-type nature, of cyber risk makes it perhaps uniquely challenging. Complementary top-down statistical, and bottom-up analytical approaches are discussed. Focusing on data breach severity, measured in private information items ('ids') extracted, we exploit relatively mature open data for U.S. data breaches. We show that this extremely heavy-tailed risk is worsening for external attacker ('hack') events - both in frequency and severity. Writing in Q2-2018, the median predicted number of ids breached in the U.S. due to hacking, for the last 6 months of 2018, is 0.5 billion. But with a 5% chance that the figure exceeds 7 billion - doubling the historical total. 'Fortunately' the total breach in that period turned out to be near the median.

stat.AP

The fair reward problem: the illusion of success and how to solve it

Humanity has been fascinated by the pursuit of fortune since time immemorial, and many successful outcomes benefit from strokes of luck. But success is subject to complexity, uncertainty, and change - and at times becoming increasingly unequally distributed. This leads to tension and confusion over to what extent people actually get what they deserve (i.e., fairness/meritocracy). Moreover, in many fields, humans are over-confident and pervasively confuse luck for skill (I win, it's skill; I lose, it's bad luck). In some fields, there is too much risk taking; in others, not enough. Where success derives in large part from luck - and especially where bailouts skew the incentives (heads, I win; tails, you lose) - it follows that luck is rewarded too much. This incentivizes a culture of gambling, while downplaying the importance of productive effort. And, short term success is often rewarded, irrespective, and potentially at the detriment, of the long-term system fitness. However, much success is truly meritocratic, and the problem is to discern and reward based on merit. We call this the fair reward problem. To address this, we propose three different measures to assess merit: (i) raw outcome; (ii) risk adjusted outcome, and (iii) prospective. We emphasize the need, in many cases, for the deductive prospective approach, which considers the potential of a system to adapt and mutate in novel futures. This is formalized within an evolutionary system, comprised of five processes, inter alia handling the exploration-exploitation trade-off. Several human endeavors - including finance, politics, and science -are analyzed through these lenses, and concrete solutions are proposed to support a prosperous and meritocratic society.

econ.GN

The ARMA Point Process and its Estimation

We introduce the ARMA (autoregressive-moving-average) point process, which is a Hawkes process driven by a Neyman-Scott process with Poisson immigration. It contains both the Hawkes and Neyman-Scott process as special cases and naturally combines self-exciting and shot-noise cluster mechanisms, useful in a variety of applications. The name ARMA is used because the ARMA point process is an appropriate analogue of the ARMA time series model for integer-valued series. As such, the ARMA point process framework accommodates a flexible family of models sharing methodological and mathematical similarities with ARMA time series. We derive an estimation procedure for ARMA point processes, as well as the integer ARMA models, based on an MCEM (Monte Carlo Expectation Maximization) algorithm. This powerful framework for estimation accommodates trends in immigration, multiple parametric specifications of excitement functions, as well as cases where marks and immigrants are not observed.

math.ST

Classification of cryptocurrency coins and tokens by the dynamics of their market capitalisations

We empirically verify that the market capitalisations of coins and tokens in the cryptocurrency universe follow power-law distributions with significantly different values, with the tail exponent falling between 0.5 and 0.7 for coins, and between 1.0 and 1.3 for tokens. We provide a rationale for this, based on a simple proportional growth with birth & death model previously employed to describe the size distribution of firms, cities, webpages, etc. We empirically validate the model and its main predictions, in terms of proportional growth (Gibrat's law) of the coins and tokens. Estimating the main parameters of the model, the theoretical predictions for the power-law exponents of coin and token distributions are in remarkable agreement with the empirical estimations, given the simplicity of the model. Our results clearly characterize coins as being "entrenched incumbents" and tokens as an "explosive immature ecosystem", largely due to massive and exuberant Initial Coin Offering activity in the token space. The theory predicts that the exponent for tokens should converge to 1 in the future, reflecting a more reasonable rate of new entrants associated with genuine technological innovations.

physics.soc-ph

Are Bitcoin Bubbles Predictable? Combining a Generalized Metcalfe's Law and the LPPLS Model

We develop a strong diagnostic for bubbles and crashes in bitcoin, by analyzing the coincidence (and its absence) of fundamental and technical indicators. Using a generalized Metcalfe's law based on network properties, a fundamental value is quantified and shown to be heavily exceeded, on at least four occasions, by bubbles that grow and burst. In these bubbles, we detect a universal super-exponential unsustainable growth. We model this universal pattern with the Log-Periodic Power Law Singularity (LPPLS) model, which parsimoniously captures diverse positive feedback phenomena, such as herding and imitation. The LPPLS model is shown to provide an ex-ante warning of market instabilities, quantifying a high crash hazard and probabilistic bracket of the crash time consistent with the actual corrections; although, as always, the precise time and trigger (which straw breaks the camel's back) being exogenous and unpredictable. Looking forward, our analysis identifies a substantial but not unprecedented overvaluation in the price of bitcoin, suggesting many months of volatile sideways bitcoin prices ahead (from the time of writing, March 2018).

econ.EM

The Extreme Risk of Personal Data Breaches & The Erosion of Privacy

Personal data breaches from organisations, enabling mass identity fraud, constitute an \emph{extreme risk}. This risk worsens daily as an ever-growing amount of personal data are stored by organisations and on-line, and the attack surface surrounding this data becomes larger and harder to secure. Further, breached information is distributed and accumulates in the hands of cyber criminals, thus driving a cumulative erosion of privacy. Statistical modeling of breach data from 2000 through 2015 provides insights into this risk: A current maximum breach size of about 200 million is detected, and is expected to grow by fifty percent over the next five years. The breach sizes are found to be well modeled by an \emph{extremely heavy tailed} truncated Pareto distribution, with tail exponent parameter decreasing linearly from 0.57 in 2007 to 0.37 in 2015. With this current model, given a breach contains above fifty thousand items, there is a ten percent probability of exceeding ten million. A size effect is unearthed where both the frequency and severity of breaches scale with organisation size like $s^{0.6}$. Projections indicate that the total amount of breached information is expected to double from two to four billion items within the next five years, eclipsing the population of users of the Internet. This massive and uncontrolled dissemination of personal identities raises fundamental concerns about privacy.

stat.AP

Of Disasters and Dragon Kings: A Statistical Analysis of Nuclear Power Incidents & Accidents

We provide, and perform a risk theoretic statistical analysis of, a dataset that is 75 percent larger than the previous best dataset on nuclear incidents and accidents, comparing three measures of severity: INES (International Nuclear Event Scale), radiation released, and damage dollar losses. The annual rate of nuclear accidents, with size above 20 Million US$, per plant, decreased from the 1950s until dropping significantly after Chernobyl (April, 1986). The rate is now roughly stable at 0.002 to 0.003, i.e., around 1 event per year across the current fleet. The distribution of damage values changed after Three Mile Island (TMI; March, 1979), where moderate damages were suppressed but the tail became very heavy, being described by a Pareto distribution with tail index 0.55. Further, there is a runaway disaster regime, associated with the "dragon-king" phenomenon, amplifying the risk of extreme damage. In fact, the damage of the largest event (Fukushima; March, 2011) is equal to 60 percent of the total damage of all 174 accidents in our database since 1946. In dollar losses we compute a 50% chance that (i) a Fukushima event (or larger) occurs in the next 50 years, (ii) a Chernobyl event (or larger) occurs in the next 27 years and (iii) a TMI event (or larger) occurs in the next 10 years. Finally, we find that the INES scale is inconsistent. To be consistent with damage, the Fukushima disaster would need to have an INES level of 11, rather than the maximum of 7.

physics.soc-ph

Estimation of the Hawkes Process With Renewal Immigration Using the EM Algorithm

We introduce the Hawkes process with renewal immigration and make its statistical estimation possible with two Expectation Maximization (EM) algorithms. The standard Hawkes process introduces immigrant points via a Poisson process, and each immigrant has a subsequent cluster of associated offspring of multiple generations. We generalize the immigration to come from a Renewal process; introducing dependence between neighbouring clusters, and allowing for over/under dispersion in cluster locations. This complicates evaluation of the likelihood since one needs to know which subset of the observed points are immigrants. Two EM algorithms enable estimation here: The first is an extension of an existing algorithm that treats the entire branching structure - which points are immigrants, and which point is the parent of each offspring - as missing data. The second considers only if a point is an immigrant or not as missing data and can be implemented with linear time complexity. Both algorithms are found to be consistent in simulation studies. Further, we show that misspecifying the immigration process introduces signficant bias into model estimation-- especially the branching ratio, which quantifies the strength of self excitation. Thus, this extended model provides a valuable alternative model in practice.

stat.AP

Effective Measure of Endogeneity for the Autoregressive Conditional Duration Point Processes via Mapping to the Self-Excited Hawkes Process

In order to disentangle the internal dynamics from exogenous factors within the Autoregressive Conditional Duration (ACD) model, we present an effective measure of endogeneity. Inspired from the Hawkes model, this measure is defined as the average fraction of events that are triggered due to internal feedback mechanisms within the total population. We provide a direct comparison of the Hawkes and ACD models based on numerical simulations and show that our effective measure of endogeneity for the ACD can be mapped onto the "branching ratio" of the Hawkes model.

q-fin.ST