SearcharxivSearch

arXiv subjects

Gordon J Ross

Publications and source records attributed to Gordon J Ross.

5 recordsLinked to original sources

Did Mary Shelley Write Frankenstein? A Stylometric Analysis

The novel Frankenstein was published anonymously in 1818, and was first credited to Mary Shelley in a French translation of 1821. Since its publication, several claims - both contemporaneous and recent - have been made suggesting that Frankenstein was actually written by Mary's husband, Percy Bysshe Shelley. We review the background of this controversy and then apply modern techniques from computational stylometry to determine who the true author is. Based on our analysis, we find extremely substantial evidence that Mary Shelley is indeed the true author of Frankenstein, and that it is very improbable that Percy Bysshe Shelley played a heavy role in composing the text. While our finding confirms mainstream scholarly opinion regarding Frankenstein, our analysis is the first application of stylometric techniques to this question and provides strong objective grounds for favouring Shelley by freeing the question from some of the politics which have traditionally accompanied it.

stat.AP

The Ancestor Hawkes Process with an Application to Group Chat Data

The Hawkes process is used to model point process data where events occur in clusters and bursts. In a standard multivariate Hawkes process, every event that occurs in a dimension has an equal impact on the process intensity. However, this assumption is unrealistic in applications such as the modelling of message cascades where the effect of an event depends on whether it was the initiator or a member of a particular cluster. To alleviate this, we introduce a new Hawkes process model, the Ancestor Hawkes process, which allows the impact of each event to vary based on its origin. The relevance of the Ancestor Hawkes process is showcased on real data from a 9-person group chat, where our proposed approach reveals individual response preferences. Crucially, this is achieved in a privacy-conscious manner, as only the sender and the time at which a message was sent -- but not its content -- are utilised. These nuances of messaging cascades are missed by the standard Hawkes process, but are relevant for studying latent interaction structure and for personalised notification management.

stat.ME

Bayesian Estimation of the ETAS Model for Earthquake Occurrences

The Epidemic Type Aftershock Sequence (ETAS) model is one of the most widely-used approaches to seismic forecasting. However most studies of ETAS use point estimates for the model parameters, which ignores the inherent uncertainty that arises from estimating these from historical earthquake catalogs, resulting in misleadingly optimistic forecasts. In contrast, Bayesian statistics allows parameter uncertainty to be explicitly represented, and fed into the forecast distribution. Despite its growing popularity in seismology, the application of Bayesian statistics to the ETAS model has been limited by the complex nature of the resulting posterior distribution which makes it infeasible to apply on catalogs containing more than a few hundred earthquakes. To combat this, we develop a new framework for estimating the ETAS model in a fully Bayesian manner, which can be efficiently scaled up to large catalogs containing thousands of earthquakes. We also provide easy-to-use software which implements our method.

stat.AP

Understanding the Heavy Tailed Dynamics in Human Behavior

The recent availability of electronic datasets containing large volumes of communication data has made it possible to study human behavior on a larger scale than ever before. From this, it has been discovered that across a diverse range of data sets, the inter-event times between consecutive communication events obey heavy tailed power law dynamics. Explaining this has proved controversial, and two distinct hypotheses have emerged. The first holds that these power laws are fundamental, and arise from the mechanisms such as priority queuing that humans use to schedule tasks. The second holds that they are a statistical artifact which only occur in aggregated data when features such as circadian rhythms and burstiness are ignored. We use a large social media data set to test these hypotheses, and find that although models that incorporate circadian rhythms and burstiness do explain part of the observed heavy tails, there is residual unexplained heavy tail behavior which suggests a more fundamental cause. Based on this, we develop a new quantitative model of human behavior which improves on existing approaches, and gives insight into the mechanisms underlying human interactions.

physics.soc-ph

Sequential Change Detection in the Presence of Unknown Parameters

It is commonly required to detect change points in sequences of random variables. In the most difficult setting of this problem, change detection must be performed sequentially with new observations being constantly received over time. Further, the parameters of both the pre- and post- change distributions may be unknown. In a recent paper by Hawkins and Zamba (2005), the sequential generalised likelihood ratio test was introduced for detecting changes in this context, under the assumption that the observations follow a Gaussian distribution. However, we show that the asymptotic approximation used in their test statistic leads to it being conservative even when a large numbers of observations is available. We propose an improved procedure which is more efficient, in the sense of detecting changes faster, in all situations. We also show that similar issues arise in other parametric change detection contexts, which we illustrate by introducing a novel monitoring procedure for sequences of Exponentially distributed random variable, which is an important topic in time-to-failure modelling.

stat.ME