SearcharxivSearch

arXiv subjects

Trevor Fenner

Publications and source records attributed to Trevor Fenner.

At least 19 recordsLinked to original sources

A stochastic differential equation approach to the analysis of the UK 2017 and 2019 general election polls

Human dynamics and sociophysics build on statistical models that can shed light on and add to our understanding of social phenomena. We propose a generative model based on a stochastic differential equation that enables us to model the opinion polls leading up to the UK 2017 and 2019 general elections, and to make predictions relating to the actual result of the elections. After a brief analysis of the time series of the poll results, we provide empirical evidence that the gamma distribution, which is often used in financial modelling, fits the marginal distribution of this time series. We demonstrate that the proposed poll-based forecasting model may improve upon predictions based solely on polls. The method uses the Euler-Maruyama method to simulate the time series, measuring the prediction error with the mean absolute error and the root mean square error, and as such could be used as part of a toolkit for forecasting elections.

physics.soc-ph

Fast Generation of Unlabelled Free Trees using Weight Sequences

In this paper, we introduce a new representation for ordered trees, the weight sequence representation. We then use this to construct new representations for both rooted trees and free trees, namely the canonical weight sequence representation. We construct algorithms for generating the weight sequence representations for all rooted and free trees of order n, and then add a number of modifications to improve the efficiency of the algorithms. Python implementations of the algorithms incorporate further improvements by using generators to avoid having to store the long lists of trees returned by the recursive calls, as well as caching the lists for rooted trees of small order, thereby eliminating many of the recursive calls. We further show how the algorithm can be modifed to generate adjacency list and adjacency matrix representations for free trees. We compared the run-times of our Python implementation for generating free trees with the Python implementation of the well-known WROM algorithm taken from NetworkX. The implementation of our algorithm is over four times as fast as the implementation of the WROM algorithm. The run-times for generating adjacency lists and matrices are somewhat longer than those for weight sequences, but are still over three times as fast as the corresponding implementations of the WROM algorithm.

cs.DS

Supercards, Sunshines and Caterpillar Graphs

The vertex-deleted subgraph G-v, obtained from the graph G by deleting the vertex v and all edges incident to v, is called a card of G. The deck of G is the multiset of its unlabelled cards. The number of common cards b(G,H) of G and H is the cardinality of the multiset intersection of the decks of G and H. A supercard G+ of G and H is a graph whose deck contains at least one card isomorphic to G and at least one card isomorphic to H. We show how maximum sets of common cards of G and H correspond to certain sets of permutations of the vertices of a supercard, which we call maximum saturating sets. We apply the theory of supercards and maximum saturating sets to the case when G is a sunshine graph and H is a caterpillar graph. We show that, for large enough n, there exists some maximum saturating set that contains at least b(G,H)-2 automorphisms of G+, and that this subset is always isomorphic to either a cyclic or dihedral group. We prove that b(G,H)<=2(n+1)/5 for large enough n, and that there exists a unique family of pairs of graphs that attain this bound. We further show that, in this case, the corresponding maximum saturating set is isomorphic to the dihedral group.

math.CO

A Problem in Human Dynamics: Modelling the Population Density of a Social Space

Human dynamics and sociophysics suggest statistical models that may explain and provide us with better insight into social phenomena. Here we tackle the problem of determining the distribution of the population density of a social space over time by modelling the dynamics of agents entering and exiting the space as a birth-death process. We show that, for a simple agent-based model in which the probabilities of entering and exiting the space depends on the number of agents currently present in the space, the population density of the space follows a gamma distribution. We also provide empirical evidence supporting the validity of the model by applying it to a data set of occupancy traces of a common space in an office building.

physics.soc-ph

Characterisation of the $χ$-index and the $rec$-index

Axiomatic characterisation of a bibliometric index provides insight into the properties that the index satisfies and facilitates the comparison of different indices. A geometric generalisation of the $h$-index, called the $χ$-index, has recently been proposed to address some of the problems with the $h$-index, in particular, the fact that it is not scale invariant, i.e., multiplying the number of citations of each publication by a positive constant may change the relative ranking of two researchers. While the square of the $h$-index is the area of the largest square under the citation curve of a researcher, the square of the $χ$-index, which we call the $rec$-index (or {\em rectangle}-index), is the area of the largest rectangle under the citation curve. Our main contribution here is to provide a characterisation of the $rec$-index via three properties: {\em monotonicity}, {\em uniform citation} and {\em uniform equivalence}. Monotonicity is a natural property that we would expect any bibliometric index to satisfy, while the other two properties constrain the value of the $rec$-index to be the area of the largest rectangle under the citation curve. The $rec$-index also allows us to distinguish between {\em influential} researchers who have relatively few, but highly-cited, publications and {\em prolific} researchers who have many, but less-cited, publications.

cs.DL

A stochastic differential equation approach to the analysis of the UK 2016 EU referendum polls

Human dynamics and sociophysics suggest statistical models that may explain and provide us with better insight into social phenomena. Here we propose a generative model based on a stochastic differential equation that allows us to analyse the polls leading up to the UK 2016 EU referendum. After a preliminary analysis of the time series of poll results, we provide empirical evidence that the beta distribution, which is a natural choice when modelling proportions, fits the marginal distribution of this time series. We also provide evidence of the predictive power of the proposed model.

physics.soc-ph

A multiplicative process for generating a beta-like survival function with application to the UK 2016 EU referendum results

Human dynamics and sociophysics suggest statistical models that may explain and provide us with better insight into social phenomena. Contextual and selection effects tend to produce extreme values in the tails of rank-ordered distributions of both census data and district-level election outcomes. Models that account for this nonlinearity generally outperform linear models. Fitting nonlinear functions based on rank-ordering census and election data therefore improves the fit of aggregate voting models. This may help improve ecological inference, as well as election forecasting in majoritarian systems. We propose a generative multiplicative decrease model that gives rise to a rank-order distribution, and facilitates the analysis of the recent UK EU referendum results. We supply empirical evidence that the beta-like survival function, which can be generated directly from our model, is a close fit to the referendum results, and also may have predictive value when covariate data are available.

physics.soc-ph

A multiplicative process for generating the rank-order distribution of UK election results

Human dynamics and sociophysics suggest statistical models that may explain and provide us with a better understanding of social phenomena. Here we propose a generative multiplicative decrease model that gives rise to a rank-order distribution and allows us to analyse the results of the last three UK parliamentary elections. We provide empirical evidence that the additive Weibull distribution, which can be generated from our model, is a close fit to the electoral data, offering a novel interpretation of the recent election results.

physics.soc-ph

A stochastic evolutionary model generating a mixture of exponential distributions

Recent interest in human dynamics has stimulated the investigation of the stochastic processes that explain human behaviour in various contexts, such as mobile phone networks and social media. In this paper, we extend the stochastic urn-based model proposed in \cite{FENN15} so that it can generate mixture models,in particular, a mixture of exponential distributions. The model is designed to capture the dynamics of survival analysis, traditionally employed in clinical trials, reliability analysis in engineering, and more recently in the analysis of large data sets recording human dynamics. The mixture modelling approach, which is relatively simple and well understood, is very effective in capturing heterogeneity in data. We provide empirical evidence for the validity of the model, using a data set of popular search engine queries collected over a period of 114 months. We show that the survival function of these queries is closely matched by the exponential mixture solution for our model.

physics.soc-ph

A stochastic evolutionary model for capturing human dynamics

The recent interest in human dynamics has led researchers to investigate the stochastic processes that explain human behaviour in various contexts. Here we propose a generative model to capture the dynamics of survival analysis, traditionally employed in clinical trials and reliability analysis in engineering. We derive a general solution for the model in the form of a product, and then a continuous approximation to the solution via the renewal equation describing age-structured population dynamics. This enables us to model a wide range of survival distributions, according to the choice of the mortality distribution. We provide empirical evidence for the validity of the model from a longitudinal data set of popular search engine queries over 114 months, showing that the survival function of these queries is closely matched by the solution for our model with power-law mortality.

physics.soc-ph

A stochastic evolutionary model for survival dynamics

The recent interest in human dynamics has led researchers to investigate the stochastic processes that explain human behaviour in different contexts. Here we propose a generative model to capture the essential dynamics of survival analysis, traditionally employed in clinical trials and reliability analysis in engineering. In our model, the only implicit assumption made is that the longer an actor has been in the system, the more likely it is to have failed. We derive a power-law distribution for the process and provide preliminary empirical evidence for the validity of the model from two well-known survival analysis data sets.

physics.soc-ph

A bibliometric index based on the complete list of cited publications

We propose a new index, the $j$-index, which is defined for an author as the sum of the square roots of the numbers of citations to each of the author's publications. The idea behind the $j$-index it to remedy a drawback of the $h$-index $-$ that the $h$-index does not take into account the full citation record of a researcher. The square root function is motivated by our desire to avoid the possible bias that may occur with a simple sum when an author has several very highly cited papers. We compare the $j$-index to the $h$-index, the $g$-index and the total citation count for three subject areas using several association measures. Our results indicate that that the association between the $j$-index and the other indices varies according to the subject area. One explanation of this variation may be due to the proportion of citations to publications of the researcher that are in the $h$-core. The $j$-index is {\em not} an $h$-index variant, and as such is intended to complement rather than necessarily replace the $h$-index and other bibliometric indicators, thus providing a more complete picture of a researcher's achievements.

cs.DL

A Discrete Evolutionary Model for Chess Players' Ratings

The Elo system for rating chess players, also used in other games and sports, was adopted by the World Chess Federation over four decades ago. Although not without controversy, it is accepted as generally reliable and provides a method for assessing players' strengths and ranking them in official tournaments. It is generally accepted that the distribution of players' rating data is approximately normal but, to date, no stochastic model of how the distribution might have arisen has been proposed. We propose such an evolutionary stochastic model, which models the arrival of players into the rating pool, the games they play against each other, and how the results of these games affect their ratings. Using a continuous approximation to the discrete model, we derive the distribution for players' ratings at time $t$ as a normal distribution, where the variance increases in time as a logarithmic function of $t$. We validate the model using published rating data from 2007 to 2010, showing that the parameters obtained from the data can be recovered through simulations of the stochastic model. The distribution of players' ratings is only approximately normal and has been shown to have a small negative skew. We show how to modify our evolutionary stochastic model to take this skewness into account, and we validate the modified model using the published official rating data.

physics.soc-ph

A Methodology for Learning Players' Styles from Game Records

We describe a preliminary investigation into learning a Chess player's style from game records. The method is based on attempting to learn features of a player's individual evaluation function using the method of temporal differences, with the aid of a conventional Chess engine architecture. Some encouraging results were obtained in learning the styles of two recent Chess world champions, and we report on our attempt to use the learnt styles to discriminate between the players from game records by trying to detect who was playing white and who was playing black. We also discuss some limitations of our approach and propose possible directions for future research. The method we have presented may also be applicable to other strategic games, and may even be generalisable to other domains where sequences of agents' actions are recorded.

cs.AI

Modelling the Navigation Potential of a Web Page

Suppose that you are navigating in ``hyperspace'' and you have reached a web page with several outgoing links you could choose to follow. Which link should you choose in such an online scenario? When you are not sure where the information you require resides, you will initiate a navigation session. This involves pruning some of the links and following one of the others, where more pruning is likely to happen the deeper you navigate. In terms of decision making, the utility of navigation diminishes with distance until finally the utility drops to zero and the session is terminated. Under this model of navigation, we call the number of nodes that are available after pruning, for browsing within a session, the {\em potential gain} of the starting web page. Thus the parameters that effect the potential gain are the local branching factor with respect to the starting web page and the discount factor. We first consider the case when the discounting factor is geometric. We show that the distribution of the effective number of links that the user can follow at each navigation step after pruning, i.e. the number of nodes added to the potential gain at that step, is given by the {\em erf} function. We derive an approximation to the potential gain of a web page and show that this is numerically a very accurate estimate. We then consider a harmonic discounting factor and show that, in this case, the potential gain at each step is closely related to the probability density function for the Poisson distribution. The potential gain has been applied to web navigation where it helps the user to choose a good starting point for initiating a navigation session. Another application is in social network analysis, where the potential gain could provide a novel measure of centrality.

physics.soc-ph

A Stochastic Evolutionary Growth Model for Social Networks

We present a stochastic model for a social network, where new actors may join the network, existing actors may become inactive and, at a later stage, reactivate themselves. Our model captures the evolution of the network, assuming that actors attain new relations or become active according to the preferential attachment rule. We derive the mean-field equations for this stochastic model and show that, asymptotically, the distribution of actors obeys a power-law distribution. In particular, the model applies to social networks such as wireless local area networks, where users connect to access-points, and peer-to-peer networks where users connect to each other. As a proof of concept, we demonstrate the validity of our model empirically by analysing a public log containing traces from a wireless network at Dartmouth College over a period of three years. Analysing the data processed according to our model, we demonstrate that the distribution of user accesses is asymptotically a power-law distribution.

physics.soc-ph

A Model for Collaboration Networks Giving Rise to a Power Law Distribution with an Exponential Cutoff

Recently several authors have proposed stochastic evolutionary models for the growth of complex networks that give rise to power-law distributions. These models are based on the notion of preferential attachment leading to the ``rich get richer'' phenomenon. Despite the generality of the proposed stochastic models, there are still some unexplained phenomena, which may arise due to the limited size of networks such as protein, e-mail, actor and collaboration networks. Such networks may in fact exhibit an exponential cutoff in the power-law scaling, although this cutoff may only be observable in the tail of the distribution for extremely large networks. We propose a modification of the basic stochastic evolutionary model, so that after a node is chosen preferentially, say according to the number of its inlinks, there is a small probability that this node will become inactive. We show that as a result of this modification, by viewing the stochastic process in terms of an urn transfer model, we obtain a power-law distribution with an exponential cutoff. Unlike many other models, the current model can capture instances where the exponent of the distribution is less than or equal to two. As a proof of concept, we demonstrate the consistency of our model empirically by analysing the Mathematical Research collaboration network, the distribution of which is known to follow a power law with an exponential cutoff.

physics.soc-ph

A Stochastic Evolutionary Model Exhibiting Power-Law Behaviour with an Exponential Cutoff

Recently several authors have proposed stochastic evolutionary models for the growth of complex networks that give rise to power-law distributions. These models are based on the notion of preferential attachment leading to the ``rich get richer'' phenomenon. Despite the generality of the proposed stochastic models, there are still some unexplained phenomena, which may arise due to the limited size of networks such as protein and e-mail networks. Such networks may in fact exhibit an exponential cutoff in the power-law scaling, although this cutoff may only be observable in the tail of the distribution for extremely large networks. We propose a modification of the basic stochastic evolutionary model, so that after a node is chosen preferentially, say according to the number of its inlinks, there is a small probability that this node will be discarded. We show that as a result of this modification, by viewing the stochastic process in terms of an urn transfer model, we obtain a power-law distribution with an exponential cutoff. Unlike many other models, the current model can capture instances where the exponent of the distribution is less than or equal to two. As a proof of concept, we demonstrate the consistency of our model by analysing a yeast protein interaction network, the distribution of which is known to follow a power law with an exponential cutoff.

cond-mat.soft