SearcharxivSearch

arXiv subjects

Paul Ormerod

Publications and source records attributed to Paul Ormerod.

At least 19 recordsLinked to original sources

Text as data: a machine learning-based approach to measuring uncertainty

The Economic Policy Uncertainty index had gained considerable traction with both academics and policy practitioners. Here, we analyse news feed data to construct a simple, general measure of uncertainty in the United States using a highly cited machine learning methodology. Over the period January 1996 through May 2020, we show that the series unequivocally Granger-causes the EPU and there is no Granger-causality in the reverse direction

econ.EM

Explaining herding and volatility in the cyclical price dynamics of urban housing markets using a large scale agent-based model

Urban housing markets, along with markets of other assets, universally exhibit periods of strong price increases followed by sharp corrections. The mechanisms generating such non-linearities are not yet well understood. We develop an agent-based model populated by a large number of heterogeneous households. The agents' behavior is compatible with economic rationality, with the trend-following behavior found to be essential in replicating market dynamics. The model is calibrated using several large and distributed datasets of the Greater Sydney region (demographic, economic and financial) across three specific and diverse periods since 2006. The model is not only capable of explaining price dynamics during these periods, but also reproduces the novel behavior actually observed immediately prior to the market peak in 2017, namely a sharp increase in the variability of prices. This novel behavior is related to a combination of trend-following aptitude of the household agents (rational herding) and their propensity to borrow.

q-fin.CP

Text as Data: Real-time Measurement of Economic Welfare

Economists are showing increasing interest in the use of text as an input to economic research. Here, we analyse online text to construct a real time metric of welfare. For purposes of description, we call it the Feel Good Factor (FGF). The particular example used to illustrate the concept is confined to data from the London area, but the methodology is readily generalisable to other geographical areas. The FGF illustrates the use of online data to create a measure of welfare which is not based, as GDP is, on value added in a market-oriented economy. There is already a large literature which measures wellbeing/happiness. But this relies on conventional survey approaches, and hence on the stated preferences of respondents. In unstructured online media text, users reveal their emotions in ways analogous to the principle of revealed preference in consumer demand theory. The analysis of online media offers further advantages over conventional survey-based measures of sentiment or well-being. It can be carried out in real time rather than with the lags which are involved in survey approaches. In addition, it is very much cheaper.

econ.GN

Understanding the Great Recession Using Machine Learning Algorithms

Nyman and Ormerod (2017) show that the machine learning technique of random forests has the potential to give early warning of recessions. Applying the approach to a small set of financial variables and replicating as far as possible a genuine ex ante forecasting situation, over the period since 1990 the accuracy of the four-step ahead predictions is distinctly superior to those actually made by the professional forecasters. Here we extend the analysis by examining the contributions made to the Great Recession of the late 2000s by each of the explanatory variables. We disaggregate private sector debt into its household and non-financial corporate components. We find that both household and non-financial corporate debt were key determinants of the Great Recession. We find a considerable degree of non-linearity in the explanatory models. In contrast, the public sector debt to GDP ratio appears to have made very little contribution. It did rise sharply during the Great Recession, but this was as a consequence of the sharp fall in economic activity rather than it being a cause. We obtain similar results for both the United States and the United Kingdom.

econ.GN

Predicting Economic Recessions Using Machine Learning Algorithms

Even at the beginning of 2008, the economic recession of 2008/09 was not being predicted. The failure to predict recessions is a persistent theme in economic forecasting. The Survey of Professional Forecasters (SPF) provides data on predictions made for the growth of total output, GDP, in the United States for one, two, three and four quarters ahead since the end of the 1960s. Over a three quarters ahead horizon, the mean prediction made for GDP growth has never been negative over this period. The correlation between the mean SPF three quarters ahead forecast and the data is very low, and over the most recent 25 years is not significantly different from zero. Here, we show that the machine learning technique of random forests has the potential to give early warning of recessions. We use a small set of explanatory variables from financial markets which would have been available to a forecaster at the time of making the forecast. We train the algorithm over the 1970Q2-1990Q1 period, and make predictions one, three and six quarters ahead. We then re-train over 1970Q2-1990Q2 and make a further set of predictions, and so on. We did not attempt any optimisation of predictions, using only the default input parameters to the algorithm we downloaded in the package R. We compare the predictions made from 1990 to the present with the actual data. One quarter ahead, the algorithm is not able to improve on the SPF predictions. Three and six quarters ahead, the correlations between actual and predicted are low, but they are very significantly different from zero. Although the timing is slightly wrong, a serious downturn in the first half of 2009 could have been predicted six quarters ahead in late 2007. The algorithm never predicts a recession when one did not occur. We obtain even stronger results with random forest machine learning techniques in the case of the United Kingdom.

q-fin.GN

Measuring Financial Sentiment to Predict Financial Instability: A New Approach based on Text Analysis

Following the financial crisis of the late 2000s, policy makers have shown considerable interest in monitoring financial stability. Several central banks now publish indices of financial stress, which are essentially based upon market related data. In this paper, we examine the potential for improving the indices by deriving information about emotion shifts in the economy. We report on a new approach, based on the content analysis of very large text databases, and termed directed algorithmic text analysis. The algorithm identifies, very rapidly, shifts through time in the relations between two core emotional groups. The method is robust. The same word-list is used to identify the two emotion groups across different studies. Membership of the words in the lists has been validated in psychological experiments. The words consist of everyday English words with no specific economic meaning. Initial results show promise. An emotion index capturing shifts between the two emotion groups in texts potentially referring to the whole US economy improves the one-quarter ahead consensus forecasts for real GDP growth. More specifically, the same indices are shown to Granger cause both the Cleveland and St Louis Federal Reserve Indices of Financial Stress.

q-fin.GN

Nowcasting economic and social data: when and why search engine data fails, an illustration using Google Flu Trends

Obtaining an accurate picture of the current state of the economy is particularly important to central banks and finance ministries, and of epidemics to health ministries. There is increasing interest in the use of search engine data to provide such 'nowcasts' of social and economic indicators. However, people may search for a phrase because they independently want the information, or they may search simply because many others are searching for it. We consider the effect of the motivation for searching on the accuracy of forecasts made using search engine data of contemporaneous social and economic indicators. We illustrate the implications for forecasting accuracy using four episodes in which Google Flu Trends data gave accurate predictions of actual flu cases, and four in which the search data over-predicted considerably. Using a standard statistical methodology, the Bass diffusion model, we show that the independent search for information motive was much stronger in the cases of accurate prediction than in the inaccurate ones. Social influence, the fact that people may search for a phrase simply because many others are, was much stronger in the inaccurate compared to the accurate cases. Search engine data may therefore be an unreliable predictor of contemporaneous indicators when social influence on the decision to search is strong.

physics.soc-ph

Big Data, Socio-Psychological Theory, Algorithmic Text Analysis and Predicting the Michigan Consumer Sentiment Index

We describe an exercise of using Big Data to predict the Michigan Consumer Sentiment Index, a widely used indicator of the state of confidence in the US economy. We carry out the exercise from a pure ex ante perspective. We use the methodology of algorithmic text analysis of an archive of brokers' reports over the period June 2010 through June 2013. The search is directed by the social-psychological theory of agent behaviour, namely conviction narrative theory. We compare one month ahead forecasts generated this way over a 15 month period with the forecasts reported for the consensus predictions of Wall Street economists. The former give much more accurate predictions, getting the direction of change correct on 12 of the 15 occasions compared to only 7 for the consensus predictions. We show that the approach retains significant predictive power even over a four month ahead horizon.

q-fin.ST

Social network markets: the influence of network structure when consumers face decisions over many similar choices

In social network markets, the act of consumer choice in these industries is governed not just by the set of incentives described by conventional consumer demand theory, but by the choices of others in which an individual's payoff is an explicit function of the actions of others. We observe two key empirical features of outcomes in social networked markets. First, a highly right-skewed, non-Gaussian distribution of the number of times competing alternatives are selected at a point in time. Second, there is turnover in the rankings of popularity over time. We show here that such outcomes can arise either when there is no alternative which exhibits inherent superiority in its attributes, or when agents find it very difficult to discern any differences in quality amongst the alternatives which are available so that it is as if no superiority exists. These features appear to obtain, as a reasonable approximation, in many social network markets. We examine the impact of network structure on both the rank-size distribution of choices at a point in time, and on the life spans of the most popular choices. We show that a key influence on outcomes is the extent to which the network follows a hierarchical structure. It is the social network properties of the markets, the meso-level structure, which determine outcomes rather than the objective attributes of the products.

cs.SI

Ex ante prediction of cascade sizes on networks of agents facing binary outcomes

We consider in this paper the potential for ex ante prediction of the cascade size in a model of binary choice with externalities (Schelling 1973, Watts 2002). Agents are connected on a network and can be in one of two states of the world, 0 or 1. Initially, all are in state 0 and a small number of seeds are selected at random to switch to state1. A simple threshold rule specifies whether other agents switch subsequently. The cascade size (the percolation) is the proportion of all agents which eventually switches to state 1. We select information on the connectivity of the initial seeds, the connectivity of the agents to which they are connected, the thresholds of these latter agents, and the thresholds of the agents to which these are connected. We obtain results for random, small world and scale -free networks with different network parameters and numbers of initial seeds. The results are robust with respect to these factors. We perform least squares regression of the logit transformation of the cascade size (Hosmer and Lemeshow 1989) on these potential explanatory variables. We find considerable explanatory power for the ex ante prediction of cascade sizes. For the random networks, on average 32 per cent of the variance of the cascade sizes is explained, 40 per cent for the small world and 46 per cent for the scale-free. The connectivity variables are hardly ever significant in the regressions, whether relating to the seeds themselves or to the agents connected to the seeds. In contrast, the information on the thresholds of agents contains much more explanatory power. This supports the conjecture of Watts and Dodds (2007.) that large cascades are driven by a small mass of easily influenced agents.

cs.SI

Predictability and Prediction for an Experimental Cultural Market

Individuals are often influenced by the behavior of others, for instance because they wish to obtain the benefits of coordinated actions or infer otherwise inaccessible information. In such situations this social influence decreases the ex ante predictability of the ensuing social dynamics. We claim that, interestingly, these same social forces can increase the extent to which the outcome of a social process can be predicted very early in the process. This paper explores this claim through a theoretical and empirical analysis of the experimental music market described and analyzed in [1]. We propose a very simple model for this music market, assess the predictability of market outcomes through formal analysis of the model, and use insights derived through this analysis to develop algorithms for predicting market share winners, and their ultimate market shares, in the very early stages of the market. The utility of these predictive algorithms is illustrated through analysis of the experimental music market data sets [2].

nlin.AO

Predictability in an unpredictable artificial cultural market

In social, economic and cultural situations in which the decisions of individuals are influenced directly by the decisions of others, there appears to be an inherently high level of ex ante unpredictability. In cultural markets such as films, songs and books, well-informed experts routinely make predictions which turn out to be incorrect. We examine the extent to which the existence of social influence may, somewhat paradoxically, increase the extent to which winners can be identified at a very early stage in the process. Once the process of choice has begun, only a very small number of decisions may be necessary to give a reasonable prospect of being able to identify the eventual winner. We illustrate this by an analysis of the music download experiments of Salganik et.al. (2006). We derive a rule for early identification of the eventual winner. Although not perfect, it gives considerable practical success. We validate the rule by applying it to similar data not used in the process of constructing the rule.

physics.soc-ph

Were people imitating others or exercising rational choice in on-line searches for 'swine flu'?

Two general patterns have been identified for the adoption and subsequent abandonment of ideas or products within a population. One is symmetric, so that concepts or products which are adopted rapidly then decline rapidly from their peak, and those which are slower to move to their peak decline more slowly. The other is asymmetric, where the decline from the peak is considerably slower than is the rise to the peak, and vice versa. We posit that these contrasting patterns arise from two fundamentally different modes of behaviour which are used by humans in making choices in different contexts. Namely, choice based on the imitation of the choices of others versus purposeful selection based upon the inherent attributes of the concept or product. We illustrate the proposition with the example of internet searches for the phrase 'swine flu' in a wide range of countries across the world. The methodology offers a general heuristic for distinguishing between these two general and contrasting modes of behavioural choice.

physics.soc-ph

An evolutionary model of long tailed distributions in the social sciences

Studies of collective human behavior in the social sciences, often grounded in details of actions by individuals, have much to offer `social' models from the physical sciences concerning elegant statistical regularities. Drawing on behavioral studies of social influence, we present a parsimonious, stochastic model, which generates an entire family of real-world right-skew socio-economic distributions, including exponential, winner-take-all, power law tails of varying exponents and power laws across the whole data. The widely used Albert-Barabasi model of preferential attachment is simply a special case of this much more general model. In addition, the model produces the continuous turnover observed empirically within those distributions. Previous preferential attachment models have generated specific distributions with turnover using arbitrary add-on rules, but turnover is an inherent feature of our model. The model also replicates an intriguing new relationship, observed across a range of empirical studies, between the power law exponent and the proportion of data represented.

physics.soc-ph

Shelf space strategy in long-tail markets

The Internet is known to have had a powerful impact on on-line retailer strategies in markets characterised by long-tail distribution of sales. Such retailers can exploit the long tail of the market, since they are effectively without physical limit on the number of choices on offer. Here we examine two extensions of this phenomenon. First, we introduce turnover into the long-tail distribution of sales. Although over any given period such as a week or a month, the distribution is right-skewed and often power law distributed, over time there is considerable turnover in the rankings of sales of individual products. Second, we establish some initial results on the implications for shelf-space strategy of physical retailers in such markets.

q-fin.GN

Random matrix theory and the evolution of business cycle synchronisation 1886-2006

The major study by Bordo and Helbing (2003) analyses the business cycle in Western economies 1881-2001. They examine four distinct periods in economic history, and conclude that there is a secular trend towards greater synchronisation for much of the 20th century. Their analysis, in common with the standard economic literature on business cycle synchronisation, relies upon the estimation of an empirical correlation matrix of time series data of macroeconomic aggregates. However because of the small number of observations and economies, the empirical correlation matrix may contain considerable noise. Random matrix theory was developed to overcome this problem. I use random matrix theory, and the associated technique of agglomerative hierarchical clustering, to examine the evolution of business cycle synchronisation between the capitalist economies in the long-run. Contrary to the findings of Bordo and Helbing, it is not possible to speak of a 'secular trend' towards greater synchronisation over the period as a whole. During the pre-First World War period, the cross-country correlations of annual real GDP growth are indistinguishable from those which could be generated by a purely random matrix. The periods 1920-38 and 1948-72 do show a certain degree of synchronisation, but it is very weak. In particular, the cycles of the major economies cannot be said to be synchronised. Such synchronisation as exists in the overall data is due to meaningful co-movements in sub-groups. So the degree of synchronisation has evolved fitfully. It is only in the most recent 1973-2006 period that we can speak meaningfully of anything resembling an international business cycle.

q-fin.ST

The evolution of EU business cycle synchronisation 1981-2007

Most of the analytical techniques used in the business cycle synchronisation literature rely upon the estimation of an empirical correlation matrix of time series data of macroeconomic aggregates, real GDP usually being the key variable. But the small number of available observations and small number of economies mean that the empirical correlation matrix may contain considerable noise. Random matrix theory was developed in physics to overcome this problem. The largest eigenvalue of the correlation matrix informs us directly about the degree to which movements of the economies are genuinely correlated. The evolution of business cycle synchronisation can be analysed with the temporal evolution of the largest eigenvalue over a fixed window of data. I analyse quarterly real GDP data 1981Q1-2008Q1 for the core EU economies - Germany, France, Italy, Spain, Netherlands, Belgium - along with the UK, which is a member of the EU but not the Euro, and the US as a comparator. The core EU economies have shown varying but strong synchronisation over the whole period. In contrast, the UK and the US are much more synchronised with each other than they are with the core EU economies.

q-fin.ST

Global recessions as a cascade phenomenon with heterogenous, interacting agents

I examine global recessions as a cascade phenomenon. In other words, how recessions arising in one or more countries might percolate across a network of connected economies. A heterogeneous agent based model is set up in which the agents are Western economies. A country has a probability of entering a recession in any given year and one of emerging from it the next. In addition, the agents have a unique threshold propensity to import a recession from the agents with which they have the strongest connections. They are connected on a small world topology, and an agent's neighbours at any time are either in (state 1) or out (state 0) of recession. If the weighted sum exceeds the threshold, the agent goes into recession. Annual real GDP growth for 17 Western countries 1871-2006 is used as the data set. The distribution of the number of countries in recession in any given year is exponential, as is the duration of recessions within individual countries. The model is calibrated against these two facts, plus the 'wait time' between recessions. It is able to replicate them successfully. The network structure is essential for the agents to replicate the stylised facts. The country-specific probabilities of entering and emerging from recession by themselves give results very different to the actual data.

q-fin.GN