SearcharxivSearch

arXiv subjects

Ole Peters

Publications and source records attributed to Ole Peters.

At least 19 recordsLinked to original sources

Reinforcement learning with non-ergodic reward increments: robustness via ergodicity transformations

Envisioned application areas for reinforcement learning (RL) include autonomous driving, precision agriculture, and finance, which all require RL agents to make decisions in the real world. A significant challenge hindering the adoption of RL methods in these domains is the non-robustness of conventional algorithms. In particular, the focus of RL is typically on the expected value of the return. The expected value is the average over the statistical ensemble of infinitely many trajectories, which can be uninformative about the performance of the average individual. For instance, when we have a heavy-tailed return distribution, the ensemble average can be dominated by rare extreme events. Consequently, optimizing the expected value can lead to policies that yield exceptionally high returns with a probability that approaches zero but almost surely result in catastrophic outcomes in single long trajectories. In this paper, we develop an algorithm that lets RL agents optimize the long-term performance of individual trajectories. The algorithm enables the agents to learn robust policies, which we show in an instructive example with a heavy-tailed return distribution and standard RL benchmarks. The key element of the algorithm is a transformation that we learn from data. This transformation turns the time series of collected returns into one for whose increments expected value and the average over a long trajectory coincide. Optimizing these increments results in robust policies.

cs.LG

The Two Growth Rates of the Economy

Economic growth is measured as the rate of relative change in gross domestic product (GDP) per capita. Yet, when incomes follow random multiplicative growth, the ensemble-average (GDP per capita) growth rate is higher than the time-average growth rate achieved by each individual in the long run. This mathematical fact is the starting point of ergodicity economics. Using the atypically high ensemble-average growth rate as the principal growth measure creates an incomplete picture. Policymaking would be better informed by reporting both ensemble-average and time-average growth rates. We analyse rigorously these growth rates and describe their evolution in the United States and France over the last fifty years. The difference between the two growth rates gives rise to a natural measure of income inequality, equal to the mean logarithmic deviation. Despite being estimated as the average of individual income growth rates, the time-average growth rate is independent of income mobility.

econ.GN

Leverage efficiency

Peters (2011a) defined an optimal leverage which maximizes the time-average growth rate of an investment held at constant leverage. It was hypothesized that this optimal leverage is attracted to 1, such that, e.g., leveraging an investment in the market portfolio cannot yield long-term outperformance. This places a strong constraint on the stochastic properties of prices of traded assets, which we call "leverage efficiency." Market conditions that deviate from leverage efficiency are unstable and may create leverage-driven bubbles. Here we expand on the hypothesis and its implications. These include a theory of noise that explains how systemic stability rules out smooth price changes at any pricing frequency; a resolution of the so-called equity premium puzzle; a protocol for central bank interest rate setting to avoid leverage-driven price instabilities; and a method for detecting fraudulent investment schemes by exploiting differences between the stochastic properties of their prices and those of legitimately-traded assets. To submit the hypothesis to a rigorous test we choose price data from different assets: the S&P500 index, Bitcoin, Berkshire Hathaway Inc., and Bernard L. Madoff Investment Securities LLC. Analysis of these data supports the hypothesis.

q-fin.GN

What are we weighting for? A mechanistic model for probability weighting

Behavioural economics provides labels for patterns in human economic behaviour. Probability weighting is one such label. It expresses a mismatch between probabilities used in a formal model of a decision (i.e. model parameters) and probabilities inferred from real people's decisions (the same parameters estimated empirically). The inferred probabilities are called "decision weights." It is considered a robust experimental finding that decision weights are higher than probabilities for rare events, and (necessarily, through normalisation) lower than probabilities for common events. Typically this is presented as a cognitive bias, i.e. an error of judgement by the person. Here we point out that the same observation can be described differently: broadly speaking, probability weighting means that a decision maker has greater uncertainty about the world than the observer. We offer a plausible mechanism whereby such differences in uncertainty arise naturally: when a decision maker must estimate probabilities as frequencies in a time series while the observer knows them a priori. This suggests an alternative presentation of probability weighting as a principled response by a decision maker to uncertainties unaccounted for in an observer's model.

econ.TH

An evolutionary advantage of cooperation

Cooperation is a persistent behavioral pattern of entities pooling and sharing resources. Its ubiquity in nature poses a conundrum. Whenever two entities cooperate, one must willingly relinquish something of value to the other. Why is this apparent altruism favored in evolution? Classical solutions assume a net fitness gain in a cooperative transaction which, through reciprocity or relatedness, finds its way back from recipient to donor. We seek the source of this fitness gain. Our analysis rests on the insight that evolutionary processes are typically multiplicative and noisy. Fluctuations have a net negative effect on the long-time growth rate of resources but no effect on the growth rate of their expectation value. This is an example of non-ergodicity. By reducing the amplitude of fluctuations, pooling and sharing increases the long-time growth rate for cooperating entities, meaning that cooperators outgrow similar non-cooperators. We identify this increase in growth rate as the net fitness gain, consistent with the concept of geometric mean fitness in the biological literature. This constitutes a fundamental mechanism for the evolution of cooperation. Its minimal assumptions make it a candidate explanation of cooperation in settings too simple for other fitness gains, such as emergent function and specialization, to be probable. One such example is the transition from single cells to early multicellular life.

nlin.AO

The sum of log-normal variates in geometric Brownian motion

Geometric Brownian motion (GBM) is a key model for representing self-reproducing entities. Self-reproduction may be considered the definition of life [5], and the dynamics it induces are of interest to those concerned with living systems from biology to economics. Trajectories of GBM are distributed according to the well-known log-normal density, broadening with time. However, in many applications, what's of interest is not a single trajectory but the sum, or average, of several trajectories. The distribution of these objects is more complicated. Here we show two different ways of finding their typical trajectories. We make use of an intriguing connection to spin glasses: the expected free energy of the random energy model is an average of log-normal variates. We make the mapping to GBM explicit and find that the free energy result gives qualitatively correct behavior for GBM trajectories. We then also compute the typical sum of lognormal variates using Ito calculus. This alternative route is in close quantitative agreement with numerical work.

cond-mat.stat-mech

The time interpretation of expected utility theory

Ergodicity economics is a new branch of economic theory that notes the conceptual difference between time averages and expectation values, which coincide only for ergodic observables. It postulates that individual agents maximise the time average growth rate of wealth, known widely as growth optimality. This contrasts with the dominant behavioural model in economics, expected utility theory, in which agents maximise expectation values of changes in psychologically transformed wealth. Historically, growth optimality was explored for additive and multiplicative gambles. Here we apply it to a general class of wealth dynamics, extending the range of economic situations where it may be used. Moreover, we show a correspondence between growth optimality and expected utility theory, in which the ergodicity transformation in the former is identified as the utility function in the latter. This correspondence offers a theoretical basis for choosing utility functions and predicts that wealth dynamics are strong determinants of risk preferences.

econ.GN

A recipe for irreproducible results

Recent studies have shown that many results published in peer-reviewed scientific journals are not reproducible. This raises the following question: why is it so easy to fool myself into believing that a result is reliable when in fact it is not? Using Brownian motion as a toy model, we show how this can happen if ergodicity is assumed where it is unwarranted. A measured value can appear stable when judged over time, although it is not stable across the ensemble: a different result will be obtained each time the experiment is run.

physics.data-an

Far from equilibrium: Wealth reallocation in the United States

Studies of wealth inequality often assume that an observed wealth distribution reflects a system in equilibrium. This constraint is rarely tested empirically. We introduce a simple model that allows equilibrium but does not assume it. To geometric Brownian motion (GBM) we add reallocation: all individuals contribute in proportion to their wealth and receive equal shares of the amount collected. We fit the reallocation rate parameter required for the model to reproduce observed wealth inequality in the United States from 1917 to 2012. We find that this rate was positive until the 1980s, after which it became negative and of increasing magnitude. With negative reallocation, the system cannot equilibrate. Even with the positive reallocation rates observed, equilibration is too slow to be practically relevant. Therefore, studies which assume equilibrium must be treated skeptically. By design they are unable to detect the dramatic conditions found here when data are analysed without this constraint.

econ.GN

Insurance makes wealth grow faster

Voluntary insurance contracts constitute a puzzle because they increase the expectation value of one party's wealth, whereas both parties must sign for such contracts to exist. Classically, the puzzle is resolved by introducing non-linear utility functions, which encode asymmetric risk preferences; or by assuming the parties have asymmetric information. Here we show the puzzle goes away if contracts are evaluated by their effect on the time-average growth rate of wealth. Our solution assumes only knowledge of wealth dynamics. Time averages and expectation values differ because wealth changes are non-ergodic. Our reasoning is generalisable: business happens when both parties grow faster.

q-fin.RM

Evaluating gambles using dynamics

Gambles are random variables that model possible changes in monetary wealth. Classic decision theory transforms money into utility through a utility function and defines the value of a gamble as the expectation value of utility changes. Utility functions aim to capture individual psychological characteristics, but their generality limits predictive power. Expectation value maximizers are defined as rational in economics, but expectation values are only meaningful in the presence of ensembles or in systems with ergodic properties, whereas decision-makers have no access to ensembles and the variables representing wealth in the usual growth models do not have the relevant ergodic properties. Simultaneously addressing the shortcomings of utility and those of expectations, we propose to evaluate gambles by averaging wealth growth over time. No utility function is needed, but a dynamic must be specified to compute time averages. Linear and logarithmic "utility functions" appear as transformations that generate ergodic observables for purely additive and purely multiplicative dynamics, respectively. We highlight inconsistencies throughout the development of decision theory, whose correction clarifies that our perspective is legitimate. These invalidate a commonly cited argument for bounded utility functions.

econ.GN

Ergodicity breaking in geometric Brownian motion

Geometric Brownian motion (GBM) is a model for systems as varied as financial instruments and populations. The statistical properties of GBM are complicated by non-ergodicity, which can lead to ensemble averages exhibiting exponential growth while any individual trajectory collapses according to its time-average. A common tactic for bringing time averages closer to ensemble averages is diversification. In this letter we study the effects of diversification using the concept of ergodicity breaking.

math-ph

Menger 1934 revisited

Karl Menger's 1934 paper on the St. Petersburg paradox contains mathematical errors that invalidate his conclusion that unbounded utility functions, specifically Bernoulli's logarithmic utility, fail to resolve modified versions of the St. Petersburg paradox.

q-fin.RM

The time resolution of the St. Petersburg paradox

A resolution of the St. Petersburg paradox is presented. In contrast to the standard resolution, utility is not required. Instead, the time-average performance of the lottery is computed. The final result can be phrased mathematically identically to Daniel Bernoulli's resolution, which uses logarithmic utility, but is derived using a conceptually different argument. The advantage of the time resolution is the elimination of arbitrary utility functions.

math.PR

Universality under conditions of self-tuning

We study systems with a continuous phase transition that tune their parameters to maximize a quantity that diverges solely at a unique critical point. Varying the size of these systems with dynamically adjusting parameters, the same finite-size scaling is observed as in systems where all relevant parameters are fixed at their critical values. This scheme is studied using a self-tuning variant of the Ising model. It is contrasted with a scheme where systems approach criticality through a target value for the order parameter that vanishes with increasing system size. In the former scheme, the universal exponents are observed in naive finite-size scaling studies, whereas in the latter they are not.

cond-mat.stat-mech

Optimal leverage from non-ergodicity

In modern portfolio theory, the balancing of expected returns on investments against uncertainties in those returns is aided by the use of utility functions. The Kelly criterion offers another approach, rooted in information theory, that always implies logarithmic utility. The two approaches seem incompatible, too loosely or too tightly constraining investors' risk preferences, from their respective perspectives. The conflict can be understood on the basis that the multiplicative models used in both approaches are non-ergodic which leads to ensemble-average returns differing from time-average returns in single realizations. The classic treatments, from the very beginning of probability theory, use ensemble-averages, whereas the Kelly-result is obtained by considering time-averages. Maximizing the time-average growth rates for an investment defines an optimal leverage, whereas growth rates derived from ensemble-average returns depend linearly on leverage. The latter measure can thus incentivize investors to maximize leverage, which is detrimental to time-average growth and overall market stability. The Sharpe ratio is insensitive to leverage. Its relation to optimal leverage is discussed. A better understanding of the significance of time-irreversibility and non-ergodicity and the resulting bounds on leverage may help policy makers in reshaping financial risk controls.

q-fin.RM

Tuning- and order parameter in the SOC ensemble

The one-dimensional Oslo model is studied under self-organized criticality (SOC) conditions and under absorbing state (AS) conditions. While the activity signals the phase transition under AS conditions by a sudden increase, this is not the case under SOC conditions. The scaling parameters of the activity are found to be identical under SOC and AS conditions, but in SOC the activity lacks a pickup.

cond-mat.stat-mech

Reply to "Comment on `Self-organized Criticality and Absorbing States: Lessons from the Ising Model'"

In [Braz. J. Phys. 30, 27 (2000)] Dickman et al. suggested that self-organized criticality can be produced by coupling the activity of an absorbing state model to a dissipation mechanism and adding an external drive. We analyzed the proposed mechanism in [Phys. Rev. E 73, 025106R (2006)] and found that if this mechanism is at work, the finite-size scaling found in self-organized criticality will depend on the details of the implementation of dissipation and driving. In the preceding comment [Phys. Rev. E XX, XXXX (2008)] Alava et al. show that one avalanche exponent in the AS approach becomes independent of dissipation and driving. In our reply we clarify their findings and put them in the context of the original article.

cond-mat.stat-mech