SearcharxivSearch

arXiv subjects

Matteo Marsili

Publications and source records attributed to Matteo Marsili.

At least 19 recordsLinked to original sources

Open-ended innovation arm-race in zero-sum games

This note discusses zero-sum games with open-ended innovation, whereby each player may introduce new strategies. The innovation process is modelled as a draw of new strategies form a distribution. It is argued that, when the cost of innovation is vanishingly small, this setting can lead to an everlasting innovation arm-race. In particular, the introduction of new technologies of the advanced player increases the marginal utility for technological innovation of the backward one, whereas innovation of the backward player disincentivizes the more advanced one to innovate.

physics.soc-ph

Exploring the conditions for sustainability with open-ended innovation

Can sustained open-ended technological progress preserve natural resources in a finite planet? We address this question on the basis of a stylized model with genuine open-ended technological innovation, where an innovation event corresponds to a random draw of a technology in the space of the parameters that define how it impacts the environment and how it interacts with the population. Technological innovation is endogenous because an innovation may invade if it satisfies constraints which depend on the state of the environment and of the population. We find that open-ended innovation leads either to a sustainable future where global population saturates and the environment is preserved, or to exploding population and a vanishing environment. What drives the transition between these two phases is not the level of environmental impact of technologies, but rather the demographic effects of technologies and labor productivity. Low demographic impact and high labor productivity (as in several western countries today) result in a Schumpeterian dynamics where new "greener" technologies displace older ones, thereby reducing the overall environmental impact. In this scenario, global population saturates to a finite value, imposing strong selective pressure on technological innovation. When technologies contribute significantly to demographic growth and/or labor productivity is low, technological innovation runs unrestrained, population grows unbounded, while the environment collapses. As such, our model captures subtle feedback effects between technological progress, demography and sustainability that rationalize and align with empirical observations of a demographic transition and the environmental Kuznets curve, without deriving it from profit maximization based on individual incentives.

physics.soc-ph

Lost in Retraining: Roaming the Parameter Space of Exponential Families Under Closed-Loop Learning

Closed-loop learning is the process of repeatedly estimating a model from data generated from the model itself. It is receiving great attention due to the possibility that large neural network models may, in the future, be primarily trained with data generated by artificial neural networks themselves. We study this process for models that belong to exponential families, deriving equations of motions that govern the dynamics of the parameters. We show that maximum likelihood estimation of the parameters endows sufficient statistics with the martingale property and that as a result the process converges to absorbing states that amplify initial biases present in the data. However, we show that this outcome may be prevented if the data contains at least one data point generated from a ground truth model, by relying on maximum a posteriori estimation or by introducing regularisation.

cs.LG

The pursuit of happiness

Happiness, in the U.S. Declaration of Independence, was understood quite differently from today's popular notions of personal pleasure. Happiness implies a flourishing life - one of virtue, purpose, and contribution to the common good. This paper studies populations of individuals - that we call homo-felix - who maximise an objective function that we call happiness. The happiness of one individual depends on the payoffs that they receive in games they play with their peers as well as on the happiness of the peers they interact with. Individuals care more or less about others depending on whether that makes them more or less happy. This paper analyses the happiness feedback loops that result from these interactions in simple settings. We find that individuals tend to care more about individuals who are happier than what they would be by being selfish. In simple 2 x 2 game theoretic settings, we show that homo-felix can converge to a variety of equilibria which includes but goes beyond Nash equilibria. In an n-persons public good game we show that the non-cooperative Nash equilibrium is marginally unstable and a single individual who develops prosocial behaviour is able to drive almost the whole population to a cooperative state.

econ.TH

Three state random energy model

We introduce a spin-1 version of the random energy model with crystal field. Crystal field controls the density of 0 spins in the system. We solve the model in the micro-canonincal ensemble. The model has a spin-glass transition at a finite temperature for all strengths of the crystal field. By introducing the magnetic field we also obtain the de Almeida Thouless line for the model. The spin-glass transition persists in the presence of external field. We also find that the magnetisation shows non-monotonic behaviour for high positive crystal field strengths. The zero magnetic field specific heat and magnetic susceptibility also exhibit a cusp beyond a threshold value of the crystal field.

cond-mat.dis-nn

Absolute abstraction: a renormalisation group approach

Abstraction is the process of extracting the essential features from raw data while ignoring irrelevant details. It is well known that abstraction emerges with depth in neural networks, where deep layers capture abstract characteristics of data by combining lower level features encoded in shallow layers (e.g. edges). Yet we argue that depth alone is not enough to develop truly abstract representations. We advocate that the level of abstraction crucially depends on how broad the training set is. We address the issue within a renormalisation group approach where a representation is expanded to encompass a broader set of data. We take the unique fixed point of this transformation -- the Hierarchical Feature Model -- as a candidate for a representation which is absolutely abstract. This theoretical picture is tested in numerical experiments based on Deep Belief Networks and auto-encoders trained on data of different breadth. These show that representations in neural networks approach the Hierarchical Feature Model as the data get broader and as depth increases, in agreement with theoretical predictions.

cs.LG

Is stochastic thermodynamics the key to understanding the energy costs of computation?

The relationship between the thermodynamic and computational characteristics of dynamical physical systems has been a major theoretical interest since at least the 19th century, and has been of increasing practical importance as the energetic cost of digital devices has exploded over the last half century. One of the most important thermodynamic features of real-world computers is that they operate very far from thermal equilibrium, in finite time, with many quickly (co-)evolving degrees of freedom. Such computers also must almost always obey multiple physical constraints on how they work. For example, all modern digital computers are periodic processes, governed by a global clock. Another example is that many computers are modular, hierarchical systems, with strong restrictions on the connectivity of their subsystems. This properties hold both for naturally occurring computers, like brains or Eukaryotic cells, as well as digital systems. These features of real-world computers are absent in 20th century analyses of the thermodynamics of computational processes, which focused on quasi-statically slow processes. However, the field of stochastic thermodynamics has been developed in the last few decades - and it provides the formal tools for analyzing systems that have exactly these features of real-world computers. We argue here that these tools, together with other tools currently being developed in stochastic thermodynamics, may help us understand at a far deeper level just how the fundamental physical properties of dynamic systems are related to the computation that they perform.

cond-mat.stat-mech

Application of spin glass ideas in social sciences, economics and finance

Classical economics has developed an arsenal of methods, based on the idea of representative agents, to come up with precise numbers for next year's GDP, inflation and exchange rates, among (many) other things. Few, however, will disagree with the fact that the economy is a complex system, with a large number of strongly heterogeneous, interacting units of different types (firms, banks, households, public institutions) and different sizes. Now, the main issue in economics is precisely the emergent organization, cooperation and coordination of such a motley crowd of micro-units. Treating them as a unique ``representative'' firm or household clearly risks throwing the baby with the bathwater. As we have learnt from statistical physics, understanding and characterizing such emergent properties can be difficult. Because of feedback loops of different signs, heterogeneities and non-linearities, the macro-properties are often hard to anticipate. In particular, these situations generically lead to a very large number of possible equilibria, or even the lack thereof. Spin-glasses and other disordered systems give a concrete example of such difficulties. In order to tackle these complex situations, new theoretical and numerical tools have been invented in the last 50 years, including of course the replica method and replica symmetry breaking, and the cavity method, both static and dynamic. In this chapter we review the application of such ideas and methods in economics and social sciences. Of particular interest are the proliferation (and fragility) of equilibria, the analogue of satisfiability phase transitions in games and random economies, and condensation (or concentration) effects in opinion, wealth, etc

physics.soc-ph

Multiscale Relevance of Natural Images

We use an agnostic information-theoretic approach to investigate the statistical properties of natural images. We introduce the Multiscale Relevance (MSR) measure to assess the robustness of images to compression at all scales. Starting in a controlled environment, we characterize the MSR of synthetic random textures as function of image roughness H and other relevant parameters. We then extend the analysis to natural images and find striking similarities with critical (H = 0) random textures. We show that the MSR is more robust and informative of image content than classical methods such as power spectrum analysis. Finally, we confront the MSR to classical measures for the calibration of common procedures such as color mapping and denoising. Overall, the MSR approach appears to be a good candidate for advanced image analysis and image processing, while providing a good level of physical interpretability.

physics.data-an

A Bayesian theory of market impact

The available liquidity at any time in financial markets falls largely short of the typical size of the orders that institutional investors would trade. In order to reduce the impact on prices due to the execution of large orders, traders in financial markets split large orders into a series of smaller ones, which are executed sequentially. The resulting sequence of trades is called a meta-order. Empirical studies have revealed a non-trivial set of statistical laws on how meta-orders affect prices, which include i) the square-root behaviour of the expected price variation with the total volume traded, ii) its crossover to a linear regime for small volumes, and iii) a reversion of average prices towards its initial value, after the sequence of trades is over. Here we recover this phenomenology within a minimal theoretical framework where the market sets prices by incorporating all information on the direction and speed of trade of the meta-order in a Bayesian manner. The simplicity of this derivation lends further support to the robustness and universality of market impact laws. In particular, it suggests that the square-root impact law originates from the over-estimation of order flows originating from meta-orders.

q-fin.TR

A simple probabilistic neural network for machine understanding

We discuss probabilistic neural networks with a fixed internal representation as models for machine understanding. Here understanding is intended as mapping data to an already existing representation which encodes an {\em a priori} organisation of the feature space. We derive the internal representation by requiring that it satisfies the principles of maximal relevance and of maximal ignorance about how different features are combined. We show that, when hidden units are binary variables, these two principles identify a unique model -- the Hierarchical Feature Model (HFM) -- which is fully solvable and provides a natural interpretation in terms of features. We argue that learning machines with this architecture enjoy a number of interesting properties, like the continuity of the representation with respect to changes in parameters and data, the possibility to control the level of compression and the ability to support functions that go beyond generalisation. We explore the behaviour of the model with extensive numerical experiments and argue that models where the internal representation is fixed reproduce a learning modality which is qualitatively different from that of traditional models such as Restricted Boltzmann Machines.

cond-mat.dis-nn

Quantifying Relevance in Learning and Inference

Learning is a distinctive feature of intelligent behaviour. High-throughput experimental data and Big Data promise to open new windows on complex systems such as cells, the brain or our societies. Yet, the puzzling success of Artificial Intelligence and Machine Learning shows that we still have a poor conceptual understanding of learning. These applications push statistical inference into uncharted territories where data is high-dimensional and scarce, and prior information on "true" models is scant if not totally absent. Here we review recent progress on understanding learning, based on the notion of "relevance". The relevance, as we define it here, quantifies the amount of information that a dataset or the internal representation of a learning machine contains on the generative model of the data. This allows us to define maximally informative samples, on one hand, and optimal learning machines on the other. These are ideal limits of samples and of machines, that contain the maximal amount of information about the unknown generative process, at a given resolution (or level of compression). Both ideal limits exhibit critical features in the statistical sense: Maximally informative samples are characterised by a power-law frequency distribution (statistical criticality) and optimal learning machines by an anomalously large susceptibility. The trade-off between resolution (i.e. compression) and relevance distinguishes the regime of noisy representations from that of lossy compression. These are separated by a special point characterised by Zipf's law statistics. This identifies samples obeying Zipf's law as the most compressed loss-less representations that are optimal in the sense of maximal relevance. Criticality in optimal learning machines manifests in an exponential degeneracy of energy levels, that leads to unusual thermodynamic properties.

cs.LG

A random energy approach to deep learning

We study a generic ensemble of deep belief networks which is parametrized by the distribution of energy levels of the hidden states of each layer. We show that, within a random energy approach, statistical dependence can propagate from the visible to deep layers only if each layer is tuned close to the critical point during learning. As a consequence, efficiently trained learning machines are characterised by a broad distribution of energy levels. The analysis of Deep Belief Networks and Restricted Boltzmann Machines on different datasets confirms these conclusions.

cond-mat.dis-nn

The rise and fall of hubs in Self-Organized Critical learning networks

Information processing networks are the result of local rewiring rules. In many instances, such rules promote links where the activity at the two end nodes is positively correlated. The conceptual problem we address is what network architecture prevails under such rules and how does the resulting network, in turn, constrain the dynamics. We focus on a simple toy model that captures the interplay between link self-reinforcement and a Self-Organised Critical dynamics in a simple way. Our main finding is that, under these conditions, a core of densely connected nodes forms spontaneously. Moreover, we show that the appearance of such clustered state can be dynamically regulated by a fatigue mechanism, eventually giving rise to non-trivial avalanche exponents.

cond-mat.stat-mech

Information thermodynamics of financial markets: the Glosten-Milgrom model

The Glosten-Milgrom model describes a single asset market, where informed traders interact with a market maker, in the presence of noise traders. We derive an analogy between this financial model and a Szil\'ard information engine by {\em i)} showing that the optimal work extraction protocol in the latter coincides with the pricing strategy of the market maker in the former and {\em ii)} defining a market analogue of the physical temperature from the analysis of the distribution of market orders. Then we show that the expected gain of informed traders is bounded above by the product of this market temperature with the amount of information that informed traders have, in exact analogy with the corresponding formula for the maximal expected amount of work that can be extracted from a cycle of the information engine. This suggests that recent ideas from information thermodynamics may shed light on financial markets, and lead to generalised inequalities, in the spirit of the extended second law of thermodynamics.

cond-mat.stat-mech

Bayesian Inference of Minimally Complex Models with Interactions of Arbitrary Order

Finding the model that best describes a high-dimensional dataset is a daunting task, even more so if one aims to consider all possible high-order patterns of the data, going beyond pairwise models. For binary data, we show that this task becomes feasible when restricting the search to a family of simple models, that we call Minimally Complex Models (MCMs). MCMs are maximum entropy models that have interactions of arbitrarily high order grouped into independent components of minimal complexity. They are simple in information-theoretic terms, which means they can only fit well certain types of data patterns and are therefore easy to falsify. We show that Bayesian model selection restricted to these models is computationally feasible and has many advantages. First, the model evidence, which balances goodness-of-fit against complexity, can be computed efficiently without any parameter fitting, enabling very fast explorations of the space of MCMs. Second, the family of MCMs is invariant under gauge transformations, which can be used to develop a representation-independent approach to statistical modeling. For small systems (up to 15 variables), combining these two results allows us to select the best MCM among all, even though the number of models is already extremely large. For larger systems, we propose simple heuristics to find optimal MCMs in reasonable times. Besides, inference and sampling can be performed without any computational effort. Finally, because MCMs have interactions of any order, they can reveal the presence of important high-order dependencies in the data, providing a new approach to explore high-order dependencies in complex systems. We apply our method to synthetic data and real-world examples, illustrating how MCMs portray the structure of dependencies among variables in a simple manner, extracting falsifiable predictions on symmetries and invariance from the data.

cs.AI

Characterising authors on the extent of their paper acceptance: A case study of the Journal of High Energy Physics

New researchers are usually very curious about the recipe that could accelerate the chances of their paper getting accepted in a reputed forum (journal/conference). In search of such a recipe, we investigate the profile and peer review text of authors whose papers almost always get accepted at a venue (Journal of High Energy Physics in our current work). We find authors with high acceptance rate are likely to have a high number of citations, high $h$-index, higher number of collaborators etc. We notice that they receive relatively lengthy and positive reviews for their papers. In addition, we also construct three networks -- co-reviewer, co-citation and collaboration network and study the network-centric features and intra- and inter-category edge interactions. We find that the authors with high acceptance rate are more `central' in these networks; the volume of intra- and inter-category interactions are also drastically different for the authors with high acceptance rate compared to the other authors. Finally, using the above set of features, we train standard machine learning models (random forest, XGBoost) and obtain very high class wise precision and recall. In a followup discussion we also narrate how apart from the author characteristics, the peer-review system might itself have a role in propelling the distinction among the different categories which could lead to potential discrimination and unfairness and calls for further investigation by the system admins.

cs.DL

Optimal Work Extraction and the Minimum Description Length Principle

We discuss work extraction from classical information engines (e.g., Szil\'ard) with $N$-particles, $q$ partitions, and initial arbitrary non-equilibrium states. In particular, we focus on their {\em optimal} behaviour, which includes the measurement of a set of quantities $\Phi$ with a feedback protocol that extracts the maximal average amount of work. We show that the optimal non-equilibrium state to which the engine should be driven before the measurement is given by the normalised maximum-likelihood probability distribution of a statistical model that admits $\Phi$ as sufficient statistics. Furthermore, we show that the minimax universal code redundancy $\mathcal{R}^*$ associated to this model, provides an upper bound to the work that the demon can extract on average from the cycle, in units of $k_{\rm B}T$. We also find that, in the limit of $N$ large, the maximum average extracted work cannot exceed $H[\Phi]/2$, i.e. one half times the Shannon entropy of the measurement. Our results establish a connection between optimal work extraction in stochastic thermodynamics and optimal universal data compression, providing design principles for optimal information engines. In particular, they suggest that: (i) optimal coding is thermodynamically efficient, and (ii) it is essential to drive the system into a critical state in order to achieve optimal performance.

cond-mat.stat-mech