Searcharxiv⌕ Search

arXiv subjects

Matteo Marsili

Publications and source records attributed to Matteo Marsili.

At least 55 records · Page 3Linked to original sources

Anomalies in the peer-review system: A case study of the journal of High Energy Physics

Peer-review system has long been relied upon for bringing quality research to the notice of the scientific community and also preventing flawed research from entering into the literature. The need for the peer-review system has often been debated as in numerous cases it has failed in its task and in most of these cases editors and the reviewers were thought to be responsible for not being able to correctly judge the quality of the work. This raises a question "Can the peer-review system be improved?" Since editors and reviewers are the most important pillars of a reviewing system, we in this work, attempt to address a related question - given the editing/reviewing history of the editors or re- viewers "can we identify the under-performing ones?", with citations received by the edited/reviewed papers being used as proxy for quantifying performance. We term such review- ers and editors as anomalous and we believe identifying and removing them shall improve the performance of the peer- review system. Using a massive dataset of Journal of High Energy Physics (JHEP) consisting of 29k papers submitted between 1997 and 2015 with 95 editors and 4035 reviewers and their review history, we identify several factors which point to anomalous behavior of referees and editors. In fact the anomalous editors and reviewers account for 26.8% and 14.5% of the total editors and reviewers respectively and for most of these anomalous reviewers the performance degrades alarmingly over time.

cs.DL↗

When does inequality freeze an economy?

Inequality and its consequences are the subject of intense recent debate. Using a simplified model of the economy, we address the relation between inequality and liquidity, the latter understood as the frequency of economic exchanges. Assuming a Pareto distribution of wealth for the agents, that is consistent with empirical findings, we find an inverse relation between wealth inequality and overall liquidity. We show that an increase in the inequality of wealth results in an even sharper concentration of the liquid financial resources. This leads to a congestion of the flow of goods and the arrest of the economy when the Pareto exponent reaches one.

econ.GN↗

Identifying relevant positions in proteins by Critical Variable Selection

Evolution in its course found a variety of solutions to the same optimisation problem. The advent of high-throughput genomic sequencing has made available extensive data from which, in principle, one can infer the underlying structure on which biological functions rely. In this paper, we present a new method aimed at extracting sites encoding structural and func- tional properties from a set of protein primary sequences, namely a Multiple Sequence Alignment. The method, called Critical Variable Selection, is based on the idea that subsets of relevant sites cor- respond to subsequences that occur with a particularly broad frequency distribution in the dataset. By applying this algorithm to in silico sequences, to the Response Regulator Receiver and to the Voltage Sensor Domain of Ion Channels, we show that this procedure recovers not only information encoded in single site statistics and pairwise correlations but it also captures dependencies going beyond pairwise correlations. The method proposed here is complementary to Statistical Coupling Analysis, in that the most relevant sites predicted by the two methods markedly differ. We find robust and consistent results for datasets as small as few hundred sequences, that reveal a hidden hierarchy of sites that is consistent with present knowledge on biologically relevant sites and evo- lutionary dynamics. This suggests that Critical Variable Selection is able to identify in a Multiple Sequence Alignment a core of sites encoding functional and structural information.

q-bio.QM↗

Criticality of mostly informative samples: A Bayesian model selection approach

We discuss a Bayesian model selection approach to high dimensional data in the deep under sampling regime. The data is based on a representation of the possible discrete states $s$, as defined by the observer, and it consists of $M$ observations of the state. This approach shows that, for a given sample size $M$, not all states observed in the sample can be distinguished. Rather, only a partition of the sampled states $s$ can be resolved. Such partition defines an {\em emergent} classification $q_s$ of the states that becomes finer and finer as the sample size increases, through a process of {\em symmetry breaking} between states. This allows us to distinguish between the $resolution$ of a given representation of the observer defined states $s$, which is given by the entropy of $s$, and its $relevance$ which is defined by the entropy of the partition $q_s$. Relevance has a non-monotonic dependence on resolution, for a given sample size. In addition, we characterise most relevant samples and we show that they exhibit power law frequency distributions, generally taken as signatures of "criticality". This suggests that "criticality" reflects the relevance of a given representation of the states of a complex system, and does not necessarily require a specific mechanism of self-organisation to a critical point.

physics.data-an↗

Phenotypic constraints promote latent versatility and carbon efficiency in metabolic networks

System-level properties of metabolic networks may be the direct product of natural selection or arise as a by-product of selection on other properties. Here we study the effect of direct selective pressure for growth or viability in particular environments on two properties of metabolic networks: latent versatility to function in additional environments and carbon usage efficiency. Using a Markov Chain Monte Carlo (MCMC) sampling based on Flux Balance Analysis (FBA), we sample from a known biochemical universe random viable metabolic networks that differ in the number of directly constrained environments. We find that the latent versatility of sampled metabolic networks increases with the number of directly constrained environments and with the size of the networks. We then show that the average carbon wastage of sampled metabolic networks across the constrained environments decreases with the number of directly constrained environments and with the size of the networks. Our work expands the growing body of evidence about nonadaptive origins of key functional properties of biological networks.

q-bio.MN↗

Trade-offs in delayed information transmission in biochemical networks

In order to transmit biochemical signals, biological regulatory systems dissipate energy with concomitant entropy production. Additionally, signaling often takes place in challenging environmental conditions. In a simple model regulatory circuit given by an input and a delayed output, we explore the trade-offs between information transmission and the system's energetic efficiency. We determine the maximally informative network, given a fixed amount of entropy production and delayed response, exploring both the case with and without feedback. We find that feedback allows the circuit to overcome energy constraints and transmit close to the maximum available information even in the dissipationless limit. Negative feedback loops, characteristic of shock responses, are optimal at high dissipation. Close to equilibrium positive feedback loops, known for their stability, become more informative. Asking how the signaling network should be constructed to best function in the worst possible environment, rather than an optimally tuned one or in steady state, we discover that at large dissipation the same universal motif is optimal in all of these conditions.

q-bio.MN↗

Contour map of estimation error for Expected Shortfall

The contour map of estimation error of Expected Shortfall (ES) is constructed. It allows one to quantitatively determine the sample size (the length of the time series) required by the optimization under ES of large institutional portfolios for a given size of the portfolio, at a given confidence level and a given estimation error.

q-fin.RM↗

Condensation phenomena in fat-tailed distributions: a characterization by means of an order parameter

Condensation phenomena are ubiquitous in nature and are found in condensed matter, disordered systems, networks, finance, etc. In the present work we investigate one of the best frameworks in which condensation phenomena take place, namely, the sum of independent and fat-tailed distributed random variables. For large deviations of the sum, this system undergoes a phase transition and shifts from a democratic phase to a condensed phase, where a single variable (the condensate) carries a finite fraction of the sum. This phenomenon yields the failure of the standard results of the Large Deviation Theory. In this work we exploit the Density Functional Method to overcome the limitation of the Large Deviation Theory and characterize the condensation transition in terms of an order parameter, i.e. the Inverse Participation Ratio (IPR). This procedure leads us to investigate the system in the large-deviation regime where both the sum and the IPR are constrained, observing new phase transitions. As a sample application, the case of condensation phenomena in financial time-series is briefly discussed.

cond-mat.stat-mech↗

Statistical Mechanics of Competitive Resource Allocation using Agent-based Models

Demand outstrips available resources in most situations, which gives rise to competition, interaction and learning. In this article, we review a broad spectrum of multi-agent models of competition (El Farol Bar problem, Minority Game, Kolkata Paise Restaurant problem, Stable marriage problem, Parking space problem and others) and the methods used to understand them analytically. We emphasize the power of concepts and tools from statistical mechanics to understand and explain fully collective phenomena such as phase transitions and long memory, and the mapping between agent heterogeneity and physical disorder. As these methods can be applied to any large-scale model of competitive resource allocation made up of heterogeneous adaptive agent with non-linear interaction, they provide a prospective unifying paradigm for many scientific disciplines.

physics.soc-ph↗

On the concentration of large deviations for fat tailed distributions, with application to financial data

Large deviations for fat tailed distributions, i.e. those that decay slower than exponential, are not only relatively likely, but they also occur in a rather peculiar way where a finite fraction of the whole sample deviation is concentrated on a single variable. The regime of large deviations is separated from the regime of typical fluctuations by a phase transition where the symmetry between the points in the sample is spontaneously broken. For stochastic processes with a fat tailed microscopic noise, this implies that while typical realizations are well described by a diffusion process with continuous sample paths, large deviation paths are typically discontinuous. For eigenvalues of random matrices with fat tailed distributed elements, a large deviation where the trace of the matrix is anomalously large concentrates on just a single eigenvalue, whereas in the thin tailed world the large deviation affects the whole distribution. These results find a natural application to finance. Since the price dynamics of financial stocks is characterized by fat tailed increments, large fluctuations of stock prices are expected to be realized by discrete jumps. Interestingly, we find that large excursions of prices are more likely realized by continuous drifts rather than by discontinuous jumps. Indeed, auto-correlations suppress the concentration of large deviations. Financial covariance matrices also exhibit an anomalously large eigenvalue, the market mode, as compared to the prediction of random matrix theory. We show that this is explained by a large deviation with excess covariance rather than by one with excess volatility.

cond-mat.stat-mech↗

$L_p$ regularized portfolio optimization

Investors who optimize their portfolios under any of the coherent risk measures are naturally led to regularized portfolio optimization when they take into account the impact their trades make on the market. We show here that the impact function determines which regularizer is used. We also show that any regularizer based on the norm $L_p$ with $p>1$ makes the sensitivity of coherent risk measures to estimation error disappear, while regularizers with $p<1$ do not. The $L_1$ norm represents a border case: its "soft" implementation does not remove the instability, but rather shifts its locus, whereas its "hard" implementation (equivalent to a ban on short selling) eliminates it. We demonstrate these effects on the important special case of Expected Shortfall (ES) that is on its way to becoming the next global regulatory market risk measure.

q-fin.PM↗

The Interrupted Power Law and The Size of Shadow Banking

Using public data (Forbes Global 2000) we show that the asset sizes for the largest global firms follow a Pareto distribution in an intermediate range, that is ``interrupted'' by a sharp cut-off in its upper tail, where it is totally dominated by financial firms. This flattening of the distribution contrasts with a large body of empirical literature which finds a Pareto distribution for firm sizes both across countries and over time. Pareto distributions are generally traced back to a mechanism of proportional random growth, based on a regime of constant returns to scale. This makes our findings of an ``interrupted'' Pareto distribution all the more puzzling, because we provide evidence that financial firms in our sample should operate in such a regime. We claim that the missing mass from the upper tail of the asset size distribution is a consequence of shadow banking activity and that it provides an (upper) estimate of the size of the shadow banking system. This estimate -- which we propose as a shadow banking index -- compares well with estimates of the Financial Stability Board until 2009, but it shows a sharper rise in shadow banking activity after 2010. Finally, we propose a proportional random growth model that reproduces the observed distribution, thereby providing a quantitative estimate of the intensity of shadow banking activity.

q-fin.GN↗

Fishing out collective memory of migratory schools

Animals form groups for many reasons but there are costs and benefit associated with group formation. One of the benefits is collective memory. In groups on the move, social interactions play a crucial role in the cohesion and the ability to make consensus decisions. When migrating from spawning to feeding areas fish schools need to retain a collective memory of the destination site over thousand of kilometers and changes in group formation or individual preference can produce sudden changes in migration pathways. We propose a modelling framework, based on stochastic adaptive networks, that can reproduce this collective behaviour. We assume that three factors control group formation and school migration behaviour: the intensity of social interaction, the relative number of informed individuals and the preference that each individual has for the particular migration area. We treat these factors independently and relate the individuals' preferences to the experience and memory for certain migration sites. We demonstrate that removal of knowledgable individuals or alteration of individual preference can produce rapid changes in group formation and collective behavior. For example, intensive fishing targeting the migratory species and also their preferred prey can reduce both terms to a point at which migration to the destination sites is suddenly stopped. The conceptual approaches represented by our modelling framework may therefore be able to explain large-scale changes in fish migration and spatial distribution.

q-bio.QM↗

On sampling and modeling complex systems

The study of complex systems is limited by the fact that only few variables are accessible for modeling and sampling, which are not necessarily the most relevant ones to explain the systems behavior. In addition, empirical data typically under sample the space of possible states. We study a generic framework where a complex system is seen as a system of many interacting degrees of freedom, which are known only in part, that optimize a given function. We show that the underlying distribution with respect to the known variables has the Boltzmann form, with a temperature that depends on the number of unknown variables. In particular, when the unknown part of the objective function decays faster than exponential, the temperature decreases as the number of variables increases. We show in the representative case of the Gaussian distribution, that models are predictable only when the number of relevant variables is less than a critical threshold. As a further consequence, we show that the information that a sample contains on the behavior of the system is quantified by the entropy of the frequency with which different states occur. This allows us to characterize the properties of maximally informative samples: in the under-sampling regime, the most informative frequency size distributions have power law behavior and Zipf's law emerges at the crossover between the under sampled regime and the regime where the sample contains enough statistics to make inference on the behavior of the system. These ideas are illustrated in some applications, showing that they can be used to identify relevant variables or to select most informative representations of data, e.g. in data clustering.

physics.data-an↗

Time-dependent information transmission in a model regulatory circuit

Many biological regulatory systems process signals out of steady state and respond with a physiological delay. A simple model of regulation which respects these features shows how the ability of a delayed output to transmit information is limited: at short times by the timescale of the dynamic input, at long times by that of the dynamic output. We find that topologies of maximally informative networks correspond to commonly occurring biological circuits linked to stress response and that circuits functioning out of steady state may exploit absorbing states to transmit information optimally.

q-bio.MN↗

What do leaders know?

The ability of a society to make the right decisions on relevant matters relies on its capability to properly aggregate the noisy information spread across the individuals it is made of. In this paper we study the information aggregation performance of a stylized model of a society whose most influential individuals - the leaders - are highly connected among themselves and uninformed. Agents update their state of knowledge in a Bayesian manner by listening to their neighbors. We find analytical and numerical evidence of a transition, as a function of the noise level in the information initially available to agents, from a regime where information is correctly aggregated to one where the population reaches consensus on the wrong outcome with finite probability. Furthermore, information aggregation depends in a non-trivial manner on the relative size of the clique of leaders, with the limit of a vanishingly small clique being singular.

physics.soc-ph↗

The Effect of Nonstationarity on Models Inferred from Neural Data

Neurons subject to a common non-stationary input may exhibit a correlated firing behavior. Correlations in the statistics of neural spike trains also arise as the effect of interaction between neurons. Here we show that these two situations can be distinguished, with machine learning techniques, provided the data are rich enough. In order to do this, we study the problem of inferring a kinetic Ising model, stationary or nonstationary, from the available data. We apply the inference procedure to two data sets: one from salamander retinal ganglion cells and the other from a realistic computational cortical network model. We show that many aspects of the concerted activity of the salamander retinal neurons can be traced simply to the external input. A model of non-interacting neurons subject to a non-stationary external field outperforms a model with stationary input with couplings between neurons, even accounting for the differences in the number of model parameters. When couplings are added to the non-stationary model, for the retinal data, little is gained: the inferred couplings are generally not significant. Likewise, the distribution of the sizes of sets of neurons that spike simultaneously and the frequency of spike patterns as function of their rank (Zipf plots) are well-explained by an independent-neuron model with time-dependent external input, and adding connections to such a model does not offer significant improvement. For the cortical model data, robust couplings, well correlated with the real connections, can be inferred using the non-stationary model. Adding connections to this model slightly improves the agreement with the data for the probability of synchronous spikes but hardly affects the Zipf plot.

q-bio.QM↗

The Social Climbing Game

The structure of a society depends, to some extent, on the incentives of the individuals they are composed of. We study a stylized model of this interplay, that suggests that the more individuals aim at climbing the social hierarchy, the more society's hierarchy gets strong. Such a dependence is sharp, in the sense that a persistent hierarchical order emerges abruptly when the preference for social status gets larger than a threshold. This phase transition has its origin in the fact that the presence of a well defined hierarchy allows agents to climb it, thus reinforcing it, whereas in a "disordered" society it is harder for agents to find out whom they should connect to in order to become more central. Interestingly, a social order emerges when agents strive harder to climb society and it results in a state of reduced social mobility, as a consequence of ergodicity breaking, where climbing is more difficult.

physics.soc-ph↗