SearcharxivSearch

arXiv subjects

Germinal Cocho

Publications and source records attributed to Germinal Cocho.

14 recordsLinked to original sources

Beta Rank Function: A Smooth Double-Pareto-Like Distribution

The Beta Rank Function (BRF) $x(u) =A(1-u)^b/u^a$, where $u$ is the normalized and continuous rank of an observation $x$, has wide applications in fitting real-world data from social science to biological phenomena. The underlying probability density function (pdf) $f_X(x)$ does not usually have a closed expression except for specific parameter values. We show however that it is approximately a unimodal skewed and asymmetric two-sided power law/double Pareto/log-Laplacian distribution. The BRF pdf has simple properties when the independent variable is log-transformed: $f_{Z=\log(X)}(z)$ . At the peak it makes a smooth turn and it does not diverge, lacking the sharp angle observed in the double Pareto or Laplace distribution. The peak position of $f_Z(z)$ is $z_0=\log A+(a-b)\log(\sqrt{a}+\sqrt{b})-(a\log(a)-b\log(b))/2 $; the probability is partitioned by the peak to the proportion of $\sqrt{b}/(\sqrt{a}+\sqrt{b})$ (left) and $\sqrt{a}/(\sqrt{a}+\sqrt{b})$ (right); the functional form near the peak is controlled by the cubic term in the Taylor expansion when $a\ne b$; the mean of $Z$ is $E[Z]=\log A+a-b$; the decay on left and right sides of the peak is approximately exponential with forms $e^{\frac{z-\log A}{b} }/b$ and $e^{ -\frac{z-\log A}{a}}/a$. These results are confirmed by numerical simulations. Properties of $f_X(x)$ without log-transforming the variable are much more complex, though the approximate double Pareto behavior, $(x/A)^{1/b}/(bx)$ (for $x A$) is simple. Our results elucidate the relationship between BRF and log-normal distributions when $a=b$ and explain why the BRF is ubiquitous and versatile. Based on the pdf, we suggest a quick way to elucidate if a real data set follows a one-sided power-law, a log-normal, a two-sided power-law or a BRF. We illustrate our results with two examples: urban populations and financial returns.

stat.ME

Rank-frequency distribution of natural languages: a difference of probabilities approach

The time variation of the rank $k$ of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for $k>50$ the law of large numbers predominates. The dynamics of $k$ is described stochastically through a master equation governing the time evolution of its probability density, which is approximated by a Fokker-Planck equation that is solved analytically. The difference between the data and the asymptotic solution is identified with the transient solution, and good agreement is obtained.

physics.soc-ph

Rank dynamics of word usage at multiple scales

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore whether word use is similar across languages, and if so, whether these generic features appear at different scales of language structure. Here we use the Google Books $N$-grams dataset to analyze the temporal evolution of word usage in several languages. We apply measures proposed recently to study rank dynamics, such as the diversity of $N$-grams in a given rank, the probability that an $N$-gram changes rank between successive time intervals, the rank entropy, and the rank complexity. Using different methods, results show that there are generic properties for different languages at different scales, such as a core of words necessary to minimally understand a language. We also propose a null model to explore the relevance of linguistic structure across multiple scales, concluding that $N$-gram statistics cannot be reduced to word statistics. We expect our results to be useful in improving text prediction algorithms, as well as in shedding light on the large-scale features of language use, beyond linguistic and cultural differences across human populations.

physics.soc-ph

Trajectory stability in the traveling salesman problem

Two generalizations of the traveling salesman problem in which sites change their position in time are presented. The way the rank of different trajectory lengths changes in time is studied using the rank diversity. We analyze the statistical properties of rank distributions and rank dynamics and give evidence that the shortest and longest trajectories are more predictable and robust to change, that is, more stable.

nlin.AO

Population patterns in World's administrative units

While there has been an extended discussion concerning city population distribution, little has been said about administrative units. Even though there might be a correspondence between cities and administrative divisions, they are conceptually different entities and the correspondence breaks as artificial divisions form and evolve. In this work we investigate the population distribution of second level administrative units for 150 countries and propose the Discrete Generalized Beta Distribution (DGBD) rank-size function to describe the data. After testing the goodness of fit of this two parameter function against power law, which is the most common model for city population, DGBD is a good statistical model for 73% of our data sets and better than power law in almost every case. Particularly, DGBD is better than power law for fitting country population data. The fitted parameters of this function allow us to construct a phenomenological characterization of countries according to the way in which people are distributed inside them. We present a computational model to simulate the formation of administrative divisions and give numerical evidence that DGBD arises from it. This model along with the DGBD function prove adequate to reproduce and describe local unit evolution and its effect on population distribution.

stat.AP

Universal temporal features of rankings in competitive sports and games

Many complex phenomena, from the selection of traits in biological systems to hierarchy formation in social and economic entities, show signs of competition and heterogeneous performance in the temporal evolution of their components, which may eventually lead to stratified structures such as the wealth distribution worldwide. However, it is still unclear whether the road to hierarchical complexity is determined by the particularities of each phenomena, or if there are universal mechanisms of stratification common to many systems. Human sports and games, with their (varied but simplified) rules of competition and measures of performance, serve as an ideal test bed to look for universal features of hierarchy formation. With this goal in mind, we analyse here the behaviour of players and team rankings over time for several sports and games. Even though, for a given time, the distribution of performance ranks varies across activities, we find statistical regularities in the dynamics of ranks. Specifically the rank diversity, a measure of the number of elements occupying a given rank over a length of time, has the same functional form in sports and games as in languages, another system where competition is determined by the use or disuse of grammatical structures. Our results support the notion that hierarchical phenomena may be driven by the same underlying mechanisms of rank formation, regardless of the nature of their components. Moreover, such regularities can in principle be used to predict lifetimes of rank occupancy, thus increasing our ability to forecast stratification in the presence of competition.

physics.soc-ph

Beyond Zipf's Law: The Lavalette Rank Function and its Properties

Although Zipf's law is widespread in natural and social data, one often encounters situations where one or both ends of the ranked data deviate from the power-law function. Previously we proposed the Beta rank function to improve the fitting of data which does not follow a perfect Zipf's law. Here we show that when the two parameters in the Beta rank function have the same value, the Lavalette rank function, the probability density function can be derived analytically. We also show both computationally and analytically that Lavalette distribution is approximately equal, though not identical, to the lognormal distribution. We illustrate the utility of Lavalette rank function in several datasets. We also address three analysis issues on the statistical testing of Lavalette fitting function, comparison between Zipf's law and lognormal distribution through Lavalette function, and comparison between lognormal distribution and Lavalette distribution.

physics.data-an

Rank diversity of languages: Generic behavior in computational linguistics

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity for books published in six European languages since 1800, and find that it follows a universal lognormal distribution. Based on the mean and standard deviation associated with the lognormal distribution, we define three different word regimes of languages: "heads" consist of words which almost do not change their rank in time, "bodies" are words of general use, while "tails" are comprised by context-specific words and vary their rank considerably in time. The heads and bodies reflect the size of language cores identified by linguists for basic communication. We propose a Gaussian random walk model which reproduces the rank variation of words in time and thus the diversity. Rank diversity of words can be understood as the result of random variations in rank, where the size of the variation depends on the rank itself. We find that the core size is similar for all languages studied.

cs.CL

Fractional exclusion statistics and the Random Matrix Boson Ensemble

The k-body Gaussian Embedded Ensemble of Random Matrices is considered for N bosons distributed on two single-particle levels. When k = N, the ensemble is equivalent to the Gaussian Orthogonal Ensemble (GOE), and when k = 2 it corresponds to the Two-body Random Ensemble (TBRE) for bosons. It is shown that the energy spectrum leads to a rank function which is of the form of a discrete generalized beta distribution. The same distribution is obtained assuming N non-interacting quasiparticles that obey the fractional exclusion statistics introduced by Haldane two decades ago.

physics.atom-ph

Tail universalities in rank distributions as an algebraic problem: the beta-like function

Although power laws of the Zipf type have been used by many workers to fit rank distributions in different fields like in economy, geophysics, genetics, soft-matter, networks etc., these fits usually fail at the tails. Some distributions have been proposed to solve the problem, but unfortunately they do not fit at the same time both ending tails. We show that many different data in rank laws, like in granular materials, codons, author impact in scientific journal, etc. are very well fitted by a beta-like function. Then we propose that such universality is due to the fact that a system made from many subsystems or choices, imply stretched exponential frequency-rank functions which qualitatively and quantitatively can be fitted with the proposed beta-like function distribution in the limit of many random variables. We prove this by transforming the problem into an algebraic one: finding the rank of successive products of a given set of numbers.

physics.data-an

Scale-free foraging by primates emerges from their interaction with a complex environment

Scale-free foraging patterns are widespread among animals. These may be the outcome of an optimal searching strategy to find scarce randomly distributed resources, but a less explored alternative is that this behaviour may result from the interaction of foraging animals with a particular distribution of resources. We introduce a simple foraging model where individuals follow mental maps and choose their displacements according to a maximum efficiency criterion, in a spatially disordered environment containing many trees with a heterogeneous size distribution. We show that a particular tree size frequency distribution induces non-Gaussian movement patterns with multiple spatial scales (Lévy walks). These results are consistent with tree size variation and Spider monkey (Ateles geoffroyi) foraging patterns. We discuss the consequences that our results may have for the patterns of seed dispersal by foraging primates.

q-bio.PE

Levy walk patterns in the foraging movements of spider monkeys (Ateles geoffroyi)

Scale invariant patterns have been found in different biological systems, in many cases resembling what physicists have found in other nonbiological systems. Here we describe the foraging patterns of free-ranging spider monkeys (Ateles geoffroyi) in the forest of the Yucatan Peninsula, Mexico and find that these patterns resemble what physicists know as Levy walks. First, the length of a trajectory s constituent steps, or continuous moves in the same direction, is best described by a power-law distribution in which the frequency of ever larger steps decreases as a negative power function of their length. The rate of this decrease is very close to that predicted by a previous analytical Levy walk model to be an optimal strategy to search for scarce resources distributed at random Viswanathan et al 1999). Second, the frequency distribution of the duration of stops or waiting times also approximates a power-law function. Finally, the mean square displacement during the monkeys first foraging trip increases more rapidly than would be expected from a random walk with constant step length, but within the range predicted for Levy walks. In view of these results, we analyze the different exponents characterizing the trajectories described by females and males, and by monkeys on their own or when part of a subgroup. We discuss the origin of these patterns and their implications for the foraging ecology of spider monkeys.

physics.bio-ph

Translocation Properties of Primitive Molecular Machines and Their Relevance to the Structure of the Genetic Code

We address the question, related with the origin of the genetic code, of why are there three bases per codon in the translation to protein process. As a followup to our previous work, we approach this problem by considering the translocation properties of primitive molecular machines, which capture basic features of ribosomal/messenger RNA interactions, while operating under prebiotic conditions. Our model consists of a short one-dimensional chain of charged particles(rRNA antecedent) interacting with a polymer (mRNA antecedent) via electrostatic forces. The chain is subject to external forcing that causes it to move along the polymer which is fixed in a quasi one dimensional geometry. Our numerical and analytic studies of statistical properties of random chain/polymer potentials suggest that, under very general conditions, a dynamics is attained in which the chain moves along the polymer in steps of three monomers. By adjusting the model in order to consider present day genetic sequences, we show that the above property is enhanced for coding regions. Intergenic sequences display a behavior closer to the random situation. We argue that this dynamical property could be one of the underlying causes for the three base codon structure of the genetic code.

physics.bio-ph