SearcharxivSearch

arXiv subjects

Oscar Fontanelli

Publications and source records attributed to Oscar Fontanelli.

12 recordsLinked to original sources

Non-Zipfian Distribution of Stopwords or Function Words and Subset Selection Models

Stopwords and function words are relatively less informative for the content of a language and more often play a structural role in a sentence. Stopwords are ubiquitous words and may contain verbs, adjectives and adverbs. On the other hand, function words are strictly prepositions, conjunctions, pronouns, determiners, qualifiers, articles, interrogatives, and a limited number of auxiliary verbs. In contrast to the well known Zipf's law for rank-frequency plot for all words, the rank-frequency plots for stopwords or function words are best fitted by the Beta Rank Function (BRF). On the other hand, the rank-frequency plots of non-stopwords or non-function-words also deviate from the Zipf's law, but are better described by a quadratic function of log-token-count over log-rank than by BRF. Based on the observed rank of stopwords or function words in the full word list, we propose a stopword/function word/subset selection model that the probability for being selected, as a function of the word's rank $r$, is a decreasing Hill's function ($1/(1+(r/r_{mid})^γ)$); whereas the probability for not being selected is the standard Hill's function ($1/(1+(r_{mid}/r)^γ)$). We validate this selection probability model by a direct estimation from an independent collection of texts. We also show analytically that this model leads to a BRF rank-frequency distribution for stopwords or function words when the original full word list follows the Zipf's law, as well as explaining the quadratic fitting function for the non-stopwords or non-function-words. A corollary of these results is that Zipf's law is not expected to be true for telegraphic speech in early childhood language learners or in agrammatism patients.

cs.CL

Quadratic Term Correction on Heaps' Law

Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a straight line in log-log scale. However, it has been observed that even in log-log scale, the type-token curve is still slightly concave, invalidating the power-law relation. At the next-order approximation, we have shown, by twenty English novels or writings (some are translated from another language to English), that quadratic functions in log-log scale fit the type-token data perfectly. Regression analyses of log(type)-log(token) data with both a linear and quadratic term consistently lead to a linear coefficient of slightly larger than 1, and a quadratic coefficient around -0.02. Using the ``random drawing colored ball from the bag with replacement" model, we have shown that the curvature of the log-log scale is identical to a ``pseudo-variance" which is negative. Although a pseudo-variance calculation may encounter numeric instability when the number of tokens is large, due to the large values of pseudo-weights, this formalism provides a rough estimation of the curvature when the number of tokens is small.

cs.CL

Modeling Two-Scale Rank Distributions via Redistribution Dynamics or an Analytic Derivation of the Beta Rank Function

Beta Rank Function (BRF) is a two-sided distribution characterized by a smooth peak and double powerlaw decay, widely used to model empirical data exhibiting deviations from pure power laws. In this paper, we introduce a novel two-step generative process that produces data exactly following the BRF distribution. The first step involves any mechanism generating a power-law distribution, while the second step applies a regressive redistribution process that reallocates resources from poorer to richer entities, thereby amplifying inequality. This approach represents the first analytic derivation of an exact BRF distribution from a generative mechanism. We validate the model through applications to income and urban population distributions. Beyond exact generation, this framework offers new insights into the systemic origins of deviations from power laws frequently observed in complex systems, linking rank distributions to underlying feedback and redistribution dynamics.

physics.soc-ph

The Contact and Mobility Networks of Mexico City

Mexico City, the largest city in Mexico, is also one of the largest cities in the world. It has over 9 million inhabitants and concentrates the vast majority of government and business centers. In this work we describe algorithms that use anonymized location data from mobile devices to construct Mexico City's contact and mobility networks aiming to help the analysis of the city's complexity by understanding movement and physical interaction patterns between its inhabitants. We show the effectiveness and usefulness of our approach by building networks with data collected in February 2020 and performing a general descriptive analysis on them. We found that contact networks in Mexico City are very sparse, characterized by a largest connected component, and with a heavy-tailed degree distribution. On the other hand, we observed that paths conformed by the highest-degrree nodes of mobility networks resemble Mexico City's street network; moreover, we found interesting qualitative differences in the degree distribution of these networks between weekends and weekdays. We present these results along with the release of contact and mobility networks.

physics.soc-ph

Human mobility patterns in Mexico City and their links with socioeconomic variables during the COVID-19 pandemic

The availability of cellphone geolocation data provides a remarkable opportunity to study human mobility patterns and how these patterns are affected by the recent pandemic. Two simple centrality metrics allow us to measure two different aspects of mobility in origin-destination networks constructed with this type of data: variety of places connected to a certain node (degree) and number of people that travel to or from a given node (strength). In this contribution, we present an analysis of node degree and strength in daily origin-destination networks for Greater Mexico City during 2020. Unlike what is observed in many complex networks, these origin-destination networks are not scale free. Instead, there is a characteristic scale defined by the distribution peak; centrality distributions exhibit a skewed two-tail distribution with power law decay on each side of the peak. We found that high mobility areas tend to be closer to the city center, have higher population and better socioeconomic conditions. Areas with anomalous behavior are almost always on the periphery of the city, where we can also observe qualitative difference in mobility patterns between east and west. Finally, we study the effect of mobility restrictions due to the outbreak of the COVID-19 pandemics on these mobility patterns.

cs.SI

Intermunicipal Travel Networks of Mexico (2020-2021)

We present a collection of networks that describe the travel patterns between municipalities in Mexico between 2020 and 2021. Using anonymized mobile device geo-location data we constructed directed, weighted networks representing the (normalized) volume of travels between municipalities. We analysed changes in global (graph total weight sum), local (centrality measures), and mesoscale (community structure) network features. We observe that changes in these features are associated with factors such as Covid-19 restrictions and population size. In general, events in early 2020 (when initial Covid-19 restrictions were implemented) induced more intense changes in network features, whereas later events had a less notable impact in network features. We believe these networks will be useful for researchers and decision makers in the areas of transportation, infrastructure planning, epidemic control and network science at large.

cs.SI

Analyzing time series activity of Twitter political spambots

The presence and complexity of political Twitter bots has increased in recent years, making it a very difficult task to recognize these accounts from real, human users. We intended to provide an answer to the following question: are temporal patterns of activity qualitatively different in fake and human accounts? We collected a large sample of tweets during the post-electoral conflict in the US in 2020 and performed supervised and non-supervised statistical learning technique sto quantify the predictive power of time-series features for human-bot recognition. Our results show that there are no substantial differences, suggesting that political bots are nowadays very capable of mimicking human behaviour. This finding reveals the need for novel, more sophisticated bot-detection techniques.

cs.SI

Modeling the Popularity of Twitter Hashtags with Master Equations

In this work we introduce a model based on master equations to describe the time evolution of the popularity of topics and hashtags on the Twitter social network. Specifically, we model the number of times a certain hashtag appears on the network as a function of time. In our model, the behavior of this quantity depends on the degree distribution of the network and the extrinsic interest the community has for the topic or hashtag. From the master equation, we are able to obtain explicit solutions for the mean and variance. We propose a gamma kernel function to model the topic popularity, which is quite simple and yields reasonable results. Finally, we validate the plausibility of the model by analyzing actual Twitter data obtained through the public API.

cs.SI

Distribuciones de probabilidad en las ciencias de la complejidad: una perspectiva contemporánea

Science in the 21st century seems to be governed by novel approaches involving interdisciplinary work, systemic perspectives and complexity theory concepts. These new paradigms force us to leave aside our elder mechanistic approaches and embrace new starting points based on stochasticity, chaoticity, statistics and probability. In this work we review the fundamental ideas of complexity theory and the classic probabilistic models to study complex systems, based on the law of large numbers, central limit theorems and stable distributions. We also talk about power laws as the most common model for phenomena showing long tail distributions and we explore the principal difficulties that arise in practice with this kind of models. We show a novel alternative for the descripition of this type of phenomena and lastly we show two examples that illustrate the applications of this new model.

physics.soc-ph

Beta Rank Function: A Smooth Double-Pareto-Like Distribution

The Beta Rank Function (BRF) $x(u) =A(1-u)^b/u^a$, where $u$ is the normalized and continuous rank of an observation $x$, has wide applications in fitting real-world data from social science to biological phenomena. The underlying probability density function (pdf) $f_X(x)$ does not usually have a closed expression except for specific parameter values. We show however that it is approximately a unimodal skewed and asymmetric two-sided power law/double Pareto/log-Laplacian distribution. The BRF pdf has simple properties when the independent variable is log-transformed: $f_{Z=\log(X)}(z)$ . At the peak it makes a smooth turn and it does not diverge, lacking the sharp angle observed in the double Pareto or Laplace distribution. The peak position of $f_Z(z)$ is $z_0=\log A+(a-b)\log(\sqrt{a}+\sqrt{b})-(a\log(a)-b\log(b))/2 $; the probability is partitioned by the peak to the proportion of $\sqrt{b}/(\sqrt{a}+\sqrt{b})$ (left) and $\sqrt{a}/(\sqrt{a}+\sqrt{b})$ (right); the functional form near the peak is controlled by the cubic term in the Taylor expansion when $a\ne b$; the mean of $Z$ is $E[Z]=\log A+a-b$; the decay on left and right sides of the peak is approximately exponential with forms $e^{\frac{z-\log A}{b} }/b$ and $e^{ -\frac{z-\log A}{a}}/a$. These results are confirmed by numerical simulations. Properties of $f_X(x)$ without log-transforming the variable are much more complex, though the approximate double Pareto behavior, $(x/A)^{1/b}/(bx)$ (for $x A$) is simple. Our results elucidate the relationship between BRF and log-normal distributions when $a=b$ and explain why the BRF is ubiquitous and versatile. Based on the pdf, we suggest a quick way to elucidate if a real data set follows a one-sided power-law, a log-normal, a two-sided power-law or a BRF. We illustrate our results with two examples: urban populations and financial returns.

stat.ME

Population patterns in World's administrative units

While there has been an extended discussion concerning city population distribution, little has been said about administrative units. Even though there might be a correspondence between cities and administrative divisions, they are conceptually different entities and the correspondence breaks as artificial divisions form and evolve. In this work we investigate the population distribution of second level administrative units for 150 countries and propose the Discrete Generalized Beta Distribution (DGBD) rank-size function to describe the data. After testing the goodness of fit of this two parameter function against power law, which is the most common model for city population, DGBD is a good statistical model for 73% of our data sets and better than power law in almost every case. Particularly, DGBD is better than power law for fitting country population data. The fitted parameters of this function allow us to construct a phenomenological characterization of countries according to the way in which people are distributed inside them. We present a computational model to simulate the formation of administrative divisions and give numerical evidence that DGBD arises from it. This model along with the DGBD function prove adequate to reproduce and describe local unit evolution and its effect on population distribution.

stat.AP

Beyond Zipf's Law: The Lavalette Rank Function and its Properties

Although Zipf's law is widespread in natural and social data, one often encounters situations where one or both ends of the ranked data deviate from the power-law function. Previously we proposed the Beta rank function to improve the fitting of data which does not follow a perfect Zipf's law. Here we show that when the two parameters in the Beta rank function have the same value, the Lavalette rank function, the probability density function can be derived analytically. We also show both computationally and analytically that Lavalette distribution is approximately equal, though not identical, to the lognormal distribution. We illustrate the utility of Lavalette rank function in several datasets. We also address three analysis issues on the statistical testing of Lavalette fitting function, comparison between Zipf's law and lognormal distribution through Lavalette function, and comparison between lognormal distribution and Lavalette distribution.

physics.data-an