SearcharxivSearch

arXiv subjects

Pablo Rosillo-Rodes

Publications and source records attributed to Pablo Rosillo-Rodes.

7 recordsLinked to original sources

Dynastic dynamics: Modelling powerful naming choices with stochastic prestige

The naming choices of powerful rulers encode millennia of cultural, political, and institutional influences. Consequently, the long-term onomastic dynamics within some of the most powerful dynasties cannot be explained by simple mechanisms such as frequency-dependent reinforcement or random name reuse. Here, we propose a minimal model that encapsulates these influences in a single highly stochastic variable, prestige, and assume that name selection is driven by the prestige accumulated by previous rulers bearing the same name within each dynasty. Using an extensive dataset spanning ten dynasties, we show that the model reproduces the pronounced inequalities in name frequencies, the long-term persistence of dominant names, and the abrupt rises in popularity observed in historical records. Our results suggest that the naming traditions of powerful rulers preserve institutional memory through reinforcement while remaining sensitive to rare historical events that can reshape the hierarchy of names across generations.

physics.soc-ph

Simon's model does not produce Zipf's law: The fundamental rich-get-richer mechanism for any power-law size ranking

Many complex systems are composed of disparate, interacting types of varying sizes: Species abundances in ecosystems, firm sizes in markets, city populations in countries, word counts in language, etc. A longstanding mystery of complex systems is Zipf's law, which is the empirical observation that component size decreases as the inverse of component rank -- $S \propto r^{-1}$ -- and its generalization $S \propto r^{-\alpha}$ for $\alpha \ge 0$. Herbert Simon's 1955 theoretical rich-get-richer mechanism for system growth has prevailed as capturing the essential process. But Simon's analysis is in fact flawed: In the limit of zero innovation, the model leads to a winner-takes-all system with $\alpha \rightarrow \infty$, rather than $\alpha \rightarrow 1$. Here, for pure rich-get-richer systems, we derive the time-dependent innovation rate $\rho_t$ that correctly produces power-law size rankings across all $\alpha \ge 0$. To produce Zipf's law, we uncover that $\rho_t$ must decay as the inverse of the log of the number of types, $1/\ln N$. We then show that our time-dependent innovation rate governs type emergence in any system obeying a power-law size-ranking, independent of the underlying mechanism. We demonstrate agreement between our model's output and word rankings in a collection of famous novels, while Simon's model fails. Going forward, our dynamic innovation rate mechanism provides the fundamental, Drosophila-like model for all rich-get-richer systems.

physics.soc-ph

Statistical laws and linguistics differ in naturalistic video and fictional conversations

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is through statistical patterns such as Heaps' law, which holds that vocabulary size scales with document length. Little work on Heaps' law has looked at conversation and considered how language features impact scaling. We measure Heaps' law for conversations recorded in two distinct mediums: 1. Strangers brought together on video chat and 2. Fictional characters in movies. We find that scaling of vocabulary size differs by parts of speech, suggesting a less efficient purpose in communication by medium.

cs.CL

Complete asymptotic type-token relationship for growing complex systems with inverse power-law count rankings

The growth dynamics of complex systems often exhibit statistical regularities involving power-law relationships. For real finite complex systems formed by countable tokens (animals, words) as instances of distinct types (species, dictionary entries), an inverse power-law scaling $S \sim r^{-\alpha}$ between type count $S$ and type rank $r$, widely known as Zipf's law, is widely observed to varying degrees of fidelity. A secondary, summary relationship is Heaps' law, which states that the number of types scales sublinearly with the total number of observed tokens present in a growing system. Here, we propose an idealized model of a growing system that (1) deterministically produces arbitrary inverse power-law count rankings for types, and (2) allows us to determine the exact asymptotics of the type-token relationship. Our argument improves upon and remedies earlier work. We obtain a unified asymptotic expression for all values of $\alpha$, which corrects the special cases of $\alpha = 1$ and $\alpha \gg 1$. Our approach relies solely on the form of count rankings, avoids unnecessary approximations, and does not involve any stochastic mechanisms or sampling processes. We thereby demonstrate that a general type-token relationship arises solely as a consequence of Zipf's law.

physics.soc-ph

Entropy and type-token ratio in gigaword corpora

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six massive linguistic datasets in English, Spanish, and Turkish, consisting of books, news articles, and tweets. These gigaword corpora correspond to languages with distinct morphological features and differ in registers and genres, thus constituting a varied testbed for a quantitative approach to lexical diversity. We unveil an empirical functional relation between entropy and type-token ratio of texts of a given corpus and language, which is a consequence of the statistical laws observed in natural language. Further, in the limit of large text lengths we find an analytical expression for this relation relying on both Zipf and Heaps laws that agrees with our empirical findings.

cs.CL

Computational lexical analysis of Flamenco genres

Flamenco, recognized by UNESCO as part of the Intangible Cultural Heritage of Humanity, is a profound expression of cultural identity rooted in Andalusia, Spain. However, there is a lack of quantitative studies that help identify characteristic patterns in this long-lived music tradition. In this work, we present a computational analysis of Flamenco lyrics, employing natural language processing and machine learning to categorize over 2000 lyrics into their respective Flamenco genres, termed as $\textit{palos}$. Using a Multinomial Naive Bayes classifier, we find that lexical variation across styles enables to accurately identify distinct $\textit{palos}$. More importantly, from an automatic method of word usage, we obtain the semantic fields that characterize each style. Further, applying a metric that quantifies the inter-genre distance we perform a network analysis that sheds light on the relationship between Flamenco styles. Remarkably, our results suggest historical connections and $\textit{palo}$ evolutions. Overall, our work illuminates the intricate relationships and cultural significance embedded within Flamenco lyrics, complementing previous qualitative discussions with quantitative analyses and sparking new discussions on the origin and development of traditional music genres.

cs.CL

Modeling language ideologies for the dynamics of languages in contact

In multilingual societies, it is common to encounter different language varieties. Various approaches have been proposed to discuss different mechanisms of language shift. However, current models exploring language shift in languages in contact often overlook the influence of language ideologies. Language ideologies play a crucial role in understanding language usage within a cultural community, encompassing shared beliefs, assumptions, and feelings towards specific language forms. These ideologies shed light on the social perceptions of different language varieties expressed as language attitudes. In this study, we introduce an approach that incorporates language ideologies into a model for contact varieties by considering speaker preferences as a parameter. Our findings highlight the significance of preference in language shift, which can even outweigh the influence of language prestige associated, for example, with a standard variety. Furthermore, we investigate the impact of the degree of interaction between individuals holding opposing preferences on the language shift process. Quite expectedly, our results indicate that when communities with different preferences mix, the coexistence of language varieties becomes less likely. However, variations in the degree of interaction between individuals with contrary preferences notably lead to non-trivial transitions from states of coexistence of varieties to the extinction of a given variety, followed by a return to coexistence, ultimately culminating in the dominance of the previously extinct variety. By studying finite-size effects, we observe that the duration of coexistence states increases exponentially with network size. Ultimately, our work constitutes a quantitative approach to the study of language ideologies in sociolinguistics.

physics.soc-ph