Searcharxiv⌕ Search

arXiv subjects

Eduardo G. Altmann

Publications and source records attributed to Eduardo G. Altmann.

At least 19 recordsLinked to original sources

Fractals in rate-induced tipping

When parameters of a dynamical system change sufficiently fast, critical transitions can take place even in the absence of bifurcations. This phenomenon is known as rate-induced tipping and has been reported in a variety of systems, from simple ordinary differential equations and maps to mathematical models in climate sciences and ecology. In most examples, the transition happens at a critical rate of parameter change, a rate-induced tipping point, and is associated with a simple unstable orbit (edge state). In this work, we show how this simple picture changes when non-attracting fractal sets exist in the autonomous system, a ubiquitous situation in non-linear dynamics. We show that these fractals in phase space induce fractals in parameter space, which control the rates and parameter changes that result in tipping. We explain how such rate-induced fractals appear and how the fractal dimensions of the different sets are related to each other. We illustrate our general theory in three paradigmatic systems: a piecewise linear one-dimensional map, the two-dimensional Hénon map, and a forced pendulum.

nlin.CD↗

Code evolution for link prediction in complex networks

The problem of predicting links in complex networks appears in different disciplines and has led to a variety of ingenious human-designed methods. We use this rich program space to explore the performance and behavior of automated code-evolution systems tasked to obtain machine-designed methods for link prediction. Despite being trained on limited data, algorithms evolved through code evolution outperform human-designed methods (with an average AUC score of 0.915 vs. 0.783, computed over 580 networks) and show improved computational efficiency, allowing them to be applied to networks with millions of links. The discovered methods follow approaches that have been employed in human-designed methods, but contain key innovations in the selection and combination of node- and link-features. This illustrates the role modern large language models and genetic algorithms can play in algorithmic innovation and scientific discovery more generally.

cs.SI↗

Meso-scale structures in signed networks

Meso-scale structures in signed networks have been studied under the limiting assumption of the validity of social balance theory, which predicts positive connections within groups and negative connections between groups. Here, we propose and apply a methodology that overcomes this limitation and is able to find and characterize also the different possible unbalanced structures in signed networks. Applying our methodology to 24 empirical networks, from social-political, financial, and biological domains, we find that unbalanced meso-scale structures are prevalent in real-world networks, including cases with substantial balance at the micro-scale of triangles. In particular, we find that assortativity often prevails regardless of the interaction sign and that core-periphery structures are typical in online social networks. Our findings highlight the complexity of meso-scale relational structures, the importance of using computational methods that are a priori agnostic to specific patterns, and the importance of independently evaluating micro- and meso-scale predictions of social balance theory.

cs.SI↗

Inference of epidemic networks: the effect of different data types

We investigate how the properties of epidemic networks change depending on the availability of different types of data on a disease outbreak. This is achieved by introducing mathematical and computational methods that estimate the probability of transmission trees by combining generative models that jointly determine the number of infected hosts, the probability of infection between them depending on location and genetic information, and their time of infection and sampling. We introduce a suitable Markov Chain Monte Carlo method that we show to sample trees according to their probability. Statistics performed over the sampled trees lead to probabilistic estimations of network properties and other quantities of interest, such as the number of unobserved hosts and the depth of the infection tree. We confirm the validity of our approach by comparing the numerical results with analytically solvable examples. Finally, we apply our methodology to data from COVID-19 in Australia. We find that network properties that are important for the management of the outbreak depend sensitively on the type of data used in the inference.

physics.comp-ph↗

Assessing Honey Bee Colony Health Using Temperature Time Series

Honey bees face an increasing number of stressors that disrupt the natural behaviour of colonies and, in extreme cases, can lead to their collapse. Quantifying the status and resilience of colonies is essential to measure the impact of stressors and to identify colonies at risk. In this manuscript, we present and apply new methodologies to efficiently diagnose the status of a honey bee colony from widely available time series of hive and environmental temperature. Healthy hives have a remarkable ability to control temperature near the brood area. Our method exploits this fact and quantifies the status of a hive by measuring how resilient they are to extreme environmental temperatures, which act as natural stressors. Analysing 22 hives during different times of the year, including 3 hives that collapsed, we find the statistical signatures of stress that reveal whether honeybees are doing well or are at risk of failure. Based on these analyses, we propose a simple scale of hive status (stable, warning, and collapse) that can be determined based on a few temperature measurements. Our approach offers a lower-cost and practical bee-monitoring solution, providing a non-invasive way to track hive conditions and trigger interventions to save the hives from collapse.

stat.AP↗

Synthetic graphs for link prediction benchmarking

Predicting missing links in complex networks requires algorithms that are able to explore statistical regularities in the existing data. Here we investigate the interplay between algorithm efficiency and network structures through the introduction of suitably-designed synthetic graphs. We propose a family of random graphs that incorporates both micro-scale motifs and meso-scale communities, two ubiquitous structures in complex networks. A key contribution is the derivation of theoretical upper bounds for link prediction performance in our synthetic graphs, allowing us to estimate the predictability of the task and obtain an improved assessment of the performance of any method. Our results on the performance of classical methods (e.g., Stochastic Block Models, Node2Vec,GraphSage) show that the performance of all methods correlate with the theoretical predictability, that no single method is universally superior, and that each of the methods exploit different characteristics known to exist in large classes of networks. Our findings underline the need for careful consideration of graph structure when selecting a link prediction method and emphasize the value of comparing performance against synthetic benchmarks. We provide open-source code for generating these synthetic graphs, enabling further research on link prediction methods.

cs.SI↗

Statistical Laws in Complex Systems

Statistical laws describe regular patterns observed in diverse scientific domains, ranging from the magnitude of earthquakes (Gutenberg-Richter law) and metabolic rates in organisms (Kleiber's law), to the frequency distribution of words in texts (Zipf's and Herdan-Heaps' laws), and productivity metrics of cities (urban scaling laws). The origins of these laws, their empirical validity, and the insights they provide into underlying systems have been subjects of scientific inquiry for centuries. This monograph provides an unifying approach to the study of statistical laws, critically evaluating their role in the theoretical understanding of complex systems and the different data-analysis methods used to evaluate them. Through a historical review and a unified analysis, we uncover that the persistent controversies on the validity of statistical laws are predominantly rooted not in novel empirical findings but in the discordance among data-analysis techniques, mechanistic models, and the interpretations of statistical laws. Starting with simple examples and progressing to more advanced time-series and statistical methods, this monograph and its accompanying repository provide comprehensive material for researchers interested in analyzing data, testing and comparing different laws, and interpreting results in both existing and new datasets.

physics.soc-ph↗

A generative model for community types in directed networks

Large complex networks are often organized into groups or communities. In this paper, we introduce and investigate a generative model of network evolution that reproduces all four pairwise community types that exist in directed networks: assortative, core-periphery, disassortative, and the newly introduced source-basin type. We fix the number of nodes and the community membership of each node, allowing node connectivity to change through rewiring mechanisms that depend on the community membership of the involved nodes. We determine the dependence of the community relationship on the model parameters using a mean-field solution. It reveals that a difference in the swap probabilities of the two communities is a necessary condition to obtain a core-periphery relationship and that a difference in the average in-degree of the communities is a necessary condition for a source-basin relationship. More generally, our analysis reveals multiple possible scenarios for the transition between the different structure types, and sheds light on the mechanisms underlying the observation of the different types of communities in network data.

cs.SI↗

Sampling triangulations of manifolds using Monte Carlo methods

We propose a Monte Carlo method to efficiently find, count, and sample abstract triangulations of a given manifold M. The method is based on a biased random walk through all possible triangulations of M (in the Pachner graph), constructed by combining (bi-stellar) moves with suitable chosen accept/reject probabilities (Metropolis-Hastings). Asymptotically, the method guarantees that samples of triangulations are drawn at random from a chosen probability. This enables us not only to sample (rare) triangulations of particular interest but also to estimate the (extremely small) probability of obtaining them when isomorphism types of triangulations are sampled uniformly at random. We implement our general method for surface triangulations and 1-vertex triangulations of 3-manifolds. To showcase its usefulness, we present a number of experiments: (a) we recover asymptotic growth rates for the number of isomorphism types of simplicial triangulations of the 2-dimensional sphere; (b) we experimentally observe that the growth rate for the number of isomorphism types of 1-vertex triangulations of the 3-dimensional sphere appears to be singly exponential in the number of their tetrahedra; and (c) we present experimental evidence that a randomly chosen isomorphism type of 1-vertex n-tetrahedra 3-sphere triangulation, for n tending to infinity, almost surely shows a fixed edge-degree distribution which decays exponentially for large degrees, but shows non-monotonic behaviour for small degrees.

math.CO↗

Probabilistic description of dissipative chaotic scattering

We investigate the extent to which the probabilistic properties of a chaotic scattering system with dissipation can be understood from the properties of the dissipation-free system. For large energies $E$, a fully chaotic scattering leads to an exponential decay of the survival probability $P(t) \sim e^{-κt}$ with an escape rate $κ$ that decreases with $E$. Dissipation $γ>0$ leads to the appearance of different finite-time regimes in $P(t)$. We show how these different regimes can be understood for small $γ\ll 1$ and $t\gg 1/κ_0$ from the effective escape rate $κ_γ(t)=κ_0(E(t))$ (including the non-hyperbolic regime) until the energy reaches a critical value $E_c$ at which no escape is possible. More generally, we argue that for small dissipation $γ$ and long times $t$ the surviving trajectories in the dissipative system are distributed according to the conditionally invariant measure of the conservative system at the corresponding energy $E(t)<E(0)$. Quantitative predictions of our general theory are compared with numerical simulations in the Henon-Heiles model.

cond-mat.stat-mech↗

Quantifying the Dissimilarity of Texts

Quantifying the dissimilarity of two texts is an important aspect of a number of natural language processing tasks, including semantic information retrieval, topic classification, and document clustering. In this paper, we compared the properties and performance of different dissimilarity measures $D$ using three different representations of texts -- vocabularies, word frequency distributions, and vector embeddings -- and three simple tasks -- clustering texts by author, subject, and time period. Using the Project Gutenberg database, we found that the generalised Jensen--Shannon divergence applied to word frequencies performed strongly across all tasks, that $D$'s based on vector embedding representations led to stronger performance for smaller texts, and that the optimal choice of approach was ultimately task-dependent. We also investigated, both analytically and numerically, the behaviour of the different $D$'s when the two texts varied in length by a factor $h$. We demonstrated that the (natural) estimator of the Jaccard distance between vocabularies was inconsistent and computed explicitly the $h$-dependency of the bias of the estimator of the generalised Jensen--Shannon divergence applied to word frequencies. We also found numerically that the Jensen--Shannon divergence and embedding-based approaches were robust to changes in $h$, while the Jaccard distance was not.

cs.CL↗

Modelling daily weight variation in honey bee hives

A quantitative understanding of the dynamics of bee colonies is important to support global efforts to improve bee health and enhance pollination services. Traditional approaches focus either on theoretical models or data-centred statistical analyses. Here we argue that the combination of these two approaches is essential to obtain interpretable information on the state of bee colonies and show how this can be achieved in the case of time series of intra-day weight variation. We model how the foraging and food processing activities of bees affect global hive weight through a set of ordinary differential equations and show how to estimate reliable ranges for the ten parameters of this model from measurements on a single day. Our analysis of 10 hives at different times shows that crucial indicators of the health of honey bee colonies are estimated robustly and fall in ranges compatible with previously reported results. The indicators include the amount of food collected (foraging success) and the number of active foragers, which may be used to develop early warning indicators of colony failure.

physics.bio-ph↗

Non-parametric power-law surrogates

Power-law distributions are essential in computational and statistical investigations of extreme events and complex systems. The usual technique to generate power-law distributed data is to first infer the scale exponent $α$ using the observed data of interest and then sample from the associated distribution. This approach has important limitations because it relies on a fixed $α$ (e.g., it has limited applicability in testing the {\it family} of power-law distributions) and on the hypothesis of independent observations (e.g., it ignores temporal correlations and other constraints typically present in complex systems data). Here we propose a constrained surrogate method that overcomes these limitations by choosing uniformly at random from a set of sequences exactly as likely to be observed under a discrete power-law as the original sequence (i.e., regardless of $α$) and by showing how additional constraints can be imposed in the sequence (e.g., the Markov transition probability between states). This non-parametric approach involves redistributing observed prime factors to randomize values in accordance with a power-law model but without restricting ourselves to independent observations or to a particular $α$. We test our results in simulated and real data, ranging from the intensity of earthquakes to the number of fatalities in disasters.

nlin.AO↗

Multilayer Networks for Text Analysis with Multiple Data Types

We are interested in the widespread problem of clustering documents and finding topics in large collections of written documents in the presence of metadata and hyperlinks. To tackle the challenge of accounting for these different types of datasets, we propose a novel framework based on Multilayer Networks and Stochastic Block Models. The main innovation of our approach over other techniques is that it applies the same non-parametric probabilistic framework to the different sources of datasets simultaneously. The key difference to other multilayer complex networks is the strong unbalance between the layers, with the average degree of different node types scaling differently with system size. We show that the latter observation is due to generic properties of text, such as Heaps' law, and strongly affects the inference of communities. We present and discuss the performance of our method in different datasets (hundreds of Wikipedia documents, thousands of scientific papers, and thousands of E-mails) showing that taking into account multiple types of information provides a more nuanced view on topic- and document-clusters and increases the ability to predict missing links.

cs.SI↗

Structure of resonance eigenfunctions for chaotic systems with partial escape

Physical systems are often neither completely closed nor completely open, but instead they are best described by dynamical systems with partial escape or absorption. In this paper we introduce classical measures that explain the main properties of resonance eigenfunctions of chaotic quantum systems with partial escape. We construct a family of conditionally-invariant measures with varying decay rates by interpolating between the natural measures of the forward and backward dynamics. Numerical simulations in a representative system show that our classical measures correctly describe the main features of the quantum eigenfunctions: their multi-fractal phase space distribution, their product structure along stable/unstable directions, and their dependence on the decay rate. The (Jensen-Shannon) distance between classical and quantum measures goes to zero in the semiclassical limit for long- and short-lived eigenfunctions, while it remains finite for intermediate cases.

nlin.CD↗

Spatial interactions in urban scaling laws

Analyses of urban scaling laws assume that observations in different cities are independent of the existence of nearby cities. Here we introduce generative models and data-analysis methods that overcome this limitation by modelling explicitly the effect of interactions between individuals at different locations. Parameters that describe the scaling law and the spatial interactions are inferred from data simultaneously, allowing for rigorous (Bayesian) model comparison and overcoming the problem of defining the boundaries of urban regions. Results in five different datasets show that including spatial interactions typically leads to better models and a change in the exponent of the scaling law. Data and codes are provided in Ref. [1].

physics.soc-ph↗

Dynamics of transposable elements generates structure and symmetries in genetic sequences

Genetic sequences are known to possess non-trivial composition together with symmetries in the frequencies of their components. Recently, it has been shown that symmetry and structure are hierarchically intertwined in DNA, suggesting a common origin for both features. However, the mechanism leading to this relationship is unknown. Here we investigate a biologically motivated dynamics for the evolution of genetic sequences. We show that a metastable (long-lived) regime emerges in which sequences have symmetry and structure interlaced in a way that matches that of extant genomes.

q-bio.GN↗

Scaling laws and dynamics of hashtags on Twitter

In this paper we quantify the statistical properties and dynamics of the frequency of hashtag use on Twitter. Hashtags are special words used in social media to attract attention and to organize content. Looking at the collection of all hashtags used in a period of time, we identify the scaling laws underpinning the hashtag frequency distribution (Zipf's law), the number of unique hashtags as a function of sample size (Heaps' law), and the fluctuations around expected values (Taylor's law). While these scaling laws appear to be universal, in the sense that similar exponents are observed irrespective of when the sample is gathered, the volume and nature of the hashtags depends strongly on time, with the appearance of bursts at the minute scale, fat-tailed noise, and long-range correlations. We quantify this dynamics by computing the Jensen-Shannon divergence between hashtag distributions obtained $τ$ times apart and we find that the speed of change decays roughly as $1/τ$. Our findings are based on the analysis of 3.5 billion hashtags used between 2015 and 2016.

physics.soc-ph↗