SearcharxivSearch

arXiv subjects

Sune Lehmann

Publications and source records attributed to Sune Lehmann.

At least 19 recordsLinked to original sources

The limits of visitation entropy as a summary of mobility patterns

Visitation entropy, the Shannon entropy of an individual's distribution of visits across locations, is a widely used metric in the human mobility literature. Yet, its widespread use rests on assumptions that are rarely made explicit: entropy is defined over a fixed set of states, and estimating it empirically requires abundant, well-sampled observations. The limitations that arise when entropy is used outside this setting have been documented and explored in fields such as statistical physics, ecology, and cryptography. The implications for mobility studies, however, remain unclear. Here, we leverage synthetic and empirical trajectories to systematically examine the strengths and weaknesses of visitation entropy as a measure for characterizing human mobility. We show that, for a sequence of locations visited by an individual, the visitation entropy primarily reflects the number of unique locations visited and sequence length, which together explain 90.7% of its variance in empirical data. We also show that shorter sequences systematically lead to an underestimation of entropy, showing finite-sample bias to be a key limitation. As a result, comparisons of groups based on visitation entropy should be treated with caution. We find that apparent entropy gaps between genders, commuters and non-commuters, and urban and rural residents are reduced, or reversed, once sequence length and the number of unique locations are taken into account. Finally, we show that these issues can be addressed by controlling for sequence length and the number of unique locations, and by complementing entropy with structural network measures, which provide more nuanced insight into how mobility is organized beyond the aspects that entropy alone captures.

physics.soc-ph

Women's mobility networks enable more efficient travel

Our understanding of gender differences in mobility is marked by a clear tension: surveys portray women's movements as more complex than men's, while digital traces suggest less diverse travel. Here, we resolve the contradiction by modeling trajectories as networks of sequential visits, using smartphone traces linked to self-reported gender for 543,155 individuals across 10 countries. We show that the apparent conflict in the literature arises because women's mobility networks are simultaneously more clustered and more home-anchored -- a nuance obscured by aggregate metrics. This pattern arises because women tend to link multiple destinations within single trips, for trips spanning up to 150 km and multiple days. This organization yields systematically higher travel efficiency, measured as distance saved through destination chaining over monthly sequences.

physics.soc-ph

Adaptive cut reveals multiscale complexity in networks

Hierarchical clustering and community detection are important problems in machine learning and complex network analysis. A common approach to identify clusters is to simply cut dendrograms at some threshold. However, single-level cuts are often suboptimal in terms of capturing underlying structure in the data, especially when the dendrogram is unbalanced. In this paper, we present the adaptive cut, a novel method that leverages the hierarchical structure of dendrograms by employing multi-level cuts to overcome the limitations of single-level approaches. The adaptive cut optimizes an objective function using a Markov chain Monte Carlo with simulated annealing, resulting in better partitions. We demonstrate the effectiveness of the adaptive cut through applications to link clustering and modularity optimization, but note that the method is applicable to any clustering task that relies on a dendrogram and an objective function. Beyond the adaptive cut, we introduce the balancedness score, an information-theoretic metric that quantifies how balanced a dendrogram is. Balancedness predicts the potential benefits of using multi-level cuts. For the community detection examples, we evaluate our method on more than 200 real-world networks and multiple synthetic datasets, demonstrating significant improvements in partition density and modularity over traditional single-cut approaches. In addition, we show the generality of the adaptive cut by applying it across various hierarchical clustering techniques and objective functions. Our results indicate that the adaptive cut provides a robust and versatile tool for improving clustering outcomes.

physics.soc-ph

Mapping Regional Disparities in Discounted Grocery Products

Food waste represents a major challenge to global climate resilience, accounting for almost 10% of annual greenhouse gas emissions. The retail sector is a critical player, mediating product flows between producers and consumers, where supply chain inefficiencies can shape which items are put on sale. Yet how these dynamics vary across geographic contexts remains largely unexplored. Here, we analyze data from Denmark's largest retail group on near-expiry products put on sale. We uncover the geospatial variations using a dual-clustering approach. We characterize multi-scale spatial relationships in retail organization by correlating store clustering -- measured using shortest-path distances along the street network -- with product clustering based on promotion co-occurrence patterns. Using a bipartite network approach, we identify three regional store clusters, and use percolation thresholds to corroborate the scale of their spatial separation. We find that stores in rural communities put meat and dairy products on sale up to 2.2 times more frequently than metropolitan areas. In contrast, metropolitan and capital regions lean toward convenience products, which have more balanced nutritional profiles but less favorable environmental impacts. By linking geographic context to retail inventory, we provide evidence that reducing food waste requires interventions tailored to local retail dynamics, highlighting the importance of region-specific sustainability strategies.

physics.soc-ph

Modeling roles and trade-offs in multiplex networks

A multiplex social network captures multiple types of social relations among the same set of people, with each layer representing a distinct type of relationship. Understanding the structure of such systems allows us to identify how social exchanges may be driven by a person's own attributes and actions (independence), the status or resources of others (dependence), and mutual influence between entities (interdependence). Characterizing structure in multiplex networks is challenging, as the distinct layers can reflect different yet complementary roles, with interdependence emerging across multiple scales. Here, we introduce the Multiplex Latent Trade-off Model (MLT), a framework for extracting roles in multiplex social networks that accounts for independence, dependence, and interdependence. MLT defines roles as trade-offs, requiring each node to distribute its source and target roles across layers while simultaneously distributing community memberships within hierarchical, multi-scale structures. Applying the MLT approach to 176 real-world multiplex networks, composed of social, health, and economic layers, from villages in western Honduras, we see core social exchange principles emerging, while also revealing local, layer-specific, and multi-scale communities. Link prediction analyses reveal that modeling interdependence yields the greatest performance gains in the social layer, with subtler effects in health and economic layers. This suggests that social ties are structurally embedded, whereas health and economic ties are primarily shaped by individual status and behavioral engagement. Our findings offer new insights into the structure of human social systems.

cs.SI

The origins of large-scale structure in family networks

Family relations are the most fundamental of all social networks and encompass everyone. Family networks grow as individuals have children, creating connections between families, which over time create large and complex structures. While partner-choice homophily has been proposed as a key driver in this growth process, little is known about the connection between individual behavior and the emergent large-scale structure of family networks. Here, we analyze a unique population-complete family network, covering millions of individuals across several decades, enriched with demographic, educational, and geographic data from high-quality national registries. Drawing on the longitudinal coverage of our observations and using a series of growing-network models, we unravel how individual-level behavior shapes the large-scale network structure. Contrary to prevailing theories, we find that partner-choice homophily has little effect on the emergent large-scale structure. Instead, we identify two key drivers: First, partner-change behavior, where individuals leave one partner for another, creates `shortcuts' in the network akin to rewirings in the Watts-Strogatz model. These shortcuts decrease pathlengths and accelerate the emergence of meso-scale connected components. Second, we find that partner change is a self-exciting behavior, such that the probability of changing partner increases with an individual's prior number of partners. The self-exciting behavior accelerates the generation of large network components, with highly connected individuals functioning as network hubs. Accounting for this partner-change behavior, we are able to accurately capture multiple large-scale network properties of the empirical family network. Finally we show that homophily-driven behavior is not able to generate the observed network structure.

physics.soc-ph

Demography-independent behavioural dynamics influenced the spread of COVID-19 in Denmark

Understanding the factors that impact how a communicable disease like COVID-19 spreads is of central importance to mitigate future outbreaks. Traditionally, epidemic surveillance and forecasting analyses have focused on epidemiological data but recent advancements have demonstrated that monitoring behavioural changes may be equally important. Prior studies have shown that high-frequency survey data on social contact behaviour were able to improve predictions of epidemiological observables during the COVID-19 pandemic. Yet, the full potential of such highly granular survey data remains debated. Here, we utilise daily nationally representative survey data from Denmark collected during 23 months of the COVID-19 pandemic to demonstrate two central use-cases for such high-frequency survey data. First, we show that complex behavioural patterns across demographics collapse to a small number of universal key features, greatly simplifying the monitoring and analysis of adherence to outbreak-mitigation measures. Notably, the temporal evolution of the self-reported median number of face-to-face contacts follows a universal behavioural pattern across age groups, with potential to simplify analysis efforts for future outbreaks. Second, we show that these key features can be leveraged to improve deep-learning-based predictions of daily reported new infections. In particular, our models detect a strong link between aggregated self-reported social distancing and hygiene behaviours and the number of new cases in the subsequent days. Taken together, our results highlight the value of high-frequency surveys to improve our understanding of population behaviour in an ongoing public health crisis and its potential use for prediction of central epidemiological observables.

physics.soc-ph

Unveiling the Social Fabric: A Temporal, Nation-Scale Social Network and its Characteristics

Social networks shape individuals' lives, influencing everything from career paths to health. This paper presents a registry-based, multi-layer and temporal network of the entire Danish population in the years 2008-2021 (roughly 7.2 mill. individuals). Our network maps the relationships formed through family, households, neighborhoods, colleagues and classmates. We outline key properties of this multiplex network, introducing both an individual-focused perspective as well as a bipartite representation. We show how to aggregate and combine the layers, and how to efficiently compute network measures such as shortest paths in large administrative networks. Our analysis reveals how past connections reappear later in other layers, that the number of relationships aggregated over time reflects the position in the income distribution, and that we can recover canonical shortest path length distributions when appropriately weighting connections. Along with the network data, we release a Python package that uses the bipartite network representation for efficient analysis.

cs.SI

Decoupling geographical constraints from human mobility

Driven by access to large volumes of movement data, the study of human mobility has grown rapidly over the past decades. The field has shown that human mobility is scale-free, proposed models to generate scale-free moving distance distributions, and explained how the scale-free distribution arises. It has not, however, explicitly addressed how mobility is structured by geographical constraints. How mobility relates to the outlines of landmasses, lakes, and rivers; by the placement of buildings, roadways, and cities. Based on millions of moves, we show how separating the effect of geography from mobility choices, reveals a power law spanning five orders of magnitude. To do so, we incorporate geography via the `pair distribution function' that encapsulates the structure of locations on which mobility occurs. Showing how the spatial distribution of human settlements shapes human mobility, our approach bridges the gap between distance- and opportunity-based models of human mobility.

physics.soc-ph

Is it getting harder to make a hit? Evidence from 65 years of US music chart history

Since the creation of the Billboard Hot 100 music chart in 1958, the chart has been a window into the music consumption of Americans. Which songs succeed on the chart is decided by consumption volumes, which can be affected by consumer music taste, and other factors such as advertisement budgets, airplay time, the specifics of ranking algorithms, and more. Since its introduction, the chart has documented music consumerism through eras of globalization, economic growth, and the emergence of new technologies for music listening. In recent years, musicians and other hitmakers have voiced their worry that the music world is changing: Many claim that it is getting harder to make a hit but until now, the claims have not been backed using chart data. Here we show that the dynamics of the Billboard Hot 100 chart have changed significantly since the chart's founding in 1958, and in particular in the past 15 years. Whereas most songs spend less time on the chart now than songs did in the past, we show that top-1 songs have tripled their chart lifetime since the 1960s, the highest-ranked songs maintain their positions for far longer than previously, and the lowest-ranked songs are replaced more frequently than ever. At the same time, who occupies the chart has also changed over the years: In recent years, fewer new artists make it into the chart and more positions are occupied by established hit makers. Finally, investigating how song chart trajectories have changed over time, we show that historical song trajectories cluster into clear trajectory archetypes characteristic of the time period they were part of. The results are interesting in the context of collective attention: Whereas recent studies have documented that other cultural products such as books, news, and movies fade in popularity quicker in recent years, music hits seem to last longer now than in the past.

physics.soc-ph

Using Smartphones to Study Vaccination Decisions in the Wild

One of the most important tools available to limit the spread and impact of infectious diseases is vaccination. It is therefore important to understand what factors determine people's vaccination decisions. To this end, previous behavioural research made use of, (i) controlled but often abstract or hypothetical studies (e.g., vignettes) or, (ii) realistic but typically less flexible studies that make it difficult to understand individual decision processes (e.g., clinical trials). Combining the best of these approaches, we propose integrating real-world Bluetooth contacts via smartphones in several rounds of a game scenario, as a novel methodology to study vaccination decisions and disease spread. In our 12-week proof-of-concept study conducted with $N$ = 494 students, we found that participants strongly responded to some of the information provided to them during or after each decision round, particularly those related to their individual health outcomes. In contrast, information related to others' decisions and outcomes (e.g., the number of vaccinated or infected individuals) appeared to be less important. We discuss the potential of this novel method and point to fruitful areas for future research.

physics.soc-ph

Time to Cite: Modeling Citation Networks using the Dynamic Impact Single-Event Embedding Model

Understanding the structure and dynamics of scientific research, i.e., the science of science (SciSci), has become an important area of research in order to address imminent questions including how scholars interact to advance science, how disciplines are related and evolve, and how research impact can be quantified and predicted. Central to the study of SciSci has been the analysis of citation networks. Here, two prominent modeling methodologies have been employed: one is to assess the citation impact dynamics of papers using parametric distributions, and the other is to embed the citation networks in a latent space optimal for characterizing the static relations between papers in terms of their citations. Interestingly, citation networks are a prominent example of single-event dynamic networks, i.e., networks for which each dyad only has a single event (i.e., the point in time of citation). We presently propose a novel likelihood function for the characterization of such single-event networks. Using this likelihood, we propose the Dynamic Impact Single-Event Embedding model (DISEE). The \textsc{\modelabbrev} model characterizes the scientific interactions in terms of a latent distance model in which random effects account for citation heterogeneity while the time-varying impact is characterized using existing parametric representations for assessment of dynamic impact. We highlight the proposed approach on several real citation networks finding that the DISEE well reconciles static latent distance network embedding approaches with classical dynamic impact assessments.

cs.SI

Structural Similarities Between Language Models and Neural Response Measurements

Large language models (LLMs) have complicated internal dynamics, but induce representations of words and phrases whose geometry we can study. Human language processing is also opaque, but neural response measurements can provide (noisy) recordings of activation during listening or reading, from which we can extract similar representations of words and phrases. Here we study the extent to which the geometries induced by these representations, share similarities in the context of brain decoding. We find that the larger neural language models get, the more their representations are structurally similar to neural response measurements from brain imaging. Code is available at \url{https://github.com/coastalcph/brainlm}.

cs.CL

Generating fine-grained surrogate temporal networks

Temporal networks are essential for modeling and understanding systems whose behavior varies in time, from social interactions to biological systems. Often, however, real-world data are prohibitively expensive to collect in a large scale or unshareable due to privacy concerns. A promising way to bypass the problem consists in generating arbitrarily large and anonymized synthetic graphs with the properties of real-world networks, namely `surrogate networks'. Until now, the generation of realistic surrogate temporal networks has remained an open problem, due to the difficulty of capturing both the temporal and topological properties of the input network, as well as their correlations, in a scalable model. Here, we propose a novel and simple method for generating surrogate temporal networks. Our method decomposes the input network into star-like structures evolving in time. Then those structures are used as building blocks to generate a surrogate temporal network. Our model vastly outperforms current methods across multiple examples of temporal networks in terms of both topological and dynamical similarity. We further show that beyond generating realistic interaction patterns, our method is able to capture intrinsic temporal periodicity of temporal networks, all with an execution time lower than competing methods by multiple orders of magnitude. The simplicity of our algorithm makes it easily interpretable, extendable and algorithmically scalable.

cs.SI

Using Sequences of Life-events to Predict Human Lives

Over the past decade, machine learning has revolutionized computers' ability to analyze text through flexible computational models. Due to their structural similarity to written language, transformer-based architectures have also shown promise as tools to make sense of a range of multi-variate sequences from protein-structures, music, electronic health records to weather-forecasts. We can also represent human lives in a way that shares this structural similarity to language. From one perspective, lives are simply sequences of events: People are born, visit the pediatrician, start school, move to a new location, get married, and so on. Here, we exploit this similarity to adapt innovations from natural language processing to examine the evolution and predictability of human lives based on detailed event sequences. We do this by drawing on arguably the most comprehensive registry data in existence, available for an entire nation of more than six million individuals across decades. Our data include information about life-events related to health, education, occupation, income, address, and working hours, recorded with day-to-day resolution. We create embeddings of life-events in a single vector space showing that this embedding space is robust and highly structured. Our models allow us to predict diverse outcomes ranging from early mortality to personality nuances, outperforming state-of-the-art models by a wide margin. Using methods for interpreting deep learning models, we probe the algorithm to understand the factors that enable our predictions. Our framework allows researchers to identify new potential mechanisms that impact life outcomes and associated possibilities for personalized interventions.

stat.ML

Far-reaching consequences of trait preferences for animal social network structure and function

Social network structures play an important role in the lives of animals by affecting individual fitness and the spread of disease and information. Nevertheless, we still lack a good understanding of how these structures emerge from the behavior of individuals. Generative network models provide a powerful approach that can help close this gap. Empirical research has shown that trait-based social preferences (preferences for social partners with certain trait values, such as sex, body size, relatedness etc.) play a key role in the formation of social networks across species. Currently, however, we lack a good understanding of how such preferences affect network properties. In this study: 1) we develop a general and flexible generative network model that can create artificial (simulated) networks where social connection is affected by trait-based social preferences; 2) we use this model to investigate how different trait-based social preferences affect social network structure and function. We find that the preferences can affect the efficiency of the networks in terms of transmitting disease and information, and their robustness against fragmentation when individuals disappear, with the effects often - but not always - going in the direction of slower transmission and lower robustness. Furthermore, the extent and form of the effects depend on both the type of preference and the type of trait it is used with. The findings lead to new insights about the potential mechanisms driving the structural diversity of animal social networks, the importance of trait value distributions for social structure, the degree distributions of social networks, and the detectability of trait effects from network data. Overall, the study shows that trait-based social preferences can have far-reaching consequences for populations.

physics.soc-ph

Dialectograms: Machine Learning Differences between Discursive Communities

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word embeddings are complex, high-dimensional spaces and a focus on identifying differences only captures a fraction of their richness. Here, we take a step towards leveraging the richness of the full embedding space, by using word embeddings to map out how words are used differently. Specifically, we describe the construction of dialectograms, an unsupervised way to visually explore the characteristic ways in which each community use a focal word. Based on these dialectograms, we provide a new measure of the degree to which words are used differently that overcomes the tendency for existing measures to pick out low frequent or polysemous words. We apply our methods to explore the discourses of two US political subreddits and show how our methods identify stark affective polarisation of politicians and political entities, differences in the assessment of proper political action as well as disagreement about whether certain issues require political intervention at all.

cs.CL

Monitoring Public Behavior During a Pandemic Using Surveys: Proof-of-Concept Via Epidemic Modelling

Implementing a lockdown for disease mitigation is a balancing act: Non-pharmaceutical interventions can reduce disease transmission significantly, but interventions also have considerable societal costs. Therefore, decision-makers need near real-time information to calibrate the level of restrictions. We fielded daily surveys in Denmark during the second wave of the COVID-19 pandemic to monitor public response to the announced lockdown. A key question asked respondents to state their number of close contacts within the past 24 hours. Here, we establish a link between survey data, mobility data, and hospitalizations via epidemic modelling. Using Bayesian analysis, we then evaluate the usefulness of survey responses as a tool to monitor the effects of lockdown and then compare the predictive performance to that of mobility data. We find that, unlike mobility, self-reported contacts decreased significantly in all regions before the nation-wide implementation of non-pharmaceutical interventions and improved predicting future hospitalizations compared to mobility data. A detailed analysis of contact types indicates that contact with friends and strangers outperforms contact with colleagues and family members (outside the household) on the same prediction task. Representative surveys thus qualify as a reliable, non-privacy invasive monitoring tool to track the implementation of non-pharmaceutical interventions and study potential transmission paths.

physics.data-an