SearcharxivSearch

arXiv subjects

Thomas Louail

Publications and source records attributed to Thomas Louail.

15 recordsLinked to original sources

The dynamics of discovery and the Heaps-Zipf relationship

When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is referred to as Heaps' law, a regularity often associated with Zipf's law and frequently used to characterize human discovery processes. While random sampling from a Zipf-like distribution can reproduce Heaps' law, this connection relies on the assumption of temporal independence - an assumption often violated in real-world systems although frequently found in the literature. Here, we investigate how temporal correlations in token sequences affect the type-token curve. In human behaviors like music listening and web browsing, domain-specific correlations in token ordering lead to systematic deviations from the Zipf-Heaps framework, effectively decoupling the type-token plot from the rank-frequency distribution. Using a minimal one-parameter model, we reproduce a wide variety of type-token trajectories, including the extremal cases that bound all possible behaviors compatible with a given frequency distribution. Our results demonstrate that type-token growth reflects not only the empirical distribution of type frequencies, but also the domain-specific, temporal structure of the sequence - a factor often overlooked in empirical applications of scaling laws to characterize human behavior.

physics.soc-ph

Do Recommender Systems Promote Local Music? A Reproducibility Study Using Music Streaming Data

This paper examines the influence of recommender systems on local music representation, discussing prior findings from an empirical study on the LFM-2b public dataset. This prior study argued that different recommender systems exhibit algorithmic biases shifting music consumption either towards or against local content. However, LFM-2b users do not reflect the diverse audience of music streaming services. To assess the robustness of this study's conclusions, we conduct a comparative analysis using proprietary listening data from a global music streaming service, which we publicly release alongside this paper. We observe significant differences in local music consumption patterns between our dataset and LFM-2b, suggesting that caution should be exercised when drawing conclusions on local music based solely on LFM-2b. Moreover, we show that the algorithmic biases exhibited in the original work vary in our dataset, and that several unexplored model parameters can significantly influence these biases and affect the study's conclusion on both datasets. Finally, we discuss the complexity of accurately labeling local music, emphasizing the risk of misleading conclusions due to unreliable, biased, or incomplete labels. To encourage further research and ensure reproducibility, we have publicly shared our dataset and code.

cs.IR

Exploring the spatial segmentation of housing markets from online listings

The real estate market shows an inherent connection to space. Real estate agencies unevenly operate and specialize across space, price and type of properties, thereby segmenting the market into submarkets. We introduce here a methodology based on multipartite networks to detect the spatial segmentation emerging from data on housing online listings. Considering the spatial information of the listings, we build a bipartite network that connects agencies and spatial units. This bipartite network is projected into a network of spatial units, whose connections account for similarities in the agency ecosystem. We then apply clustering methods to this network to segment markets into spatially-coherent regions, which are found to be robust across different clustering detection algorithms, discretization of space and spatial scales, and across countries with case studies in France and Spain. This methodology addresses the long-standing issue of housing market segmentation, relevant in disciplines such as urban studies and spatial economics, and with implications for policymaking.

physics.soc-ph

A dominance tree approach to systems of cities

Characterizing the spatial organization of urban systems is a challenge which points to the more general problem of describing marked point processes in spatial statistics. We propose a non-parametric method that goes beyond standard tools of point pattern analysis and which is based on a mapping between the points and a "dominance tree", constructed from a recursive analysis of their Voronoi tessellation. Using toy models, we show that the height of a node in this tree encodes both its mark and the structure of its neighborhood, reflecting its importance in the system. We use historical population data in France (1876-2018) and the US (1880-2010) and show that the method highlights multiscale urban dynamics experienced by these countries. These include non-monotonous city trajectories in the US, as revealed by the evolution of their height in the tree. We show that the height of a city in the tree is less sensitive to different statistical definitions of cities than its rank in the urban hierarchy. The method also captures the attraction basins of cities at successive scales, and while in both countries these basin sizes become more homogeneous at larger scales, they are also more heterogeneous in France than in the US. Finally, we introduce a simple graphical representation - the height clock - that monitors the evolution of the role of each city in its country.

physics.soc-ph

Uplifting Interviews in Social Science with Individual Data Visualization: the case of Music Listening

Collecting accurate and fine-grain information about the music people like, dislike and actually listen to has long been a challenge for sociologists. As millions of people now use online music streaming services, research can build upon the individual listening history data that are collected by these platforms. Individual interviews in particular can benefit from such data, by allowing the interviewers to immerse themselves in the musical universe of consenting respondents, and thus ask them contextualized questions and get more precise answers. Designing a visual exploration tool allowing such an immersion is however difficult, because of the volume and heterogeneity of the listening data, the unequal "visual literacy" of the prospective users, or the interviewers' potential lack of knowledge of the music listened to by the respondents. In this case study we discuss the design and evaluation of such a tool. Designed with social scientists, its purpose is to help them in preparing and conducting semi-structured interviews that address various aspects of the listening experience. It was evaluated during thirty interviews with consenting users of a streaming platform in France.

cs.HC

Follow the guides: disentangling human and algorithmic curation in online music consumption

The role of recommendation systems in the diversity of content consumption on platforms is a much-debated issue. The quantitative state of the art often overlooks the existence of individual attitudes toward guidance, and eventually of different categories of users in this regard. Focusing on the case of music streaming, we analyze the complete listening history of about 9k users over one year and demonstrate that there is no blanket answer to the intertwinement of recommendation use and consumption diversity: it depends on users. First we compute for each user the relative importance of different access modes within their listening history, introducing a trichotomy distinguishing so-called `organic' use from algorithmic and editorial guidance. We thereby identify four categories of users. We then focus on two scales related to content diversity, both in terms of dispersion -- how much users consume the same content repeatedly -- and popularity -- how popular is the content they consume. We show that the two types of recommendation offered by music platforms -- algorithmic and editorial -- may drive the consumption of more or less diverse content in opposite directions, depending also strongly on the type of users. Finally, we compare users' streaming histories with the music programming of a selection of popular French radio stations during the same period. While radio programs are usually more tilted toward repetition than users' listening histories, they often program more songs from less popular artists. On the whole, our results highlight the nontrivial effects of platform-mediated recommendation on consumption, and lead us to speak of `filter niches' rather than `filter bubbles'. They hint at further ramifications for the study and design of recommendation systems.

cs.CY

Is spatial information in ICT data reliable?

An increasing number of human activities are studied using data produced by individuals' ICT devices. In particular, when ICT data contain spatial information, they represent an invaluable source for analyzing urban dynamics. However, there have been relatively few contributions investigating the robustness of this type of results against fluctuations of data characteristics. Here, we present a stability analysis of higher-level information extracted from mobile phone data passively produced during an entire year by 9 million individuals in Senegal. We focus on two information-retrieval tasks: (a) the identification of land use in the region of Dakar from the temporal rhythms of the communication activity; (b) the identification of home and work locations of anonymized individuals, which enable to construct Origin-Destination (OD) matrices of commuting flows. Our analysis reveal that the uncertainty of results highly depends on the sample size, the scale and the period of the year at which the data were gathered. Nevertheless, the spatial distributions of land use computed for different samples are remarkably robust: on average, we observe more than 75% of shared surface area between the different spatial partitions when considering activity of at least 100,000 users whatever the scale. The OD matrix is less stable and depends on the scale with a share of at least 75% of commuters in common when considering all types of flows constructed from the home-work locations of 100,000 users. For both tasks, better results can be obtained at larger levels of aggregation or by considering more users. These results confirm that ICT data are very useful sources for the spatial analysis of urban systems, but that their reliability should in general be tested more thoroughly.

physics.soc-ph

Human Mobility: Models and Applications

Recent years have witnessed an explosion of extensive geolocated datasets related to human movement, enabling scientists to quantitatively study individual and collective mobility patterns, and to generate models that can capture and reproduce the spatiotemporal structures and regularities in human trajectories. The study of human mobility is especially important for applications such as estimating migratory flows, traffic forecasting, urban planning, and epidemic modeling. In this survey, we review the approaches developed to reproduce various mobility patterns, with the main focus on recent developments. This review can be used both as an introduction to the fundamental modeling principles of human mobility, and as a collection of technical methods applicable to specific mobility-related problems. The review organizes the subject by differentiating between individual and population mobility and also between short-range and long-range mobility. Throughout the text the description of the theory is intertwined with real-world applications.

physics.soc-ph

Crowdsourcing the Robin Hood effect in cities

Socioeconomic inequalities in cities are embedded in space and result in neighborhood effects, whose harmful consequences have proved very hard to counterbalance efficiently by planning policies alone. Considering redistribution of money flows as a first step toward improved spatial equity, we study a bottom-up approach that would rely on a slight evolution of shopping mobility practices. Building on a database of anonymized credit card transactions in Madrid and Barcelona, we quantify the mobility effort required to reach a reference situation where commercial income is evenly shared among neighborhoods. The redirections of shopping trips preserve key properties of human mobility, including travel distances. Surprisingly, for both cities only a small fraction ($\sim 5 \%$) of trips need to be altered to reach equity situations, improving even other sustainability indicators. The method could be implemented in mobile applications that would assist individuals in reshaping their shopping practices, to promote the spatial redistribution of opportunities in the city.

physics.soc-ph

Headphones on the wire

We analyze a dataset providing the complete information on the effective plays of thousands of music listeners during several months. Our analysis confirms a number of properties previously highlighted by research based on interviews and questionnaires, but also uncover new statistical patterns, both at the individual and collective levels. In particular, we show that individuals follow common listening rhythms characterized by the same fluctuations, alternating heavy and light listening periods, and can be classified in four groups of similar sizes according to their temporal habits --- 'early birds', 'working hours listeners', 'evening listeners' and 'night owls'. We provide a detailed radioscopy of the listeners' interplay between repeated listening and discovery of new content. We show that different genres encourage different listening habits, from Classical or Jazz music with a more balanced listening among different songs, to Hip Hop and Dance with a more heterogeneous distribution of plays. Finally, we provide measures of how distant people are from each other in terms of common songs. In particular, we show that the number of songs $S$ a DJ should play to a random audience of size $N$ such that everyone hears at least one song he/she currently listens to, is of the form $S\sim N^α$ where the exponent depends on the music genre and is in the range $[0.5,0.8]$. More generally, our results show that the recent access to virtually infinite catalogs of songs does not promote exploration for novelty, but that most users favor repetition of the same songs.

physics.soc-ph

Comparing and modeling land use organization in cities

The advent of geolocated ICT technologies opens the possibility of exploring how people use space in cities, bringing an important new tool for urban scientists and planners, especially for regions where data is scarce or not available. Here we apply a functional network approach to determine land use patterns from mobile phone records. The versatility of the method allows us to run a systematic comparison between Spanish cities of various sizes. The method detects four major land use types that correspond to different temporal patterns. The proportion of these types, their spatial organization and scaling show a strong similarity between all cities that breaks down at a very local scale, where land use mixing is specific to each urban area. Finally, we introduce a model inspired by Schelling's segregation, able to explain and reproduce these results with simple interaction rules between different land uses.

physics.soc-ph

Influence of sociodemographic characteristics on human mobility

Human mobility has been traditionally studied using surveys that deliver snapshots of population displacement patterns. The growing accessibility to ICT information from portable digital media has recently opened the possibility of exploring human behavior at high spatio-temporal resolutions. Mobile phone records, geolocated tweets, check-ins from Foursquare or geotagged photos, have contributed to this purpose at different scales, from cities to countries, in different world areas. Many previous works lacked, however, details on the individuals' attributes such as age or gender. In this work, we analyze credit-card records from Barcelona and Madrid and by examining the geolocated credit-card transactions of individuals living in the two provinces, we find that the mobility patterns vary according to gender, age and occupation. Differences in distance traveled and travel purpose are observed between younger and older people, but, curiously, either between males and females of similar age. While mobility displays some generic features, here we show that sociodemographic characteristics play a relevant role and must be taken into account for mobility and epidemiological modelization.

physics.soc-ph

Uncovering the spatial structure of mobility networks

The extraction of a clear and simple footprint of the structure of large, weighted and directed networks is a general problem that has many applications. An important example is given by origin-destination matrices which contain the complete information on commuting flows, but are difficult to analyze and compare. We propose here a versatile method which extracts a coarse-grained signature of mobility networks, under the form of a $2\times 2$ matrix that separates the flows into four categories. We apply this method to origin-destination matrices extracted from mobile phone data recorded in thirty-one Spanish cities. We show that these cities essentially differ by their proportion of two types of flows: integrated (between residential and employment hotspots) and random flows, whose importance increases with city size. Finally the method allows to determine categories of networks, and in the mobility case to classify cities according to their commuting structure.

physics.soc-ph

Cross-checking different sources of mobility information

The pervasive use of new mobile devices has allowed a better characterization in space and time of human concentrations and mobility in general. Besides its theoretical interest, describing mobility is of great importance for a number of practical applications ranging from the forecast of disease spreading to the design of new spaces in urban environments. While classical data sources, such as surveys or census, have a limited level of geographical resolution (e.g., districts, municipalities, counties are typically used) or are restricted to generic workdays or weekends, the data coming from mobile devices can be precisely located both in time and space. Most previous works have used a single data source to study human mobility patterns. Here we perform instead a cross-check analysis by comparing results obtained with data collected from three different sources: Twitter, census and cell phones. The analysis is focused on the urban areas of Barcelona and Madrid, for which data of the three types is available. We assess the correlation between the datasets on different aspects: the spatial distribution of people concentration, the temporal evolution of people density and the mobility patterns of individuals. Our results show that the three data sources are providing comparable information. Even though the representativeness of Twitter geolocated data is lower than that of mobile phone and census data, the correlations between the population density profiles and mobility patterns detected by the three datasets are close to one in a grid with cells of 2x2 and 1x1 square kilometers. This level of correlation supports the feasibility of interchanging the three data sources at the spatio-temporal scales considered.

physics.soc-ph

From mobile phone data to the spatial structure of cities

Pervasive infrastructures, such as cell phone networks, enable to capture large amounts of human behavioral data but also provide information about the structure of cities and their dynamical properties. In this article, we focus on these last aspects by studying phone data recorded during 55 days in 31 Spanish metropolitan areas. We first define an urban dilatation index which measures how the average distance between individuals evolves during the day, allowing us to highlight different types of city structure. We then focus on hotspots, the most crowded places in the city. We propose a parameter free method to detect them and to test the robustness of our results. The number of these hotspots scales sublinearly with the population size, a result in agreement with previous theoretical arguments and measures on employment datasets. We study the lifetime of these hotspots and show in particular that the hierarchy of permanent ones, which constitute the "heart" of the city, is very stable whatever the size of the city. The spatial structure of these hotspots is also of interest and allows us to distinguish different categories of cities, from monocentric and "segregated" where the spatial distribution is very dependent on land use, to polycentric where the spatial mixing between land uses is much more important. These results point towards the possibility of a new, quantitative classification of cities using high resolution spatio-temporal data.

physics.soc-ph