SearcharxivSearch

arXiv subjects

Louis Boucherie

Publications and source records attributed to Louis Boucherie.

7 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG

Geometry and Geography of Complex Networks

Complex systems are made up of many interacting components. Network science provides the tools to analyze and understand these interactions. Community detection is a key technique in network science for uncovering the structures that shape the behavior of these networks. This thesis introduces the Adaptive Cut, a novel method that improves clustering methods by employing multi-level cuts in hierarchical dendrograms. Overcoming the limitations of traditional single-level cuts-especially in unbalanced dendrograms-the Adaptive Cut provides a multi-level cut by optimizing a Markov chain Monte Carlo with simulated annealing. In addition, we propose the Balanceness score, an information-theoretic metric that quantifies dendrogram balance and predicts the benefits of multilevel cuts. Evaluations on over 200 real and synthetic networks show significant improvements in partition density and modularity. In the second part, our analysis shows that incorporating network geometry allows redefining administrative boundaries into non-contiguous regions that better reflect social and spatial dynamics. We also discuss the representation of hierarchical data in hyperbolic space through Poincare maps, which can represent tree-like structures in low dimension. In addition, we examine how geography constrains human mobility, an aspect often overlooked in scale-free characterizations of mobility. By incorporating geography via the pair distribution function from condensed matter physics, we separate geographic constraints from mobility choices. Analyzing datasets containing millions of individual movements, we identify a universal power law that spans five orders of magnitude, thereby bridging the divide between distance-based and opportunity-driven models of human mobility.

physics.soc-ph

Adaptive cut reveals multiscale complexity in networks

Hierarchical clustering and community detection are important problems in machine learning and complex network analysis. A common approach to identify clusters is to simply cut dendrograms at some threshold. However, single-level cuts are often suboptimal in terms of capturing underlying structure in the data, especially when the dendrogram is unbalanced. In this paper, we present the adaptive cut, a novel method that leverages the hierarchical structure of dendrograms by employing multi-level cuts to overcome the limitations of single-level approaches. The adaptive cut optimizes an objective function using a Markov chain Monte Carlo with simulated annealing, resulting in better partitions. We demonstrate the effectiveness of the adaptive cut through applications to link clustering and modularity optimization, but note that the method is applicable to any clustering task that relies on a dendrogram and an objective function. Beyond the adaptive cut, we introduce the balancedness score, an information-theoretic metric that quantifies how balanced a dendrogram is. Balancedness predicts the potential benefits of using multi-level cuts. For the community detection examples, we evaluate our method on more than 200 real-world networks and multiple synthetic datasets, demonstrating significant improvements in partition density and modularity over traditional single-cut approaches. In addition, we show the generality of the adaptive cut by applying it across various hierarchical clustering techniques and objective functions. Our results indicate that the adaptive cut provides a robust and versatile tool for improving clustering outcomes.

physics.soc-ph

Cultural evolution of human beauty standards

Beauty standards shape self-perception and health through social comparison and objectification, while exposure to idealized imagery exacerbates body-image concerns. Media and fashion are central arbiters of these ideals, yet long-term, quantitative, intersectional studies on how representation has changed remain scarce. We assembled a dataset of 793199 records spanning 25 years of advertising, magazine covers, runway shows, and editorials to quantify changes in anthropometric and demographic representation. We find a paradox in the evolution of beauty ideals: while representational diversity has increased, the median model physique remains stable. This is driven by selective plus-size inclusion at the upper tail, while the typical physique continues to diverge from the US population. Intersectionally, non-white models are 4.5 times more likely to be plus-size, indicating that progress in size inclusivity falls disproportionately on multiple underrepresented identities. Stratifying the industry via a data-driven prestige hierarchy, we find that thinness is overrepresented at the top tier. Finally, comparing two regulatory interventions we observe that numeric thresholds are more effective at reducing underweight appearances. Our results quantify the cultural evolution in media and fashion, revealing that inclusion has increased; however, gains are uneven and intersectionally concentrated on size and ethnicity, whereas the prevailing thin ideal remains largely unchanged.

physics.soc-ph

Decoupling geographical constraints from human mobility

Driven by access to large volumes of movement data, the study of human mobility has grown rapidly over the past decades. The field has shown that human mobility is scale-free, proposed models to generate scale-free moving distance distributions, and explained how the scale-free distribution arises. It has not, however, explicitly addressed how mobility is structured by geographical constraints. How mobility relates to the outlines of landmasses, lakes, and rivers; by the placement of buildings, roadways, and cities. Based on millions of moves, we show how separating the effect of geography from mobility choices, reveals a power law spanning five orders of magnitude. To do so, we incorporate geography via the `pair distribution function' that encapsulates the structure of locations on which mobility occurs. Showing how the spatial distribution of human settlements shapes human mobility, our approach bridges the gap between distance- and opportunity-based models of human mobility.

physics.soc-ph

Time to Cite: Modeling Citation Networks using the Dynamic Impact Single-Event Embedding Model

Understanding the structure and dynamics of scientific research, i.e., the science of science (SciSci), has become an important area of research in order to address imminent questions including how scholars interact to advance science, how disciplines are related and evolve, and how research impact can be quantified and predicted. Central to the study of SciSci has been the analysis of citation networks. Here, two prominent modeling methodologies have been employed: one is to assess the citation impact dynamics of papers using parametric distributions, and the other is to embed the citation networks in a latent space optimal for characterizing the static relations between papers in terms of their citations. Interestingly, citation networks are a prominent example of single-event dynamic networks, i.e., networks for which each dyad only has a single event (i.e., the point in time of citation). We presently propose a novel likelihood function for the characterization of such single-event networks. Using this likelihood, we propose the Dynamic Impact Single-Event Embedding model (DISEE). The \textsc{\modelabbrev} model characterizes the scientific interactions in terms of a latent distance model in which random effects account for citation heterogeneity while the time-varying impact is characterized using existing parametric representations for assessment of dynamic impact. We highlight the proposed approach on several real citation networks finding that the DISEE well reconciles static latent distance network embedding approaches with classical dynamic impact assessments.

cs.SI

Characterizing Polarization in Social Networks using the Signed Relational Latent Distance Model

Graph representation learning has become a prominent tool for the characterization and understanding of the structure of networks in general and social networks in particular. Typically, these representation learning approaches embed the networks into a low-dimensional space in which the role of each individual can be characterized in terms of their latent position. A major current concern in social networks is the emergence of polarization and filter bubbles promoting a mindset of "us-versus-them" that may be defined by extreme positions believed to ultimately lead to political violence and the erosion of democracy. Such polarized networks are typically characterized in terms of signed links reflecting likes and dislikes. We propose the latent Signed relational Latent dIstance Model (SLIM) utilizing for the first time the Skellam distribution as a likelihood function for signed networks and extend the modeling to the characterization of distinct extreme positions by constraining the embedding space to polytopes. On four real social signed networks of polarization, we demonstrate that the model extracts low-dimensional characterizations that well predict friendships and animosity while providing interpretable visualizations defined by extreme positions when endowing the model with an embedding space restricted to polytopes.

stat.ML