SearcharxivSearch

arXiv subjects

Lucas Sage

Publications and source records attributed to Lucas Sage.

4 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG

Embeddings of Nation-Level Social Networks

Full nation-scale social networks are now emerging from countries such as the Netherlands and Denmark, but these networks present challenging technical issues in working with large, multiplex, time-dependent networks. We report on our experiences in producing dynamic node embeddings of the population network of the Netherlands. We present (a) a layer-sensitive random walk strategy which improves on traditional flattening methods for multiplex networks, (b) a temporal alignment strategy that brings annual networks into the same embedding space, without leaking information to future years, and (c) the use of Fibonacci spirals and embedding whitening techniques for more balanced and effective partitioning. We demonstrate the effectiveness of these techniques in building embedding-based models for 13 downstream tasks.

cs.SI

Inferring cumulative advantage from longitudinal records

Inequality in human success may emerge through endogenous success-breeds-success dynamics but may also originate in pre-existing differences in talent. It is widely recognized that the skew in static frequency distributions of success implied by a cumulative advantage model is also consistent with a talent model. Studies have turned to longitudinal records of success, seeking to exploit the time dimension for adjudication. Here we show that success histories suffer from a similar identification problem as static distributional evidence. We prove that for any talent model there exists an analogous path dependent model that generates the same longitudinal predictions, and vice versa. We formally identify such twins for prominent models in the literature, in both directions. These results imply that longitudinal data previously interpreted to support a talent model equally well fits a model of cumulative advantage and vice versa.

math.PR

Can ethnic tolerance curb self-reinforcing school segregation? A theoretical Agent Based Model

Schelling and Sakoda prominently proposed computational models suggesting that strong ethnic residential segregation can be the unintended outcome of a self-reinforcing dynamic driven by choices of individuals with rather tolerant ethnic preferences. There are only few attempts to apply this view to school choice, another important arena in which ethnic segregation occurs. In the current paper, we explore with an agent-based theoretical model similar to those proposed for residential segregation, how ethnic tolerance among parents can affect the level of school segregation. More specifically, we ask whether and under which conditions school segregation could be reduced if more parents hold tolerant ethnic preferences. We move beyond earlier models of school segregation in three ways. First, we model individual school choices using a random utility discrete choice approach. Second, we vary the pattern of ethnic segregation in the residential context of school choices systematically, comparing residential maps in which segregation is unrelated to parents' level of tolerance to residential maps reflecting their ethnic preferences. Thirdly, we introduce heterogeneity in tolerance levels among parents belonging to the same group. Our simulation experiments suggest that ethnic school segregation can be a very robust phenomenon, occurring even when about half of the population prefers mixed to segregated schools. However, we also identify a sweet spot in the parameter space in which a larger proportion of tolerant parents makes the biggest difference. This is the case when the preference for nearby schools weighs heavily in parents' utility function and the residential map is only moderately segregated. Further experiments are presented that unravel the underlying mechanisms.

physics.soc-ph