SearcharxivSearch

arXiv subjects

Emilio Zagheni

Publications and source records attributed to Emilio Zagheni.

17 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG

Forced Migration and Information-Seeking Behavior on Wikipedia: Insights from the Ukrainian Refugee Crisis

Gathering information about where to migrate is an important part of the migration process, especially during forced migration, when people must make rapid decisions under uncertainty. This study examines how forced migration relates to online information-seeking on Wikipedia. Focusing on the 2022 Russian invasion of Ukraine, we analyze how the resulting refugee crisis, which led to over six million Ukrainians fleeing across Europe, shaped views of Wikipedia articles about European cities. We compare changes in views of Ukrainian-language Wikipedia articles, used as a proxy for information-seeking by Ukrainians, with those in four other language editions. Our findings show that views of Ukrainian-language articles about European cities correlate more strongly with the number of Ukrainian refugees applying for temporary protection in European countries than views in other languages. Because Poland and Germany became the main destinations for refugees, we examine these countries more closely and find that applications for temporary protection in Polish and German cities are also more strongly correlated with views of their Ukrainian-language Wikipedia articles. We further analyze the timing between refugee flows to Poland and online information-seeking. Refugee border crossings occurred before increases in Ukrainian-language views of Polish city articles, indicating that information-seeking surged after displacement. This reactive pattern contrasts with the pre-departure planning typical of regular labor migration. Moreover, while official protection applications often lagged behind border crossings by weeks, Wikipedia activity rose almost immediately. Overall, Wikipedia usage offers a near real-time indicator of emerging migration patterns during crises.

cs.CY

The social stratification of internal migration and daily mobility during the COVID-19 pandemic

This study leverages mobile phone data for 5.4 million users to unveil the complex dynamics of internal migration and daily mobility in Santiago de Chile during the global COVID-19 pandemic, with a focus on socioeconomic differentials. Major findings include an increase in daily mobility among lower-income brackets compared to higher ones in 2020. In contrast, long-term relocation patterns rose primarily among higher-income groups. These shifts indicate a nuanced response to the pandemic across socioeconomic strata. Unlike in 2017, economic factors in 2020 influenced a change not only in the decision to emigrate but also in the selection of destinations, suggesting a profound transformation in mobility behaviors. Contrary to expectations, there was no evidence supporting a preference for rural over urban destinations despite the surge in emigration from Santiago during the pandemic. The study enhances our understanding of how varying socioeconomic conditions intersect with mobility decisions during crises and provides valuable insights for policymakers aiming to enact fair, informed measures in rapidly changing circumstances.

cs.CY

Return migration of German-affiliated researchers: Analyzing departure and return by gender, cohort, and discipline using Scopus bibliometric data 1996-2020

The international migration of researchers is an important dimension of scientific mobility, and has been the subject of considerable policy debate. However, tracking the migration life courses of researchers is challenging due to data limitations. In this study, we use Scopus bibliometric data on eight million publications from 1.1 million researchers who have published at least once with an affiliation address from Germany in 1996-2020. We construct the partial life histories of published researchers in this period and explore both their out-migration and the subsequent return of a subset of this group: the returnees. Our analyses shed light on the career stages and gender disparities between researchers who remain in Germany, those who emigrate, and those who eventually return. We find that the return migration streams are even more gender imbalanced, which points to the need for additional efforts to encourage female researchers to come back to Germany. We document a slightly declining trend in return migration among more recent cohorts of researchers who left Germany, which, for most disciplines, was associated with a decrease in the German collaborative ties of these researchers. Moreover, we find that the gender disparities for the most gender imbalanced disciplines are unlikely to be mitigated by return migration given the gender compositions of the cohorts of researchers who have left Germany and of those who have returned. This analysis uncovers new dimensions of migration among scholars by investigating the return migration of published researchers, which is critical for the development of science policy.

cs.DL

International Migration in Academia and Citation Performance: An Analysis of German-Affiliated Researchers by Gender and Discipline Using Scopus Publications 1996-2020

Germany has become a major country of immigration, as well as a research powerhouse in Europe. As Germany spends a higher fraction of its GDP on research and development than most countries with advanced economies, there is an expectation that Germany should be able to attract and retain international scholars who have high citation performance. Using an exhaustive set of over eight million Scopus publications, we analyze the trends in international migration to and from Germany among published researchers over the past 24 years. We assess changes in institutional affiliations for over one million researchers who have published with a German affiliation address at some point during the 1996-2020 period. We show that while Germany has been highly integrated into the global movement of researchers, with particularly strong ties to the US, the UK, and Switzerland, the country has been sending more published researchers abroad than it has attracted. While the balance has been largely negative over time, analyses disaggregated by gender, citation performance, and field of research show that compositional differences in migrant flows may help to alleviate persistent gender inequalities in selected fields.

cs.DL

Scholarly migration within Mexico: Analyzing internal migration among researchers using Scopus longitudinal bibliometric data

The migration of scholars is a major driver of innovation and of diffusion of knowledge. Although large-scale bibliometric data have been used to measure international migration of scholars, our understanding of internal migration among researchers is very limited. This is partly due to a lack of data aggregated at a suitable sub-national level. In this study, we analyze internal migration in Mexico based on over 1.1 million authorship records from the Scopus database. We trace the movements of scholars between Mexican states, and provide key demographic measures of internal migration for the 1996-2018 period. From a methodological perspective, we develop a new framework for enhancing data quality, inferring states from affiliations, and detecting moves from modal states for the purposes of studying internal migration among researchers. Substantively, we combine demographic and network science techniques to improve our understanding of internal migration patterns within country boundaries. The migration patterns between states in Mexico appear to be heterogeneous in size and direction across regions. However, while many scholars remain in their regions, there seems to be a preference for Mexico City and the surrounding states as migration destinations. We observed that over the past two decades, there has been a general decreasing trend in the crude migration intensity. However, the migration network has become more dense and more diverse, and has included greater exchanges between states along the Gulf and the Pacific Coast. Our analysis, which is mostly empirical in nature, lays the foundations for testing and developing theories that can rely on the analytical framework developed by migration scholars, and the richness of appropriately processed bibliometric data.

cs.DL

How Biased is the Population of Facebook Users? Comparing the Demographics of Facebook Users with Census Data to Generate Correction Factors

Censuses around the world are key sources of data to guide government investments and public policies. However, these sources are very expensive to obtain and are collected relatively infrequently. Over the last decade, there has been growing interest in the use of data from social media to complement traditional data sources. However, social media users are not representative of the general population. Thus, analyses based on social media data require statistical adjustments, like post-stratification, in order to remove the bias and make solid statistical claims. These adjustments are possible only when we have information about the frequency of demographic groups using social media. These data, when compared with official statistics, enable researchers to produce appropriate statistical correction factors. In this paper, we leverage the Facebook advertising platform to compile the equivalent of an aggregate-level census of Facebook users. Our compilation includes the population distribution for seven demographic attributes such as gender and age at different geographic levels for the US. By comparing the Facebook counts with official reports provided by the US Census and Gallup, we found very high correlations, especially for political leaning and race. We also identified instances where official statistics may be underestimating population counts as in the case of immigration. We use the information collected to calculate bias correction factors for all computed attributes in order to evaluate the extent to which different demographic groups are more or less represented on Facebook. We provide the first comprehensive analysis for assessing biases in Facebook users across several dimensions. This information can be used to generate bias-adjusted population estimates and demographic counts in a timely way and at fine geographic granularity in between data releases of official statistics

cs.SI

Combining social media and survey data to nowcast migrant stocks in the United States

Measuring and forecasting migration patterns, and how they change over time, has important implications for understanding broader population trends, for designing policy effectively and for allocating resources. However, data on migration and mobility are often lacking, and those that do exist are not available in a timely manner. Social media data offer new opportunities to provide more up-to-date demographic estimates and to complement more traditional data sources. Facebook, for example, can be thought of as a large digital census that is regularly updated. However, its users are not representative of the underlying population. This paper proposes a statistical framework to combine social media data with traditional survey data to produce timely `nowcasts' of migrant stocks by state in the United States. The model incorporates bias adjustment of the Facebook data, and a pooled principal component time series approach, to account for correlations across age, time and space. We illustrate the results for migrants from Mexico, India and Germany, and show that the model outperforms alternatives that rely solely on either social media or survey data.

stat.AP

The demography of the peripatetic researcher: Evidence on highly mobile scholars from the Web of Science

The policy debate around researchers' geographic mobility has been moving away from a theorized zero-sum game in which countries can be winners (brain gain) or losers (brain drain), and toward the concept of brain circulation, which implies that researchers move in and out of countries and everyone benefits. Quantifying trends in researchers' movements is key to understanding the drivers of the mobility of talent, as well as the implications of these patterns for the global system of science, and for the competitive advantages of individual countries. Existing studies have investigated bilateral flows of researchers. However, in order to understand migration systems, determining the extent to which researchers have worked in more than two countries is essential. This study focuses on the subgroup of highly mobile researchers whom we refer to as peripatetic researchers or super-movers. More specifically, our aim is to track the international movements of researchers who have published in more than two countries through changes in the main affiliation addresses of researchers in over 62 million publications indexed in the Web of Science database over the 1956-2016 period. Using this approach, we have established a longitudinal dataset on the international movements of highly mobile researchers across all subject categories, and in all disciplines of scholarship. This article contributes to the literature by offering for the first time a snapshot of the key features of highly mobile researchers, including their patterns of migration and return migration by academic age, the relative frequency of their disciplines, and the relative frequency of their countries of origin and destination. Among other findings, the results point to the emergence of a global system that includes the USA and China as two large hubs, and England and Germany as two smaller hubs for highly mobile researchers.

cs.DL

Demographic Differentials in Facebook Usage Around the World

We use data from the Facebook Advertisement Platform to study patterns of demographic disparities in usage of Facebook across countries. We address three main questions: (1) How does Facebook usage differ by age and by gender around the world? (2) How does the size of friendship networks vary by age and by gender? (3) What are the demographic characteristics of specific subgroups of Facebook users? We find that in countries in North America and northern Europe, patterns of Facebook usage differ little between older people and younger adults. In Asian countries, which have high levels of gender inequality, differences in Facebook adoption by gender disappear at older ages, possibly as a result of selectivity. We also observe that across countries, women tend to have larger networks of close friends than men, and that female users who are living away from their hometown are more likely to engage in Facebook use than their male counterparts, regardless of their region and age group. Our findings contextualize recent research on gender gaps in online usage, and offer new insights into some of the nuances of demographic differentials in the adoption and the use of digital technologies.

cs.SI

Rock, Rap, or Reggaeton?: Assessing Mexican Immigrants' Cultural Assimilation Using Facebook Data

The degree to which Mexican immigrants in the U.S. are assimilating culturally has been widely debated. To examine this question, we focus on musical taste, a key symbolic resource that signals the social positions of individuals. We adapt an assimilation metric from earlier work to analyze self-reported musical interests among immigrants in Facebook. We use the relative levels of interest in musical genres, where a similarity to the host population in musical preferences is treated as evidence of cultural assimilation. Contrary to skeptics of Mexican assimilation, we find significant cultural convergence even among first-generation immigrants, which problematizes their use as assimilative "benchmarks" in the literature. Further, 2nd generation Mexican Americans show high cultural convergence vis-à-vis both Anglos and African-Americans, with the exception of those who speak Spanish. Rather than conforming to a single assimilation path, our findings reveal how Mexican immigrants defy simple unilinear theoretical expectations and illuminate their uniquely heterogeneous character.

cs.SI

Studying Migrant Assimilation Through Facebook Interests

Migrants' assimilation is a major challenge for European societies, in part because of the sudden surge of refugees in recent years and in part because of long-term demographic trends. In this paper, we use Facebook's data for advertisers to study the levels of assimilation of Arabic-speaking migrants in Germany, as seen through the interests they express online. Our results indicate a gradient of assimilation along demographic lines, language spoken and country of origin. Given the difficulty to collect timely migration data, in particular for traits related to cultural assimilation, the methods that we develop and the results that we provide open new lines of research that computational social scientists are well-positioned to address.

cs.SI

Mater certa est, pater numquam: What can Facebook Advertising Data Tell Us about Male Fertility Rates?

In many developing countries, timely and accurate information about birth rates and other demographic indicators is still lacking, especially for male fertility rates. Using anonymous and aggregate data from Facebook's Advertising Platform, we produce global estimates of the Mean Age at Childbearing (MAC), a key indicator of fertility postponement. Our analysis indicates that fertility measures based on Facebook data are highly correlated with conventional indicators based on traditional data, for those countries for which we have statistics. For instance, the correlation of the MAC computed using Facebook and United Nations data is 0.47 (p = 4.02e-08) and 0.79 (p = 2.2e-15) for female and male respectively. Out of sample validation for a simple regression model indicates that the mean absolute percentage error is 2.3%. We use the linear model and Facebook data to produce estimates of the male MAC for countries for which we do not have data.

cs.CY

Professional Gender Gaps Across US Cities

Gender imbalances in work environments have been a long-standing concern. Identifying the existence of such imbalances is key to designing policies to help overcome them. In this work, we study gender trends in employment across various dimensions in the United States. This is done by analyzing anonymous, aggregate statistics that were extracted from LinkedIn's advertising platform. The data contain the number of male and female LinkedIn users with respect to (i) location, (ii) age, (iii) industry and (iv) certain skills. We studied which of these categories correlate the most with high relative male or female presence on LinkedIn. In addition to examining the summary statistics of the LinkedIn data, we model the gender balance as a function of the different employee features using linear regression. Our results suggest that the gender gap varies across all feature types, but the differences are most profound among industries and skills. A high correlation between gender ratios of people in our LinkedIn data set and data provided by the US Bureau of Labor Statistics serves as external validation for our results.

cs.SI

Fertility and its Meaning: Evidence from Search Behavior

Fertility choices are linked to the different preferences and constraints of individuals and couples, and vary importantly by socio-economic status, as well by cultural and institutional context. The meaning of childbearing and child-rearing, therefore, differs between individuals and across groups. In this paper, we combine data from Google Correlate and Google Trends for the U.S. with ground truth data from the American Community Survey to derive new insights into fertility and its meaning. First, we show that Google Correlate can be used to illustrate socio-economic differences on the circumstances around pregnancy and birth: e.g., searches for "flying while pregnant" are linked to high income fertility, and "paternity test" are linked to non-marital fertility. Second, we combine several search queries to build predictive models of regional variation in fertility, explaining about 75% of the variance. Third, we explore if aggregated web search data can also be used to model fertility trends.

cs.CY

A Flexible Bayesian Model for Estimating Subnational Mortality

Reliable mortality estimates at the subnational level are essential in the study of health inequalities within a country. One of the difficulties in producing such estimates is the presence of small populations, where the stochastic variation in death counts is relatively high, and so the underlying mortality levels are unclear. We present a Bayesian hierarchical model to estimate mortality at the subnational level. The model builds on characteristic age patterns in mortality curves, which are constructed using principal components from a set of reference mortality curves. Information on mortality rates are pooled across geographic space and smoothed over time. Testing of the model shows reasonable estimates and uncertainty levels when the model is applied to both simulated data which mimic US counties, and real data for French departments. The estimates produced by the model have direct applications to the study of subregional health patterns and disparities.

stat.AP

From Migration Corridors to Clusters: The Value of Google+ Data for Migration Studies

Recently, there have been considerable efforts to use online data to investigate international migration. These efforts show that Web data are valuable for estimating migration rates and are relatively easy to obtain. However, existing studies have only investigated flows of people along migration corridors, i.e. between pairs of countries. In our work, we use data about "places lived" from millions of Google+ users in order to study migration "clusters", i.e. groups of countries in which individuals have lived. For the first time, we consider information about more than two countries people have lived in. We argue that these data are very valuable because this type of information is not available in traditional demographic sources which record country-to-country migration flows independent of each other. We show that migration clusters of country triads cannot be identified using information about bilateral flows alone. To demonstrate the additional insights that can be gained by using data about migration clusters, we first develop a model that tries to predict the prevalence of a given triad using only data about its constituent pairs. We then inspect the groups of three countries which are more or less prominent, compared to what we would expect based on bilateral flows alone. Next, we identify a set of features such as a shared language or colonial ties that explain which triple of country pairs are more or less likely to be clustered when looking at country triples. Then we select and contrast a few cases of clusters that provide some qualitative information about what our data set shows. The type of data that we use is potentially available for a number of social media services. We hope that this first study about migration clusters will stimulate the use of Web data for the development of new theories of international migration that could not be tested appropriately before.

cs.SI