Searcharxiv⌕ Search

arXiv subjects

Elizaveta Sivak

Publications and source records attributed to Elizaveta Sivak.

5 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG↗

Combining the Strengths of Dutch Survey and Register Data in a Data Challenge to Predict Fertility (PreFer)

The social sciences have produced an impressive body of research on determinants of fertility outcomes, or whether and when people have children. However, the strength of these determinants and underlying theories are rarely evaluated on their predictive ability on new data. This prevents us from systematically comparing studies, hindering the evaluation and accumulation of knowledge. In this paper, we present two datasets which can be used to study the predictability of fertility outcomes in the Netherlands. One dataset is based on the LISS panel, a longitudinal survey which includes thousands of variables on a wide range of topics, including individual preferences and values. The other is based on the Dutch register data which lacks attitudinal data but includes detailed information about the life courses of millions of Dutch residents. We provide information about the datasets and the samples, and describe the fertility outcome of interest. We also introduce the fertility prediction data challenge PreFer which is based on these datasets and will start in Spring 2024. We outline the ways in which measuring the predictability of fertility outcomes using these datasets and combining their strengths in the data challenge can advance our understanding of fertility behaviour and computational social science. We further provide details for participants on how to take part in the data challenge.

cs.LG↗

Core but not peripheral online social ties is a protective factor against depression: evidence from a nationally representative sample of young adults

As social interactions are increasingly taking place in the digital environment, online friendship and its effects on various life outcomes from health to happiness attract growing research attention. In most studies, online ties are treated as representing a single type of relationship. However, our online friendship networks are not homogeneous and could include close connections, e.g. a partner, as well as people we have never met in person. In this paper, we investigate the potentially differential effects of online friendship ties on mental health. Using data from a Russian panel study (N = 4,400), we find that - consistently with previous research - the number of online friends correlates with depression symptoms. However, this is true only for networks that do not exceed Dunbar's number in size (N <= 150) and only for core but not peripheral nodes of a friendship network. The findings suggest that online friendship could encode different types of social relationships that should be treated separately while investigating the association between online social integration and life outcomes, in particular well-being or mental health.

cs.SI↗

Measuring Adolescents' Well-being: Correspondence of Naive Digital Traces to Survey Data

Digital traces are often used as a substitute for survey data. However, it is unclear whether and how digital traces actually correspond to the survey-based traits they purport to measure. This paper examines correlations between self-reports and digital trace proxies of depression, anxiety, mood, social integration and sleep among high school students. The study is based on a small but rich multilayer data set (N = 144). The data set contains mood and sleep measures, assessed daily over a 4-month period, along with survey measures at two points in time and information about online activity from VK, the most popular social networking site in Russia. Our analysis indicates that 1) the sentiments expressed in social media posts are correlated with depression; namely, adolescents with more severe symptoms of depression write more negative posts, 2) late-night posting indicates less sleep and poorer sleep quality, and 3) students who were nominated less often as somebody's friend in the survey have fewer friends on VK and their posts receive fewer "likes." However, these correlations are generally weak. These results demonstrate that digital traces can serve as useful supplements to, rather than substitutes for, survey data in studies on adolescents' well-being. These estimates of correlations between survey and digital trace data could provide useful guidelines for future research on the topic.

cs.SI↗

Gender Bias in Sharenting: Both Men and Women Mention Sons More Often Than Daughters on Social Media

Gender inequality starts before birth. Parents tend to prefer boys over girls, which is manifested in reproductive behavior, marital life, and parents' pastimes and investments in their children. While social media and sharing information about children (so-called "sharenting") have become an integral part of parenthood, it is not well-known if and how gender preference shapes online behavior of users. In this paper, we investigate public mentions of daughters and sons on social media. We use data from a popular social networking site on public posts from 635,665 users. We find that both men and women mention sons more often than daughters in their posts. We also find that posts featuring sons get more "likes" on average. Our results indicate that girls are underrepresented in parents' digital narratives about their children. This gender imbalance may send a message that girls are less important than boys, or that they deserve less attention, thus reinforcing gender inequality.

cs.SI↗