SearcharxivSearch

arXiv subjects

Charles Rahal

Publications and source records attributed to Charles Rahal.

7 recordsLinked to original sources

Births are difficult to predict even with rich survey and full-population register data

Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.

cs.LG

Gender and the Production of Research Impact

Evaluating the impact of scientific research beyond academia --- on policy, health, the economy, and cultural life --- has become a cornerstone of science policy and research-funding allocation worldwide. Yet which researchers produce the research underpinning this impact, and how this production is shaped by gender, remains poorly understood. We combine structured and unstructured records from the United Kingdom's latest Research Excellence Framework, the largest national research assessment currently in operation, with large-scale bibliometric data to quantify gender differences among the researchers underpinning documented impact. Women account for 38.16% of these contributors: underrepresented overall, but with a consistently higher share than in research-output authorships (33.63%), both overall and across all four REF panels. Impact production is also strongly gendered across domains: women are better represented in case studies concerning education, health, cultural, and civil-society impact, whereas those concerning commercialisation pathways such as patenting and manufacturing remain dominated by men. These findings reveal critical inequalities across the pathways that connect research to impact beyond academia, offering crucial evidence for policymakers and academic institutions aiming to build more equitable and representative systems for evaluating scientific contributions.

cs.DL

RobustiPy: An efficient next generation multiversal library with model selection, averaging, resampling, and explainable artificial intelligence

Scientific inference is often undermined by the vast but rarely explored "multiverse" of defensible modelling choices, which can generate results as variable as the phenomena under study. We introduce RobustiPy, an open-source Python library that systematizes multiverse analysis and model-uncertainty quantification at scale. RobustiPy unifies bootstrap-based inference, combinatorial specification search, model selection and averaging, joint-inference routines, and explainable AI methods within a modular, reproducible framework. Beyond exhaustive specification curves, it supports rigorous out-of-sample validation and quantifies the marginal contribution of each covariate. We demonstrate its utility across five simulation designs and ten empirical case studies spanning economics, sociology, psychology, and medicine, including a re-analysis of widely cited findings with documented discrepancies. Benchmarking on ~672 million simulated regressions shows that RobustiPy delivers state-of-the-art computational efficiency while expanding transparency in empirical research. By standardizing and accelerating robustness analysis, RobustiPy transforms how researchers interrogate sensitivity across the analytical multiverse, offering a practical foundation for more reproducible and interpretable computational science.

stat.ME

Capitalizing on a Crisis: A Computational Analysis of all Five Million British Firms During the Covid-19 Pandemic

The Covid-19 pandemic brought unprecedented changes to business ownership in the UK which affects a generation of entrepreneurs and their employees. Nonetheless, the impact remains poorly understood. This is because research on capital accumulation has typically lacked high-quality, individualized, population-level data. We overcome these barriers to examine who benefits from economic crises through a computationally orientated lens of firm creation. Leveraging a comprehensive cache of administrative data on every UK firm and all nine million people running them, combined with probabilistic algorithms, we conduct individual-level analyses to understand who became Covid entrepreneurs. Using these techniques, we explore characteristics of entrepreneurs--such as age, gender, region, business experience, and industry--which potentially predict Covid entrepreneurship. By employing an automated time series model selection procedure to generate counterfactuals, we show that Covid entrepreneurs were typically aged 35-49 (40.4%), men (73.1%), and had previously held roles in existing firms (59.4%). For most industries, growth was disproportionately concentrated around London. It was therefore existing corporate elites who were most able to capitalize on the Covid crisis and not, as some hypothesized, young entrepreneurs who were setting up their first businesses. In this respect, the pandemic will likely impact future wealth inequalities. Our work offers methodological guidance for future policymakers during economic crises and highlights the long-term consequences for capital and wealth inequality.

econ.GN

On the Unknowable Limits to Prediction

We propose a rigorous decomposition of predictive error, highlighting that not all 'irreducible' error is genuinely immutable. Many domains stand to benefit from iterative enhancements in measurement, construct validity, and modeling. Our approach demonstrates how apparently 'unpredictable' outcomes can become more tractable with improved data (across both target and features) and refined algorithms. By distinguishing aleatoric from epistemic error, we delineate how accuracy may asymptotically improve--though inherent stochasticity may remain--and offer a robust framework for advancing computational research.

cs.LG

Estimating the Cost of Informal Care with a Novel Two-Stage Approach to Individual Synthetic Control

Informal carers provide the majority of care for people living with challenges related to older age, long-term illness, or disability. However, the care they provide often results in a significant income penalty for carers, a factor largely overlooked in the economics literature and policy discourse. Leveraging data from the UK Household Longitudinal Study, this paper provides the first robust causal estimates of the caring income penalty using a novel individual synthetic control based method that accounts for unit-level heterogeneity in post-treatment trajectories over time. Our baseline estimates identify an average relative income gap of up to 45%, with an average decrease of {\pounds}162 in monthly income, peaking at {\pounds}192 per month after 4 years, based on the difference between informal carers providing the highest-intensity of care and their synthetic counterparts. We find that the income penalty is more pronounced for women than for men, and varies by ethnicity and age.

econ.GN

Social network-based distancing strategies to flatten the COVID 19 curve in a post-lockdown world

Social distancing and isolation have been introduced widely to counter the COVID-19 pandemic. However, more moderate contact reduction policies become desirable owing to adverse social, psychological, and economic consequences of a complete or near-complete lockdown. Adopting a social network approach, we evaluate the effectiveness of three targeted distancing strategies designed to 'keep the curve flat' and aid compliance in a post-lockdown world. These are limiting interaction to a few repeated contacts, seeking similarity across contacts, and strengthening communities via triadic strategies. We simulate stochastic infection curves that incorporate core elements from infection models, ideal-type social network models, and statistical relational event models. We demonstrate that strategic reduction of contact can strongly increase the efficiency of social distancing measures, introducing the possibility of allowing some social contact while keeping risks low. This approach provides nuanced insights to policy makers for effective social distancing that can mitigate negative consequences of social isolation.

physics.soc-ph