SearcharxivSearch

arXiv subjects

Daniel Racek

Publications and source records attributed to Daniel Racek.

4 recordsLinked to original sources

Assessing Reporting Delays in ACLED Conflict Event Data

Timely and accurate conflict event data are essential for real-time monitoring, forecasting, and policy response. Yet near-real-time conflict datasets such as the Armed Conflict Location \& Event Data Project (ACLED) are subject to reporting delays, that is, delays between event occurrence and first inclusion in the database. Such delays can introduce bias in short-term analyses and forecasts. This study provides a statistical analysis of reporting delays for African events recorded in ACLED's weekly releases from June 30, 2024, to June 1, 2025. Treating delay as a discrete time duration, we estimate grouped proportional hazards models with additive-linear and smooth terms incorporating event-level, spatial, and country-level covariates. Our results show that more than half of events are reported within two weeks, but delays vary systematically by event type, fatalities, geographic location, and political regime. Higher-fatality events are reported more quickly, while events in more restrictive political and informational environments tend to be reported more slowly. We also find substantial between-country heterogeneity, and country-specific analyses indicate that event-level effects differ across contexts. These findings show that reporting delays are structured rather than random and that real-time conflict analysis must account for them. More broadly, they provide an empirical foundation for developing nowcasting approaches to correct short-term underreporting in conflict event data.

stat.AP

Forests of Uncertaint(r)ees: Using tree-based ensembles to estimate probability distributions of future conflict

Predictions of fatalities from violent conflict on the PRIO-GRID-month (pgm) level are characterized by high levels of uncertainty, limiting their usefulness in practical applications. We discuss the two main sources of uncertainty for this prediction task, the nature of violent conflict and data limitations, embedding conflict prediction in the wider literature on uncertainty quantification in machine learning. Based on this, we develop a strategy to quantify uncertainty in conflict forecasting, shifting from traditional point predictions to full predictive distributions. Our approach combines multiple tree-based classifiers and distributional regressors in a custom AutoML setup, estimating distributions for each pgm individually. We also test the integration of regional models in spatial ensembles as a potential avenue to reduce uncertainty by lowering data requirements and accounting for systematic differences between conflict contexts. The models are able to consistently outperform a suite of benchmarks derived from conflict history in predictions up to one year in advance. Marginal differences in model-wide metrics emphasize the need to understand their behavior for a given prediction problem, in this case characterized by extremely high zero-inflatedness. Adressing this, we compliment our evaluation with a simulation experiment, which demonstrates that our models reflect meaningful performance improvements, which can be traced back to conflict-affected regions. Lastly, we show that the integration of regional models does not decrease performance, opening avenues to integrate additional data sources in the future.

stat.AP

The Politics of Language Choice: How the Russian-Ukrainian War Influences Ukrainians' Language Use on Twitter

The use of language is innately political and often a vehicle of cultural identity as well as the basis for nation building. Here, we examine language choice and tweeting activity of Ukrainian citizens based on more than 4 million geo-tagged tweets from over 62,000 users before and during the Russian-Ukrainian War, from January 2020 to October 2022. Using statistical models, we disentangle sample effects, arising from the in- and outflux of users on Twitter, from behavioural effects, arising from behavioural changes of the users. We observe a steady shift from the Russian language towards the Ukrainian language already before the war, which drastically speeds up with its outbreak. We attribute these shifts in large part to users' behavioural changes. Notably, we find that more than half of the Russian-tweeting users shift towards Ukrainian as a result of the war.

cs.CY

Factorized Structured Regression for Large-Scale Varying Coefficient Models

Recommender Systems (RS) pervade many aspects of our everyday digital life. Proposed to work at scale, state-of-the-art RS allow the modeling of thousands of interactions and facilitate highly individualized recommendations. Conceptually, many RS can be viewed as instances of statistical regression models that incorporate complex feature effects and potentially non-Gaussian outcomes. Such structured regression models, including time-aware varying coefficients models, are, however, limited in their applicability to categorical effects and inclusion of a large number of interactions. Here, we propose Factorized Structured Regression (FaStR) for scalable varying coefficient models. FaStR overcomes limitations of general regression models for large-scale data by combining structured additive regression and factorization approaches in a neural network-based model implementation. This fusion provides a scalable framework for the estimation of statistical models in previously infeasible data settings. Empirical results confirm that the estimation of varying coefficients of our approach is on par with state-of-the-art regression techniques, while scaling notably better and also being competitive with other time-aware RS in terms of prediction performance. We illustrate FaStR's performance and interpretability on a large-scale behavioral study with smartphone user data.

stat.ML