arXiv · 2608.10210
FUBU-EPSTEIN: A Large-Scale Twitter Dataset on the Jeffrey Epstein Case and Its Global Public Discourse (2019-2023)
Abstract
The criminal case of Jeffrey Epstein has generated a complex, long-running global discourse on digital platforms, characterized by punctuated attention shocks, conspiracy theories, and blame attribution. To facilitate the computational study of these dynamics, we introduce the FUBU-EPSTEIN dataset, a large-scale, multi-dimensional research corpus of 54.38 million Twitter statuses authored by 7.11 million users and collected between August 2019 and April 2023. The source corpus was captured continuously in near-real time and enriched with a directed social contact graph of 37.03 million edge rows and annotations from Qwen2.5-7B-Instruct covering sentiment, conspiracy and misinformation stance, toxicity, moral emotion, and related dimensions. For public distribution, we created a textless, de-identified derivative that retains one row for every deduplicated status, categorical and numeric annotations, coarse temporal information, and 46.15 million internal status relationships. It excludes tweet text, original post and user identifiers, handles, profile fields, exact timestamps, and reverse mappings. The resulting release supports longitudinal content and diffusion analyses while reducing disclosure and platform-content redistribution risks. To request access to raw data for collaborative research under ethical and legal safeguards, contact fubu.dataset@gmail.com.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Kreil, Tristan Manfred Stöber, Daniel Thilo Schroeder. 2026-08-10. FUBU-EPSTEIN: A Large-Scale Twitter Dataset on the Jeffrey Epstein Case and Its Global Public Discourse (2019-2023). https://arxiv.org/abs/2608.10210
Cite the original work for its findings. Save a collection to share your selection of sources.