arXiv · 2609.22375
Initial Evaluation of Potential Bias in Coverage of Humans in Wikidata
Abstract
Introduction. Open collaborative knowledge graphs such as Wikidata increasingly ground agentic artificial intelligence, information retrieval, and language modeling systems, making systematic auditing of their demographic representation and overall equity a research imperative. Methods. Herein, we present an open-source auditing platform that ingests over 10 million statement bindings representing over 6 million humans on Wikidata via QLever, and evaluates representation of gender, sexual orientation, geography, birthplace urbanicity, ethnicity, multilingual coverage of labels, descriptions, and aliases, occupation, and select intersectional pairs of these entities. It does so by making use of Chi-square goodness-of-fit tests, 95% Wilson-score confidence intervals, and disparity ratios, in light of Rubin's missingness taxonomy. Results. Women accounted for 28.71% (CI +/-0.04) of all humans in Wikidata with a stated gender. 38.26% of humans had a citizenship statement, with Western Europe and North America (WENA) representing approximately 53% of such statements. Among 1.8 million birthplaces that could be classified, rural birthplaces were observed in 2.48% of cases (in comparison to 27.4% global baseline). Fewer than 1.2% of entities carried an ethnicity statement, and non-English Wikidata descriptions covered 18.2% of items. Discussion. Our findings reveal significant missingness across the evaluated axes. Ethnicity and sexual orientation were the most critically under-documented (missing statements) while rural birthplaces and non-WENA citizenship were the most underrepresented.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Clair Kronk. 2026-09-17. Initial Evaluation of Potential Bias in Coverage of Humans in Wikidata. https://arxiv.org/abs/2609.22375
Cite the original work for its findings. Save a collection to share your selection of sources.