arXiv · 2110.04384
Evaluation of Summarization Systems across Gender, Age, and Race
Abstract
Summarization systems are ultimately evaluated by human annotators and raters. Usually, annotators and raters do not reflect the demographics of end users, but are recruited through student populations or crowdsourcing platforms with skewed demographics. For two different evaluation scenarios -- evaluation against gold summaries and system output ratings -- we show that summary evaluation is sensitive to protected attributes. This can severely bias system development and evaluation, leading us to build models that cater for some groups rather than others.
Explore related subjects
Keep this discovery
Anna Jørgensen, Anders Søgaard. 2021-10-08. Evaluation of Summarization Systems across Gender, Age, and Race. https://arxiv.org/abs/2110.04384
Cite the original work for its findings. Save a collection to share your selection of sources.