arXiv · 2505.05406
Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries
Abstract
News headlines and summaries shape how events are interpreted through selective emphasis and omission, a phenomenon commonly referred to as framing. Large language models are now routinely used to generate such content, yet existing evaluation frameworks largely overlook this dimension. We introduce Frame In, Frame Out (FIFO), the first large-scale benchmark for measuring framing presence in LLM-generated news summaries, grounded in the widely used XSum dataset. FIFO combines 15,499 jury-annotated examples with 320 expert-labeled instances ($\kappa = 0.61$) to validate and calibrate model-based annotations. Using FIFO, we analyze measured framing rates across 27 summarization models. We find that LLM-generated summaries often exhibit higher calibrated framing rates than human-written references, with substantial variation across topics and training regimes, including elevated rates in scientific and public health summaries. Our results establish framing as an underexplored and consequential dimension of summarization quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Valeria Pastorino, Nafise Sadat Moosavi. 2025-05-08. Frame In, Frame Out: Measuring Framing Bias in LLM-Generated News Summaries. https://arxiv.org/abs/2505.05406
Cite the original work for its findings. Save a collection to share your selection of sources.