arXiv · 2608.12343
Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models
Abstract
Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores mask distinct competency profiles: rankings are unstable across task types, source modalities, and geographical scopes, with a consistent penalty for Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source decontextualization, when models reason from source content rather than treating it as an object of analysis, and temporal disorientation, when responses are historically misplaced. This study introduces the first LLM history benchmark grounded in the Polish national curriculum.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adrian Trzoss, Kacper Dudzic, Wiktor Werner, Marcin Moskalewicz. 2026-06-03. Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models. https://arxiv.org/abs/2608.12343
Cite the original work for its findings. Save a collection to share your selection of sources.