arXiv · 2609.32020
A Benchmark for LLM's Understanding of Middle School and High School Science Topics
Abstract
Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs' capacity for interactive, evidence-based feedback in educational scenarios.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Noah L. Schroeder, Yessy Eka Ambarwati, Yuji Zhang, ChengXiang Zhai. 2026-09-25. A Benchmark for LLM's Understanding of Middle School and High School Science Topics. https://arxiv.org/abs/2609.32020
Cite the original work for its findings. Save a collection to share your selection of sources.