Searcharxiv⌕ Search

arXiv subjects

Noah L. Schroeder

Publications and source records attributed to Noah L. Schroeder.

3 recordsLinked to original sources

A Benchmark for LLM's Understanding of Middle School and High School Science Topics

Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were systematically evaluated using this benchmark, indicating that several smaller, locally deployable models achieved high accuracy across diverse science domains and question types. Our findings indicate that model size did not consistently predict performance, emphasizing the importance of intentional model selection for educational deployment. We then incorporated a human reviewer into the loop, reviewing the items generated by the LLMs for alignment with NGSS standards. The human review indicated that synthetically generated items were not in perfect alignment with the NGSS standards, indicating the benefits of human-in-the-loop item development, the need to explore the intersection of content and pedagogical knowledge, and the need to extend benchmarks to evaluate LLMs' capacity for interactive, evidence-based feedback in educational scenarios.

cs.AI↗

AI-Assisted Data Extraction for Systematic Reviews in Education

Systematic reviews are time-consuming endeavors that require knowledgeable human reviewers to screen studies for relevance and extract data following a specific coding scheme before any analysis or synthesis can occur. Large language models (LLMs) hold promise for substantially accelerating this process and reducing reviewer workload, yet their application within the context of systematic reviews in the field of education remains underexplored. We address this issue in two ways: through empirical studies and the iterative development of an open-source software tool. First, we conducted two empirical studies examining the efficacy of using LLMs for data extraction using data from a published review on pedagogical agents. We extracted a variety of data types from 112 studies and compared the results to data extracted by human coding. Results indicate that LLMs struggled with extracting data accurately and therefore are not ready to be used as primary data extraction tools without explicit human validation of the data extracted. These findings highlight the dire need for a human-in-the-loop (HIL) approach to AI-assisted data extraction. We then propose a HIL workflow and introduce and describe the development of a free, web-based, open-source tool designed to support user-friendly, human-validated data extraction with LLMs.

cs.HC↗

Interactive Evidence Maps for Visualizing and Understanding Systematic Reviews

Systematic reviews provide comprehensive syntheses of research fields. As a result, systematic reviews often emphasize synthesizing across the large bodies of literature rather than just describing the studies from which the conclusions were drawn. This risks an incomplete description of the sample - encouraging overgeneralization of the findings, obscuring connections between existing work, or overshadowing gaps in the literature. To address this challenge, we introduce interactive evidence maps; an accessible visualization tool that enables researchers to explore, filter, and analyze review data dynamically. Our approach leverages large language models to extract topic models that structure heterogeneous review data into an interactive, explorable knowledge map that supports deeper inspection beyond static tables and figures. We demonstrate the usefulness of interactive evidence maps using data from a published scoping review of pedagogical agents in K-12 education, and compare the results of the evidence map to those reported in the scoping review. Results show that interactive evidence maps complement traditional syntheses by enhancing transparency, supporting exploratory analysis, and revealing patterns and gaps that may not be easy to detect through narrative summaries alone.

cs.DL↗