SearcharxivSearch

arXiv subjects

Sean Savage

Publications and source records attributed to Sean Savage.

4 recordsLinked to original sources

Probing AI-generated physics solutions and preparing students to critique them

This study examines Artificial Intelligence (AI)-generated physics solutions from two connected perspectives: how prompt design shapes these solutions and how students can be prepared to critique them. Using a rotational-mechanics problem, we adapted a problem-classification framework to examine prompt variations, evaluating OpenAI's o4-mini responses with the Minnesota Assessment of Problem Solving (MAPS) rubric. Well-specified prompts improved solution completeness; underspecified and multimodal prompts exposed weaknesses in physics reasoning and correctness. In the student-evaluation phase, 24 introductory physics lab groups evaluated an o4-mini solution to this problem after either independently solving a related problem or critiquing its AI-generated solution with MAPS-based reflection questions. Problem-solving-only groups exhibited uncritical or misconception-based critiques; MAPS-guided groups identified more expert-aligned issues, including skipped numerical procedures and undefined notation. Together, our findings contribute to physics education research by showing how AI-generated solutions can ground both model-reasoning benchmarks and improved student critique of that reasoning through MAPS-based reflection.

physics.ed-ph

Using LLMs to Detect Growth in Computational Thinking in Introductory Physics

As computation becomes more central to physics education, creating scalable methods to assess authentic computational thinking (CT) in students remains a critical challenge. While student-written responses capture nuanced reasoning, they are difficult to evaluate at scale. In this study, we investigated the use of Large Language Models (LLMs) to analyze students' written explanations of computational physics problems on a pre- and post- semester survey. By first establishing a human-coded baseline, grounded in CT literature, we identified significant growth in Data Practices and Computational Problem-Solving Practices. When given the same responses, an LLM successfully mirrored the human evaluations and scaled up the detection of these key trends across a large dataset. Notably, both human raters and the LLM struggled to reliably evaluate more complex constructs such as Systems Thinking. Overall, this study demonstrates that LLMs offer a viable method to scale the evaluation of students' CT in large-enrollment physics courses

physics.ed-ph

Using Large Language Models to Analyze Engagement in Computational Thinking via Computational Physics Essays

As computational thinking (CT) becomes increasingly important to physics education, the need for authentic, project-based assessments has grown. While open-ended multimodal assignments, such as Computational Physics Essays (CPEs), help capture student reasoning and encourage active learning, they introduce a significant evaluation bottleneck. Manually grading these complex notebooks across a complex taxonomy of computational practices is resource-intensive and limits scalability in large-enrollment courses. In this study, we investigated the viability of using a multimodal Large Language Model (LLM) to automate the evaluation of 100 student-generated CPEs. Using a human-coded baseline, we systematically evaluated the model's capacity to detect student engagement across 20 distinct CT sub-practices and a holistic overall quality score. The results showed that the LLM performs very well on clearly defined tasks, achieving an 84% exact agreement with human raters on the binary sub-practices. However, more subjective constructs proved challenging, with the model reaching only a 71% agreement for the holistic quality analysis. Our findings demonstrated that while LLMs can reliably automate the detection of specific computational practices, subjective evaluation remains a hurdle.

physics.ed-ph

Using an LLM to Investigate Students' Explanations on Conceptual Physics Questions

Analyzing students' written solutions to physics questions is a major area in PER. However, gauging student understanding in college courses is bottlenecked by large class sizes, which limits assessments to a multiple-choice (MC) format for ease of grading. Although sufficient in quantifying scientifically correct conceptions, MC assessments do not uncover students' deeper ways of understanding physics. Large language models (LLMs) offer a promising approach for assessing students' written responses at scale. Our study used an LLM, validated by human graders, to classify students' written explanations to three questions on the Energy and Momentum Conceptual Survey as correct or incorrect, and organized students' incorrect explanations into emergent categories. We found that the LLM (GPT-4o) can fairly assess students' explanations, comparable to human graders (0-3% discrepancy). Furthermore, the categories of incorrect explanations were different from corresponding MC distractors, allowing for different and deeper conceptions to become accessible to educators.

physics.ed-ph