SearcharxivSearch

arXiv subjects

Amir Bralin

Publications and source records attributed to Amir Bralin.

6 recordsLinked to original sources

Probing AI-generated physics solutions and preparing students to critique them

This study examines Artificial Intelligence (AI)-generated physics solutions from two connected perspectives: how prompt design shapes these solutions and how students can be prepared to critique them. Using a rotational-mechanics problem, we adapted a problem-classification framework to examine prompt variations, evaluating OpenAI's o4-mini responses with the Minnesota Assessment of Problem Solving (MAPS) rubric. Well-specified prompts improved solution completeness; underspecified and multimodal prompts exposed weaknesses in physics reasoning and correctness. In the student-evaluation phase, 24 introductory physics lab groups evaluated an o4-mini solution to this problem after either independently solving a related problem or critiquing its AI-generated solution with MAPS-based reflection questions. Problem-solving-only groups exhibited uncritical or misconception-based critiques; MAPS-guided groups identified more expert-aligned issues, including skipped numerical procedures and undefined notation. Together, our findings contribute to physics education research by showing how AI-generated solutions can ground both model-reasoning benchmarks and improved student critique of that reasoning through MAPS-based reflection.

physics.ed-ph

Assessing AI in Introductory Physics Problem Solving

Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.

physics.ed-ph

Towards Just-in-Time Adaptive Feedback: Enhancing Student Learning via Knowledge-Grounded LLM

Educational interventions are effective tools for enhancing student learning. While Large Language Models (LLMs) allow for generating adaptive feedback at scale, current studies lack clear methodologies for providing Just-in-Time (JiT) feedback in authentic instructional settings. In this paper, we present a framework that provides adaptive feedback by grounding LLMs with domain-specific expert knowledge. Our approach collects written reasoning logic (strategy essays) from students, analyzes potential error types based on the content of that reasoning, and delivers non-intrusive feedback designed to clarify missing or incorrect concepts. We deploy this framework in a large-scale university course (N > 1000), where it improved student performance by over 80% compared to previous semesters. Lastly, we validate the framework's pedagogical utility by analyzing the learning trajectories; we demonstrate how iterative conversations with LLM facilitate shifting one's misconception to correct understanding.

cs.CL

Using Large Language Models to Analyze Engagement in Computational Thinking via Computational Physics Essays

As computational thinking (CT) becomes increasingly important to physics education, the need for authentic, project-based assessments has grown. While open-ended multimodal assignments, such as Computational Physics Essays (CPEs), help capture student reasoning and encourage active learning, they introduce a significant evaluation bottleneck. Manually grading these complex notebooks across a complex taxonomy of computational practices is resource-intensive and limits scalability in large-enrollment courses. In this study, we investigated the viability of using a multimodal Large Language Model (LLM) to automate the evaluation of 100 student-generated CPEs. Using a human-coded baseline, we systematically evaluated the model's capacity to detect student engagement across 20 distinct CT sub-practices and a holistic overall quality score. The results showed that the LLM performs very well on clearly defined tasks, achieving an 84% exact agreement with human raters on the binary sub-practices. However, more subjective constructs proved challenging, with the model reaching only a 71% agreement for the holistic quality analysis. Our findings demonstrated that while LLMs can reliably automate the detection of specific computational practices, subjective evaluation remains a hurdle.

physics.ed-ph

AI Reasoning Models for Problem Solving in Physics

Reasoning models are the new generation of Large Language Models (LLMs) capable of complex problem solving. Their reliability in solving introductory physics problems was tested by evaluating a sample of n = 5 solutions generated by one such model -- OpenAI's o3-mini -- per each problem from 20 chapters of a standard undergraduate textbook. In total, N = 408 problems were given to the model and N x n = 2,040 generated solutions examined. The model successfully solved 94% of the problems posed, excelling at the beginning topics in mechanics but struggling with the later ones such as waves and thermodynamics.

physics.ed-ph

Investigating the Design-Science Connection in a multi-week Engineering Design (ED)-based introductory physics laboratory task

Reform documents advocate for innovative pedagogical strategies to enhance student learning. A key innovation is the integration of science and engineering practices through Engineering Design (ED)-based physics laboratory tasks, where students tackle engineering design problems by applying physics principles. While this approach has its benefits, research shows that students do not always effectively apply scientific concepts, but instead rely on trial-and-error approaches, and end up 'gadgeteering' their way to a solution. This leads to what is commonly referred to as the "design-science gap" -- that students do not always consciously apply science concepts while solving a design problem. However, as obvious as the notion of a `gap' may appear, there seems to exist no consensus on the definitions of `design' and `science', further complicating the understanding of this `gap'. This qualitative study addresses the notion of the design-science gap by examining student-groups' discussions and written lab reports from a multi-week ED-based undergraduate introductory physics laboratory task. Building on our earlier studies, we developed and employed a nuanced, multi-layered coding scheme inspired by the Gioia Framework to characterize `design thinking' and `science thinking'. We discuss how student-groups engage in various aspects of design and how they apply concepts physics principles to solve the problem. In the process, we demonstrate the interconnectedness of students' design thinking and science thinking. We advocate for the usage of the term "design-science connection" as opposed to "design-science gap" to deepen both design and scientific thinking. Our findings offer valuable insights for educators in design-based science education.

physics.ed-ph