SearcharxivSearch

arXiv subjects

N. Sanjay Rebello

Publications and source records attributed to N. Sanjay Rebello.

At least 19 recordsLinked to original sources

Exploring Students' Perceptions of Using Generative AI-Assisted Problem Posing

Problem posing, a pedagogical practice that asks students to generate novel problems or meaningful variations to problems encountered, supports transfer of learning and strengthens problem-solving skills in physics. However, generating physics problems can be challenging, particularly for novice learners. This study investigates students' perceptions of an approach to facilitate their use of Generative AI in ways that maximize their benefits and limit risks. Students were introduced to Generative AI-assisted problem posing as a self-study technique. This study utilizes a phenomenological approach to investigate students' perceptions on how training shaped AI interactions, attitudes towards problem posing with Generative AI, and how students view its incorporation into personal study practices. Results of this study suggest that students perceived a positive change in their interactions with Generative AI after receiving training on prompt engineering techniques. This study also reveals that students hold generally positive views towards the problem-posing technique, with a smaller subset of students showing hesitations towards using Generative AI. These results lay the foundation to introduce and employ training methods for Generative AI more widely and to continue to incorporate Generative AI into structured study techniques, like problem posing.

physics.ed-ph

Probing AI-generated physics solutions and preparing students to critique them

This study examines Artificial Intelligence (AI)-generated physics solutions from two connected perspectives: how prompt design shapes these solutions and how students can be prepared to critique them. Using a rotational-mechanics problem, we adapted a problem-classification framework to examine prompt variations, evaluating OpenAI's o4-mini responses with the Minnesota Assessment of Problem Solving (MAPS) rubric. Well-specified prompts improved solution completeness; underspecified and multimodal prompts exposed weaknesses in physics reasoning and correctness. In the student-evaluation phase, 24 introductory physics lab groups evaluated an o4-mini solution to this problem after either independently solving a related problem or critiquing its AI-generated solution with MAPS-based reflection questions. Problem-solving-only groups exhibited uncritical or misconception-based critiques; MAPS-guided groups identified more expert-aligned issues, including skipped numerical procedures and undefined notation. Together, our findings contribute to physics education research by showing how AI-generated solutions can ground both model-reasoning benchmarks and improved student critique of that reasoning through MAPS-based reflection.

physics.ed-ph

A bottom-up taxonomy of student discourse with a Socratic AI physics tutor

Large language model (LLM) tutors are being deployed in introductory physics courses at a scale that produces transcript corpora far larger than traditional qualitative coding can absorb. A central question for physics education research (PER) is empirical and prior to any claim about effectiveness: what do students actually say to these tutors? We address this question for one Socratic AI tutor deployed in an introductory calculus-based mechanics course by building a bottom-up taxonomy of student discourse. Each student turn is assigned an emergent free-text label by an LLM coder using the surrounding conversational context; near-paraphrase labels are then consolidated into a smaller set of discourse categories using a similarity-based grouping procedure. The procedure is validated against a stratified human-coded sample. The resulting taxonomy of 357 categories is strikingly concentrated: the top 25 categories cover roughly half of all student turns, and two thematic bands: equation-handling and meta-procedural requests together dominate the head of the distribution. The substantive contribution is the taxonomy itself: a description of the discourse PER researchers can expect to encounter when students work with an AI tutor of this design, including a striking prevalence of meta-procedural turns in which students cede strategic control to the tutor

physics.ed-ph

Using LLMs to Detect Growth in Computational Thinking in Introductory Physics

As computation becomes more central to physics education, creating scalable methods to assess authentic computational thinking (CT) in students remains a critical challenge. While student-written responses capture nuanced reasoning, they are difficult to evaluate at scale. In this study, we investigated the use of Large Language Models (LLMs) to analyze students' written explanations of computational physics problems on a pre- and post- semester survey. By first establishing a human-coded baseline, grounded in CT literature, we identified significant growth in Data Practices and Computational Problem-Solving Practices. When given the same responses, an LLM successfully mirrored the human evaluations and scaled up the detection of these key trends across a large dataset. Notably, both human raters and the LLM struggled to reliably evaluate more complex constructs such as Systems Thinking. Overall, this study demonstrates that LLMs offer a viable method to scale the evaluation of students' CT in large-enrollment physics courses

physics.ed-ph

Transition Matrix Analysis Analyzing Students Use of Cognitive Resources in Physics

Conceptual surveys of multiple choice format have been developed to test the effect of pedagogical interventions on students understanding of physics knowledge. Predominantly, they are administered in pre and posttest settings and analyzed to obtain a performance gain. However, focusing on the correct answers to each question alone ignores the incorrect options which could inform us the stability and coherency of the knowledge structure of students. According to the resource model framework, each specific answer from students could reflect a specific cognitive resource being activated and implemented in the context at hand. Consequently, conceptual surveys could demonstrate how instruction could affect the activation of different cognitive resources of students in a variety of contexts when administered before and after instruction. Guided by the resources framework, we propose a transition matrix analysis to analyze data collected through conceptual surveys to investigate how instruction affects the consistency of cognitive resources activation by students. To provide proof of concept, we demonstrated how to utilize this method by analyzing students responses to a subset of questions from the DIRECT survey in a classroom study.

physics.ed-ph

Assessing AI in Introductory Physics Problem Solving

Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.

physics.ed-ph

Using Large Language Models to Analyze Engagement in Computational Thinking via Computational Physics Essays

As computational thinking (CT) becomes increasingly important to physics education, the need for authentic, project-based assessments has grown. While open-ended multimodal assignments, such as Computational Physics Essays (CPEs), help capture student reasoning and encourage active learning, they introduce a significant evaluation bottleneck. Manually grading these complex notebooks across a complex taxonomy of computational practices is resource-intensive and limits scalability in large-enrollment courses. In this study, we investigated the viability of using a multimodal Large Language Model (LLM) to automate the evaluation of 100 student-generated CPEs. Using a human-coded baseline, we systematically evaluated the model's capacity to detect student engagement across 20 distinct CT sub-practices and a holistic overall quality score. The results showed that the LLM performs very well on clearly defined tasks, achieving an 84% exact agreement with human raters on the binary sub-practices. However, more subjective constructs proved challenging, with the model reaching only a 71% agreement for the holistic quality analysis. Our findings demonstrated that while LLMs can reliably automate the detection of specific computational practices, subjective evaluation remains a hurdle.

physics.ed-ph

Game-Based vs. Simulation-Based Instruction: exploring the sequencing effect on elementary pre-service teachers' understanding of the photoelectric effect

The use of digital tools and multiple representations like educational games and interactive simulations is of great importance to physics education. This study investigates the sequencing effects of an educational video game 'Photon Jump' and the PhET Photoelectric Effect simulation on pre-service teachers' understanding of the photoelectric effect. Using a counterbalanced quasi-experimental crossover design, pre-service teachers enrolled in a physics course (N = 83) experienced both interventions in opposite orders. Conceptual understanding was measured across three standardized assessments, complemented by open-ended reflection questions on participants' preferences and willingness to use both tools for future learning. The simulation-first sequence yielded a greater significant improvement in performance p = 0.001 as compared with game-first sequence p = 0.06. participants' preferences for using the game as opposed to the simulation were dependent on the sequence that they were randomly assigned to. Findings underscore the complementary strengths of game-based and simulation-based instruction, highlighting the importance of choosing the right sequence when using multiple representations in teaching abstract physics phenomena to pre-service teachers.

physics.ed-ph

Two Paths to Learning Physics: How Games and Simulations Shape Physics Learning Among Physics and Engineering Students

The use of digital tools like educational video games and interactive simulations is of great importance to physics education. This study investigates the sequencing effects of an educational game, 'Photon Jump' and PhET Photoelectric Effect simulation on students' perception of such tools for learning about the photoelectric effect. Using a counterbalanced quasi-experimental crossover design, a total of 55 physics and engineering students from a calculus-based physics course were divided into two groups of comparable sizes and administered the game and the simulation in reverse order followed by 9 open ended reflection questions. The results show a clear preference for the game-based activity over the simulation. Additionally, gaming frequency showed no correlation with willingness to use similar tools, suggesting broad accessibility. Thematic analysis revealed that intuitive and explorative learning through the game was reinforced by the analytical aspect of the simulation.

physics.ed-ph

Relating visual attention and learning in an online instructional physics module

Learning using Computer-Assisted Instruction (CAI) demands a high level of attention given the tendency to be distracted and mind-wander. How does the online STEM instructor know when learners are having attentional problems and the extent to which these problems affect learning? In the present study, the visual attentional and cognitive state of physics graduate students were probed while they went through a multimedia instructional module to refresh their knowledge of Newton's II Law. Data from an eye tracker, webcam, egocentric glasses, screen recording, and mouse and keyboard events were integrated to record learners' attention overt attention to the learning environment (+/-) and thinking about learning content (+/-) to analyze students' attention spans during learning from this module. On average, learners were found to be on-task and on-screen for a vast majority of time, with evidence of mind wandering. The learning module improved the participants efficiency with which they answered the questions correctly on a post-test relative to the pre-test. Further, there is a positive albeit statistically non-significant correlation between the improvement from pre- to post-test efficiency and the time spent on-screen and on-task during the module.

physics.ed-ph

Cognitive Load and Situational Interest in Physics Laboratories: A Comparative Study Across Three Instructional Modalities

Understanding how an instructional approach shapes student's cognitive resources and engagement is central to improving undergraduate physics education especially for novice learners. This study examines how three instructional modalities (Inquiry-based, Design-based, and Game-based learning) affect cognitive load and situational interest in physics laboratories for non-STEM majors. Guided by the revised Cognitive Load Theory framework, two experiments were conducted across two physics domains: mechanics and electrical circuits. In each experiment, students completed three laboratory sessions, one in each instructional modality, followed by surveys measuring cognitive load and situational interest. One-way ANOVA analyses revealed significant differences across the three modalities in both experiments. Game-based laboratories consistently yielded the lowest cognitive load and the highest situational interest, while inquiry-based and design-based labs imposed higher cognitive demands, with their relative effects varying by domain. Overall, situational interest exhibited an inverse relationship with cognitive load, suggesting that reduced cognitive demands support greater engagement. These findings emphasize the value of strategically selecting and combining instructional modalities to balance cognitive load and foster meaningful engagement in physics laboratories for novice learners.

physics.ed-ph

Making Evidence Actionable in Adaptive Learning Closing the Diagnostic Pedagogical Loop

Adaptive learning often diagnoses precisely yet intervenes weakly, producing help that is mistimed or misaligned. This study presents evidence supporting an instructor-governed feedback loop that converts concept-level assessment evidence into vetted microinterventions. The adaptive learning algorithm includes three safeguards: adequacy as a hard guarantee of gap closure, attention as a budgeted limit for time and redundancy, and diversity as protection against overfitting to a single resource. We formulate intervention assignment as a binary integer program with constraints for coverage, time, difficulty windows derived from ability estimates, prerequisites encoded by a concept matrix, and anti-redundancy with diversity. Greedy selection serves low-richness and tight-latency settings, gradient-based relaxation serves rich repositories, and a hybrid switches along a richness-latency frontier. In simulation and in an introductory physics deployment with 1204 students, both solvers achieved full skill coverage for nearly all learners within bounded watch time. The gradient-based method reduced redundant coverage by about 12 percentage points relative to greedy and produced more consistent difficulty alignment, while greedy delivered comparable adequacy at lower computational cost in resource-scarce environments. Slack variables localized missing content and guided targeted curation, sustaining sufficiency across student subgroups. The result is a tractable and auditable controller that closes the diagnostic pedagogical loop and enables equitable, load-aware personalization at the classroom scale.

cs.CE

Making Evidence Actionable in Adaptive Learning

Adaptive learning often diagnoses precisely yet intervenes weakly, yielding help that is mistimed or misaligned. This study presents evidence supporting an instructor-governed feedback loop that converts concept-level assessment evidence into vetted micro-interventions. The adaptive learning algorithm contains three safeguards: adequacy as a hard guarantee of gap closure, attention as a budgeted constraint for time and redundancy, and diversity as protection against overfitting to a single resource. We formalize intervention assignment as a binary integer program with constraints for coverage, time, difficulty windows informed by ability estimates, prerequisites encoded by a concept matrix, and anti-redundancy enforced through diversity. Greedy selection serves low-richness and tight-latency regimes, gradient-based relaxation serves rich repositories, and a hybrid method transitions along a richness-latency frontier. In simulation and in an introductory physics deployment with one thousand two hundred four students, both solvers achieved full skill coverage for essentially all learners within bounded watch time. The gradient-based method reduced redundant coverage by approximately twelve percentage points relative to greedy and harmonized difficulty across slates, while greedy delivered comparable adequacy with lower computational cost in scarce settings. Slack variables localized missing content and supported targeted curation, sustaining sufficiency across subgroups. The result is a tractable and auditable controller that closes the diagnostic-pedagogical loop and delivers equitable, load-aware personalization at classroom scale.

cs.AI

Assessing Physics Students' Scientific Argumentation using Natural Language Processing

Scientific argumentation is a core science and engineering practice and a necessary 21st Century workforce skill. Due to the nature of large enrollment classes, it is difficult to individually assess students and provide feedback on their scientific argumentation. The recent developments in Natural Language Processing (NLP) and Machine Learning (ML) provide new opportunities to analyze large collections of student writing efficiently. In this study, we investigate how undergraduate students' scientific argumentation evolves across four semesters of an introductory calculus-based physics course as increasingly structured argumentation scaffolds were introduced. We investigate the use of NLP and ML, specifically topic modeling, to analyze student scientific argumentation across those semesters. We report on the emergent themes present in each semester. Our findings show a clear shift in the thematic focus of student arguments corresponding to the level of scaffolding provided. In semesters with minimal scaffolding, students' arguments emphasized procedural and surface-level features, while semesters with explicit scaffolds exhibited greater concentration around physics-principle-based themes. These results suggest that structured scaffolding supports students in constructing more conceptually grounded scientific arguments and highlights the potential of NLP and ML as scalable approaches for evaluate broad trends in students' scientific argumentation.

physics.ed-ph

Examining Student and AI Generated Personalized Analogies in Introductory Physics

Comparing abstract concepts (such as electric circuits) with familiar ideas (plumbing systems) through analogies is central to practice and communication of physics. Contemporary research highlights self-generated analogies to better facilitate students' learning than the taught ones. "Spontaneous" and "self-generated" analogies represent the two ways through which students construct personalized analogies. However, facilitating them, particularly in large enrollment courses remains a challenge, and recent developments in generative artificial intelligence (AI) promise potential to address this issue. In this qualitative study, we analyze around 800 student responses in exploring the extent to which students spontaneously leverage analogies while explaining Morse potential curve in a language suitable for second graders and self-generate analogies in their preferred everyday contexts. We also compare the student-generated spontaneous analogies with AI-generated ones prompted by students. Lastly, we explore the themes associated with students' perceived ease and difficulty in generating analogies across both cases. Results highlight that unlike AI responses, student-generated spontaneous explanations seldom employ analogies. However, when explicitly asked to explain the behavior of the curve in terms of their everyday contexts, students employ diverse analogical contexts. A combination of disciplinary knowledge, agency to generate customized explanations, and personal attributes tend to influence students' perceived ease in generating explanations across the two cases. Implications of these results on the potential of AI to facilitate students' personalized analogical reasoning, and the role of analogies in making students notice gaps in their understanding are discussed.

physics.ed-ph

Help or Hype? Students' Engagement and Perception of Using AI to Solve Physics Problems

With the rise of large language models such as ChatGPT, interest has grown in understanding how these tools influence learning in STEM education, including physics. This study explores how students use ChatGPT during a physics problem-solving task embedded in a formal assessment. We analyzed patterns of AI usage and their relationship to student performance. Findings indicate that students who engaged with ChatGPT generally performed better than those who did not. Particularly, students who provided more complete and contextual prompts experienced greater benefits. Further, students who demonstrated overall positive gains collectively asked more conceptual questions than those who exhibited overall negative gains. However, the presence of incorrect AI-generated responses also underscores the importance of critically evaluating AI output. These results suggest that while AI can be a valuable aid in problem solving, its effectiveness depends significantly on how students use it, reinforcing the need to incorporate structured AI-literacy into STEM education.

physics.ed-ph

Feedback That Clicks: Introductory Physics Students' Valued Features in AI Feedback Generated From Self-Crafted and Engineered Prompts

Since the advent of GPT-3.5 in 2022, Generative Artificial Intelligence (AI) has shown tremendous potential in STEM education, particularly in providing real-time, customized feedback to students in large-enrollment courses. A crucial skill that mediates effective use of AI is the systematic structuring of natural language instructions to AI models, commonly referred to as prompt engineering. This study has three objectives: (i) to investigate the sophistication of student-generated prompts when seeking feedback from AI on their arguments, (ii) to examine the features that students value in AI-generated feedback, and (iii) to analyze trends in student preferences for feedback generated from self-crafted prompts versus prompts incorporating prompt engineering techniques and principles of effective feedback. Results indicate that student-generated prompts typically reflect only a subset of foundational prompt engineering techniques. Despite this lack of sophistication, such as incomplete descriptions of task context, AI responses demonstrated contextual intuitiveness by accurately inferring context from the overall content of the prompt. We also identified 12 distinct features that students attribute the usefulness of AI-generated feedback, spanning four broader themes: Evaluation, Content, Presentation, and Depth. Finally, results show that students overwhelmingly prefer feedback generated from structured prompts, particularly those combining prompt engineering techniques with principles of effective feedback. Implications of these results such as integrating the principles of effective feedback in design and delivery of feedback through AI systems, and incorporating prompt engineering in introductory physics courses are discussed.

physics.ed-ph

Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis

Qualitative analysis is typically limited to small datasets because it is time-intensive. Moreover, a second human rater is required to ensure reliable findings. Artificial intelligence tools may replace human raters if we demonstrate high reliability compared to human ratings. We investigated the inter-rater reliability of state-of-the-art Large Language Models (LLMs), ChatGPT-4o and ChatGPT-4.5-preview, in rating audio transcripts coded manually. We explored prompts and hyperparameters to optimize model performance. The participants were 14 undergraduate student groups from a university in the midwestern United States who discussed problem-solving strategies for a project. We prompted an LLM to replicate manual coding, and calculated Cohen's Kappa for inter-rater reliability. After optimizing model hyperparameters and prompts, the results showed substantial agreement ($κ>0.6$) for three themes and moderate agreement on one. Our findings demonstrate the potential of GPT-4o and GPT-4.5 for efficient, scalable qualitative analysis in physics education and identify their limitations in rating domain-general constructs.

physics.ed-ph