SearcharxivSearch

arXiv subjects

Bor Gregorcic

Publications and source records attributed to Bor Gregorcic.

10 recordsLinked to original sources

Report on the Scoping Workshop on AI in Science Education Research 2025

This report summarizes the outcomes of a two-day international scoping workshop on the role of artificial intelligence (AI) in science education research. As AI rapidly reshapes scientific practice, classroom learning, and research methods, the field faces both new opportunities and significant challenges. The report clarifies key AI concepts to reduce ambiguity and reviews evidence of how AI influences scientific work, teaching practices, and disciplinary learning. It identifies how AI intersects with major areas of science education research, including curriculum development, assessment, epistemic cognition, inclusion, and teacher professional development, highlighting cases where AI can support human reasoning and cases where it may introduce risks to equity or validity. The report also examines how AI is transforming methodological approaches across quantitative, qualitative, ethnographic, and design-based traditions, giving rise to hybrid forms of analysis that combine human and computational strengths. To guide responsible integration, a systems-thinking heuristic is introduced that helps researchers consider stakeholder needs, potential risks, and ethical constraints. The report concludes with actionable recommendations for training, infrastructure, and standards, along with guidance for funders, policymakers, professional organizations, and academic departments. The goal is to support principled and methodologically sound use of AI in science education research.

physics.ed-ph

Creating a customisable Socratic AI physics tutor

This paper explores role engineering as an effective paradigm for customizing Large Language Models (LLMs) into specialized AI tutors for physics education. We demonstrate this methodology by designing a Socratic physics problem-solving tutor using Google's Gemini Gems feature, defining its pedagogical behavior through a detailed 'script' that specifies its role and persona. We present two illustrative use cases: the first demonstrates the Gem's multimodal ability to analyze a student's hand-drawn force diagram and apply notational rules from a 'Knowledge' file; the second showcases its capacity to guide conceptual reasoning in electromagnetism using its pre-trained knowledge without using specific documents provided by the instructor. Our findings show that the 'role-engineered' Gem successfully facilitates a Socratic dialogue, in stark contrast to a standard Gemini model, which tends to immediately provide direct solutions. We conclude that role engineering is a pivotal and accessible method for educators to transform a general-purpose 'solution provider' into a reliable pedagogical tutor capable of engaging students in an active reflection process. This approach offers a powerful tool for both instructors and students, while also highlighting the importance of addressing the technology's inherent limitations, such as the potential for occasional inaccuracies.

physics.ed-ph

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Multimodal large language models (MLLMs) capable of processing both text and visual inputs are increasingly being explored for uses in physics education, such as tutoring, formative assessment, and grading. This study evaluates a range of publicly available MLLMs on a set of standardized, image-based physics research-based conceptual assessments (concept inventories). We benchmark 15 models from three major providers (Anthropic, Google, and OpenAI) across 102 physics items, focusing on two main questions: (1) How well do these models perform on conceptual physics tasks involving visual representations? and (2) What are the financial costs associated with their use? The results show high variability in both performance and cost. The performance of the tested models ranges from 81.5% to as low as 21%. We also found that expensive models do not always outperform cheaper ones and that, depending on the demands of the context, cheaper models may be sufficiently capable for some tasks. This is especially relevant in contexts where financial resources are limited or for large-scale educational implementation of MLLMs. By providing these analyses, our aim is to inform teachers, institutions, and other educational stakeholders so that they can make evidence-based decisions about the selection of models for use in AI-supported physics education.

physics.ed-ph

Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories. The inventories, sourced from the PhysPort website, cover classical physics topics such as mechanics, electromagnetism, optics, and thermodynamics, as well as relativity, quantum mechanics, astronomy, mathematics, and laboratory skills. Unlike previous text-only studies, we uploaded the inventories as images to reflect what a student would see on paper, thereby assessing the system's multimodal functionality. Our results indicate variation in performance across subjects, with laboratory skills standing out as the weakest. We also observe differences across languages, with English and European languages showing the strongest performance. Notably, the relative difficulty of an inventory item is largely independent of the language of the survey. When comparing AI results to existing literature on student performance, we find that the AI system outperforms average post-instruction undergraduate students in all subject categories except laboratory skills. Furthermore, the AI performs worse on items requiring visual interpretation of images than on those that are purely text-based. While our exploratory findings show GPT-4o's potential usefulness in physics education, they highlight the critical need for instructors to foster students' ability to critically evaluate AI outputs, adapt curricula thoughtfully in response to AI advancements, and address equity concerns associated with AI integration.

physics.ed-ph

Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment

Artificial intelligence-based chatbots are increasingly influencing physics education due to their ability to interpret and respond to textual and visual inputs. This study evaluates the performance of two large multimodal model-based chatbots, ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a conceptual physics inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. Quantitative analysis shows that ChatGPT-4o outperforms both ChatGPT-4 and a large sample of university students, and demonstrates improvements in ChatGPT-4o's vision interpretation ability over its predecessor ChatGPT-4. However, qualitative analysis of ChatGPT-4o's responses reveals persistent challenges. We identified three types of difficulties in the chatbot's responses to tasks on BEMA: (1) difficulties with visual interpretation, (2) difficulties in providing correct physics laws or rules, and (3) difficulties with spatial coordination and application of physics representations. Spatial reasoning tasks, particularly those requiring the use of the right-hand rule, proved especially problematic. These findings highlight that the most broadly used large multimodal model-based chatbot, ChatGPT-4o, still exhibits significant difficulties in engaging with physics tasks involving visual representations. While the chatbot shows potential for educational applications, including personalized tutoring and accessibility support for students who are blind or have low vision, its limitations necessitate caution. On the other hand, our findings can also be leveraged to design assessments that are difficult for chatbots to solve.

physics.ed-ph

Evaluating vision-capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based models

This study investigates the performance of eight large multimodal model (LMM)-based chatbots on the Test of Understanding Graphs in Kinematics (TUG-K), a research-based concept inventory. Graphs are a widely used representation in STEM and medical fields, making them a relevant topic for exploring LMM-based chatbots' visual interpretation abilities. We evaluated both freely available chatbots (Gemini 1.0 Pro, Claude 3 Sonnet, Microsoft Copilot, and ChatGPT-4o) and subscription-based ones (Gemini 1.0 Ultra, Gemini 1.5 Pro API, Claude 3 Opus, and ChatGPT-4). We found that OpenAI's chatbots outperform all the others, with ChatGPT-4o showing the overall best performance. Contrary to expectations, we found no notable differences in the overall performance between freely available and subscription-based versions of Gemini and Claude 3 chatbots, with the exception of Gemini 1.5 Pro, available via API. In addition, we found that tasks relying more heavily on linguistic input were generally easier for chatbots than those requiring visual interpretation. The study provides a basis for considerations of LMM-based chatbot applications in STEM and medical education, and suggests directions for future research.

physics.ed-ph

ChatGPT as a tool for honing teachers' Socratic dialogue skills

In this proof-of-concept paper, we propose a specific kind of pedagogical use of ChatGPT - to help teachers practice their Socratic dialogue skills. We follow up on the previously published paper "ChatGPT and the frustrated Socrates" by re-examining ChatGPT's ability to engage in Socratic dialogue in the role of a physics student. While in late 2022 its ability to engage in such dialogue was poor, we see significant advancements in the chatbot's ability to respond to leading questions asked by a human teacher. We suggest that ChatGPT now has the potential to be used in teacher training to help pre- or in-service physics teachers hone their Socratic dialogue skills. In the paper and its supplemental material, we provide illustrative examples of Socratic dialogues with ChatGPT and present a report on a pilot activity involving pre-service physics and mathematics teachers conversing with it in a Socratic fashion.

physics.ed-ph

How understanding large language models can inform the use of ChatGPT in physics education

The paper aims to fulfil three main functions: (1) to serve as an introduction for the physics education community to the functioning of Large Language Models (LLMs), (2) to present a series of illustrative examples demonstrating how prompt-engineering techniques can impact LLMs performance on conceptual physics tasks and (3) to discuss potential implications of the understanding of LLMs and prompt engineering for physics teaching and learning. We first summarise existing research on the performance of a popular LLM-based chatbot (ChatGPT) on physics tasks. We then give a basic account of how LLMs work, illustrate essential features of their functioning, and discuss their strengths and limitations. Equipped with this knowledge, we discuss some challenges with generating useful output with ChatGPT-4 in the context of introductory physics, paying special attention to conceptual questions and problems. We then provide a condensed overview of relevant literature on prompt engineering and demonstrate through illustrative examples how selected prompt-engineering techniques can be employed to improve ChatGPT-4's output on conceptual introductory physics problems. Qualitatively studying these examples provides additional insights into ChatGPT's functioning and its utility in physics problem solving. Finally, we consider how insights from the paper can inform the use of LLMs in the teaching and learning of physics.

physics.ed-ph

Performance of ChatGPT on the Test of Understanding Graphs in Kinematics

The well-known artificial intelligence-based chatbot ChatGPT-4 has become able to process image data as input in October 2023. We investigated its performance on the Test of Understanding Graphs in Kinematics to inform the physics education community of the current potential of using ChatGPT in the education process, particularly on tasks that involve graphical interpretation. We found that ChatGPT, on average, performed similarly to students taking a high-school level physics course, but with important differences in the distribution of the correctness of its responses, as well as in terms of the displayed "reasoning" and "visual" abilities. While ChatGPT was very successful at proposing productive strategies for solving the tasks on the test and expressed correct "reasoning" in most of its responses, it had difficulties correctly "seeing" graphs. We suggest that, based on its performance, caution and a critical approach are needed if one intends to use it in the role of a tutor, a model of a student, or a tool for assisting vision-impaired persons in the context of kinematics graphs.

physics.ed-ph

Development of Habits Through Apprenticeship in a Community: Conceptual Model of Physics Teacher Preparation

We propose the DHAC conceptual model of teacher preparation (Development of Habits through Apprenticeship in a Community), which builds on the contemporary literature on teacher preparation in physics and other disciplines and strives to provide a better understanding of the process of teacher formation. Extant literature on teacher preparation suggests that pre-service teachers learn best when they are immersed in a community that allows them to develop dispositions, knowledge, and practical skills and share with the community a strong vision of what good teaching entails. However, despite having developed the requisite dispositions, knowledge, and skills in pursuing the shared vision of good teaching, the professional demands on a teacher's time are so great out of and so complex during class time that if every decision requires multiple considerations and deliberations with oneself, the productive decisions might not materialize. We therefore argue that the missing link between intentional decision-making and actual teaching practice are the teacher's habits (spontaneous responses to situational cues). Teachers unavoidably develop habits with practical experience and under the influence of knowledge and belief structures that in many ways condition the responses of teachers in their practical work. To steer new teachers away from developing unproductive habits directed towards "survival" instead of student learning, we argue that any teacher preparation program (e.g., physics) should strive to develop in pre-service teachers strong habits of mind and practice that will serve as an underlying support structure for beginning teachers. These habits form the core of the conceptual model described in the paper. We provide examples of habits, which physics teachers need to develop before starting teaching and propose mechanisms for the development of such habits.

physics.ed-ph