SearcharxivSearch

arXiv subjects

Zhongzhou Chen

Publications and source records attributed to Zhongzhou Chen.

17 recordsLinked to original sources

Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment

Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted GenAI access enables task outsourcing that undermines the validity of traditional assessments; blanket prohibitions are difficult to enforce, may push use underground, and do little to prepare students for workplaces where GenAI-supported workflows are increasingly common. This paper addresses this dilemma by proposing a framework grounded in Evidence-Centered Design (ECD) that treats GenAI as a design variable within the assessment argument rather than an external threat to it. The framework analyzes how GenAI reshapes the student model, evidence model, and task model, and uses this analysis to articulate three principled governance stances. Restrict is warranted when GenAI would contaminate the inferential link between student work products and targeted unaided proficiency. Scaffold is warranted when bounded GenAI support can support peripheral demands without revealing the target construct, preserving inferential interpretability. Require is warranted when the target construct is disciplinary AI interaction competency and tasks can be designed to elicit process artifacts, including prompts, critiques, and revisions, that make student reasoning observable, scorable, and distinguishable from AI-generated output. This framework specifies when to restrict, scaffold, or require GenAI use in STEM assessment. We present two task designs deployed in an introductory physics course and demonstrate that disciplinary AI interaction competencies are observable in student response artifacts and can be scored using defensible rubrics grounded in student data and expert knowledge. By situating GenAI governance within validity arguments, the framework offers actionable guidance for preserving learning integrity while supporting authentic preparation for AI-enabled professional environments.

cs.CY

Scalable Generation and Validation of Isomorphic Physics Problems with GenAI

Traditional synchronous STEM assessments face growing challenges including accessibility barriers, security concerns from resource-sharing platforms, and limited comparability across institutions. We present a framework for generating and evaluating large-scale isomorphic physics problem banks using Generative AI to enable asynchronous, multi-attempt assessments. Isomorphic problems test identical concepts through varied surface features and contexts, providing richer variation than conventional parameterized questions while maintaining consistent difficulty. Our generation framework employs prompt chaining and tool use to achieve precise control over structural variations (numeric values, spatial relations) alongside diverse contextual variations. For pre-deployment validation, we evaluate generated items using 17 open-source language models (LMs) (0.6B-32B) and compare against actual student performance (N>200) across three midterm exams. Results show that 73% of deployed banks achieve statistically homogeneous difficulty, and LMs pattern correlate strongly with student performance (Pearson's $ρ$ up to 0.594). Additionally, LMs successfully identify problematic variants, such as ambiguous problem texts. Model scale also proves critical for effective validation, where extremely small (<4B) and large (>14B) models exhibit floor and ceiling effects respectively, making mid-sized models optimal for detecting difficulty outliers.

cs.CY

Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use

Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted access can enable task outsourcing that undermines the validity of traditional assessments, while blanket prohibitions are difficult to enforce, may drive use underground, and do little to prepare students for workplaces where GenAI supported workflows are increasingly common. This paper proposes a student focused framework grounded in Evidence Centered Design (ECD) that specifies when to restrict, scaffold, or require GenAI use in STEM assessment. The framework extends existing AI use taxonomies by providing decision rules that link target constructs, evidence requirements, and task characteristics to governance regimes. Restriction is warranted when GenAI threatens construct relevant evidence for unaided proficiency, particularly for foundational knowledge and routine skills. Scaffolding is appropriate when bounded GenAI support reduces peripheral demands while maintaining interpretability. Requiring GenAI is appropriate when the target construct involves human AI collaboration and AI literacy. Using examples from introductory physics, we illustrate how tasks can be designed under different GenAI use policies. The framework provides guidance for preserving learning integrity while supporting preparation for AI enabled environments.

cs.CY

When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment

Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary problem: \textit{predicting when an LLM grader is likely to be correct}. This enables selective automation where high-confidence predictions are processed automatically while uncertain cases are flagged for human review. We compare three confidence estimation methods (self-reported confidence, self-consistency voting, and token probability) across seven LLMs of varying scale (4B to 120B parameters) on three educational datasets: RiceChem (long-answer chemistry), SciEntsBank, and Beetle (short-answer science). Our experiments reveal that self-reported confidence consistently achieves the best calibration across all conditions (avg ECE 0.166 vs 0.229 for self-consistency). Surprisingly, self-consistency remains 38\% worse despite requiring 5$\times$ the inference cost. Larger models exhibit substantially better calibration though gains vary by dataset and method (e.g., a 28\% ECE reduction for self-reported), with GPT-OSS-120B achieving the best calibration (avg ECE 0.100) and strong discrimination (avg AUC 0.668). We also observe that confidence is strongly top-skewed across methods, creating a ``confidence floor'' that practitioners must account for when setting thresholds. These findings suggest that simply asking LLMs to report their confidence provides a practical approach for identifying reliable grading predictions. Code is available \href{https://github.com/sonkar-lab/llm_grading_calibration}{here}.

cs.CL

Opportunities and Challenges in Harnessing Digital Technology for Effective Teaching and Learning

Most of today's educators are in no shortage of digital and online learning technologies available at their fingertips, ranging from Learning Management Systems such as Canvas, Blackboard, or Moodle, online meeting tools, online homework, and tutoring systems, exam proctoring platforms, computer simulations, and even virtual reality/augmented reality technologies. Furthermore, with the rapid development and wide availability of generative artificial intelligence (GenAI) services such as ChatGPT, we are just at the beginning of harnessing their potential to transform higher education. Yet, facing the large number of available options provided by cutting-edge technology, an imminent question on the mind of most educators is the following: how should I choose the technologies and integrate them into my teaching process so that they would best support student learning? We contemplate over these types of important and timely questions and share our reflections on evidence-based approaches to harnessing digital learning tools using a Self-regulated Engaged Learning Framework we have employed in our research in physics education that can be valuable for educators in other disciplines.

physics.ed-ph

Reliable generation of isomorphic physics problems using Generative AI with prompt-chaining and tool use

We present a method for generating large numbers of isomorphic physics problems using generative AI services such as ChatGPT, through prompt chaining and tool use. This approach enables precise control over structural variations-such as numeric values and spatial relations-while supporting diverse contextual variations in the problem body. By utilizing the Python code interpreter, the method supports automatic solution validation and simple diagram generation, addressing key limitations in existing LLM-based methods. We generated two example isomorphic problem banks and compared the outcome against two simpler prompt-based approaches. Results show that prompt-chaining produces significantly higher quality and more consistent outputs than simpler, non-chaining prompts. We also show that GenAI services can be used to validate the quality of the generated isomorphic problems. This work demonstrates a promising method for efficient and scalable problem creation accessible to the average instructor, which opens new possibilities for personalized adaptive testing and automated content development.

physics.ed-ph

Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback

This study examines the feasibility and potential advantages of using large language models, in particular GPT-4o, to perform partial credit grading of large numbers of student written responses to introductory level physics problems. Students were instructed to write down verbal explanations of their reasoning process when solving one conceptual and two numerical calculation problems on in class exams. The explanations were then graded according to a 3-item rubric with each item grades as binary (1 or 0). We first demonstrate that machine grading using GPT-4o with no examples nor reference answer can reliably agree with human graders on 70%-80% of all cases, which is equal to or higher than the level at which two human graders agree with each other. Two methods are essential for achieving this level of accuracy: 1. Adding explanation language to each rubric item that targets the errors of initial machine grading. 2. Running the grading process 5 times and taking the most frequent outcome. Next, we show that the variation in outcomes across 5 machine grading attempts as measured by the Shannon Entropy can serve as a grading confidence index, allowing a human instructor to identify ~40% of all potentially incorrect gradings by reviewing just 10 - 15% of all responses. Finally, we show that it is straightforward to use GPT-4o to write clear explanations of the partial credit grading outcomes. Those explanations can be used as feedback for students, which will allow students to understand their grades and raise different opinions when necessary. Almost all feedback messages generated were rated 3 or above on a 5-point scale by two experienced instructors. The entire grading and feedback generating process cost roughly $5 per 100 student answers, which shows immense promise for automating labor-intensive grading process by a combination of machine grading with human input and supervision.

physics.ed-ph

Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education

This paper introduces a novel approach to create a high-resolution "map" for physics learning: an "atomic" learning objectives (LOs) system designed to capture detailed cognitive processes and concepts required for problem solving in a college-level introductory physics course. Our method leverages Large Language Models (LLMs) for automated labeling of physics questions and introduces a comprehensive set of metrics to evaluate the quality of the labeling outcomes. The atomic LO system, covering nine chapters of an introductory physics course, uses a "subject-verb-object'' structure to represent specific cognitive processes. We apply this system to 131 questions from expert-curated question banks and the OpenStax University Physics textbook. Each question is labeled with 1-8 atomic LOs across three chapters. Through extensive experiments using various prompting strategies and LLMs, we compare automated LOs labeling results against human expert labeling. Our analysis reveals both the strengths and limitations of LLMs, providing insight into LLMs reasoning processes for labeling LOs and identifying areas for improvement in LOs system design. Our work contributes to the field of learning analytics by proposing a more granular approach to mapping learning objectives with questions. Our findings have significant implications for the development of intelligent tutoring systems and personalized learning pathways in STEM education, paving the way for more effective "learning GPS'' systems.

cs.CY

Achieving Human Level Partial Credit Grading of Written Responses to Physics Conceptual Question using GPT-3.5 with Only Prompt Engineering

Large language modules (LLMs) have great potential for auto-grading student written responses to physics problems due to their capacity to process and generate natural language. In this explorative study, we use a prompt engineering technique, which we name "scaffolded chain of thought (COT)", to instruct GPT-3.5 to grade student written responses to a physics conceptual question. Compared to common COT prompting, scaffolded COT prompts GPT-3.5 to explicitly compare student responses to a detailed, well-explained rubric before generating the grading outcome. We show that when compared to human raters, the grading accuracy of GPT-3.5 using scaffolded COT is 20% - 30% higher than conventional COT. The level of agreement between AI and human raters can reach 70% - 80%, comparable to the level between two human raters. This shows promise that an LLM-based AI grader can achieve human-level grading accuracy on a physics conceptual problem using prompt engineering techniques alone.

physics.ed-ph

Comparing student performance on a multi-attempt asynchronous assessment to a single-attempt synchronous assessment in introductory level physics

The current paper examines the possibility of replacing conventional synchronous single-attempt exam with more flexible and accessible multi-attempt asynchronous assessments in introductory-level physics by using large isomorphic problem banks. We compared student's performance on both numeric and conceptual problems administered on a multi-attempt, asynchronous quiz to their performance on isomorphic problems administered on a subsequent single-attempt, synchronous exam. We computed the phi coefficient and the McNemar's test statistic for the correlation matrix between paired problems on both assessments as a function of the number of attempts considered on the quiz. We found that for the conceptual problems, a multi-attempt quiz with five allowed attempts could potentially replace similar problems on a single-attempt exam, while there was a much weaker association for the numerical questions beyond two quiz attempts.

physics.ed-ph

Exploring Generative AI assisted feedback writing for students' written responses to a physics conceptual question with prompt engineering and few-shot learning

Instructor's feedback plays a critical role in students' development of conceptual understanding and reasoning skills. However, grading student written responses and providing personalized feedback can take a substantial amount of time. In this study, we explore using GPT-3.5 to write feedback to student written responses to conceptual questions with prompt engineering and few-shot learning techniques. In stage one, we used a small portion (n=20) of the student responses on one conceptual question to iteratively train GPT. Four of the responses paired with human-written feedback were included in the prompt as examples for GPT. We tasked GPT to generate feedback to the other 16 responses, and we refined the prompt after several iterations. In stage two, we gave four student researchers the 16 responses as well as two versions of feedback, one written by the authors and the other by GPT. Students were asked to rate the correctness and usefulness of each feedback, and to indicate which one was generated by GPT. The results showed that students tended to rate the feedback by human and GPT equally on correctness, but they all rated the feedback by GPT as more useful. Additionally, the successful rates of identifying GPT's feedback were low, ranging from 0.1 to 0.6. In stage three, we tasked GPT to generate feedback to the rest of the student responses (n=65). The feedback was rated by four instructors based on the extent of modification needed if they were to give the feedback to students. All the instructors rated approximately 70% of the feedback statements needing only minor or no modification. This study demonstrated the feasibility of using Generative AI as an assistant to generating feedback for student written responses with only a relatively small number of examples. An AI assistance can be one of the solutions to substantially reduce time spent on grading student written responses.

physics.ed-ph

Reforming Physics Exams Using Openly Accessible Large Isomorphic Problem Banks created with the assistance of Generative AI: an Explorative Study

This paper explores using large isomorphic problem banks to overcome many challenges of traditional exams in large STEM classes, especially the threat of content sharing websites and generative AI to the security of exam items. We first introduce an efficient procedure for creating large numbers of isomorphic physics problems, assisted by the large language model GPT-3 and several other open-source tools. We then propose that if exam items are randomly drawn from large enough problem banks, then giving students open access to problem banks prior to the exam will not dramatically impact students' performance on the exam or lead to wide-spread rote-memorization of solutions. We tested this hypothesis on two mid-term physics exams, comparing students' performance on problems drawn from open isomorphic problem banks to similar transfer problems that were not accessible to students prior to the exam. We found that on both exams, both open bank and transfer problems had the highest difficulty. The differences in percent correct were between 5% to 10%, which is comparable to the differences between different isomorphic versions of the same problem type. Item response theory analysis found that both types of problem have high discrimination (>1.5) with no significant differences. Student performance on open-bank and transfer problems are highly correlated with each other, and the correlations are stronger than average correlations between problems on the exam. Exploratory factor analysis also found that open-bank and transfer problems load on the same factor, and even formed their own factor on the second exam. Those observations all suggest that giving students open access to large isomorphic problem banks only had a small impact on students' performance on the exam but could have significant potential in reforming traditional classroom exams.

physics.ed-ph

A Multi-Level Trace Clustering Analysis Scheme for Measuring Students' Self-Regulated Learning Behavior in a Master-Based Online Learning Environment

The study introduces a new analysis scheme to analyze trace data and visualize students' self-regulated learning strategies in a mastery-based online learning modules platform. The pedagogical design of the platform resulted in fewer event types and less variability in student trace data. The current analysis scheme overcomes those challenges by conducting three levels of clustering analysis. On the event level, mixture-model fitting is employed to distinguish between abnormally short and normal assessment attempts and study events. On the module level, trace level clustering is performed with three different methods for generating distance metrics between traces, with the best performing output used in the next step. On the sequence level, trace level clustering is performed on top of module-level clusters to reveal students' change of learning strategy over time. We demonstrated that distance metrics generated based on learning theory produced better clustering results than pure data-driven or hybrid methods. The analysis showed that most students started the semester with productive learning strategies, but a significant fraction shifted to a multitude of less productive strategies in response to increasing content difficulty and stress. The observations could prompt instructors to rethink conventional course structure and implement interventions to improve self-regulation at optimal times.

physics.ed-ph

Toward more accurate measurement of the impact of online instructional design on students' ability to transfer physics problem-solving skills

In two earlier studies, we developed a new method to measure students' ability to transfer physics problem solving skills to new contexts using a sequence of online learning modules, and implemented two interventions in the form of additional learning modules designed to improve transfer ability. The current paper introduces a new data analysis scheme that could improve the accuracy of the measurement by accounting for possible differences in students' goal orientation and behavior, as well as revealing the possible mechanism by which one of the two interventions improves transfer ability. Based on a two by two framework of self-regulated learning, students with a performance-avoidance oriented goal are more likely to guess on some of the assessment attempts in order to save time, resulting in an underestimation of the student populations' transfer ability. The current analysis shows that about half of the students had frequent brief initial assessment attempts, and significantly lower correct rates on certain modules, which we think is likely to have originated at least in part from students adopting a performance-avoidance strategy. We then divided the remaining population, for which we can be certain that few students adopted a performance-avoidance strategy, based on whether they interacted with one of the intervention modules designed to develop basic problem solving skills, or passed that module on their first attempt without interacting with the instructional material. By comparing to propensity score matched populations from a previous semester, we found that the improvement in subsequent transfer performance observed in a previous study mainly came from the latter population, suggesting that the intervention served as an effective reminder for students to activate existing skills, but fell short of developing those skills among those who have yet to master it.

physics.ed-ph

Measuring the Impact of COVID-19 Induced Campus Closure on Student Self-Regulated Learning in Physics Online Learning Modules

This paper examines the impact of COVID-19 induced campus closure on university students' self-regulated learning behavior by analyzing click-stream data collected from student interactions with 70 online learning modules in a university physics course. To do so, we compared the trend of six types of actions related to the three phases of self-regulated learning before and after campus closures and between two semesters. We found that campus closure changed students' planning and goal setting strategies for completing the assignments, but didn't have a detectable impact on the outcome or the time of completion, nor did it change students' self-reflection behavior. The results suggest that most students still manage to complete assignments on time during the pandemic, while the design of online learning modules might have provided the flexibility and support for them to do so.

physics.ed-ph

Exploring the relation between students' online learning behavior and course performance by including contextual information in data analysis

This study examines whether including more contextual information in data analysis could improve our ability to identify the relation between students' online learning behavior and overall performance in an introductory physics course. We created four linear regression models correlating students' pass-fail events in a sequence of online learning modules with their normalized total course score. Each model takes into account an additional level of contextual information than the previous one, such as student learning strategy and duration of assessment attempts. Each of the latter three models is also accompanied by a visual representation of students' interaction states on each learning module. We found that the best performing model is the one that includes the most contextual information, including instruction condition, internal condition, and learning strategy. The model shows that while most students failed on the most challenging learning module, those with normal learning behavior are more likely to obtain higher total course scores, whereas students who resorted to guessing on the assessments of subsequent modules tended to receive lower total scores. Our results suggest that considering more contextual information related to each event can be an effective method to improve the quality of learning analytics, leading to more accurate and actionable recommendations for instructors.

physics.ed-ph

Measuring the Effectiveness of Learning Resources Via Student Interaction with Online Learning Modules

We present a new method for measuring the effectiveness of online learning resources, through the analysis of time-stamped log data of students' interaction with a sequence of online learning modules created based on the concept of mastery learning. Each module was designed to assess students' mastery of one topic before and after interacting with the learning resources in the module. In addition, analysis of log data provides information on students' test-taking effort when completing the assessment, and learning effort when interacting with the learning resources. Combining these three measurements provides accurate information on the quality of each online learning module, as well as detailed suggestions for future improvements. The results from data collected from all 10 modules are presented in a sequence of sunburst charts, an intuitive visual representation designed to allow the average instructor to quickly grasp the key outcomes of the data analysis, and identify less effective modules for future improvements. Online learning modules can be implemented much more frequently than either clinical experiments or classroom assessments, and can provide more interpretable data than most existing online courses, making it a valuable tool for quickly assessing the effectiveness of online learning resources.

physics.ed-ph