Searcharxiv⌕ Search

arXiv subjects

Christian D. Schunn

Publications and source records attributed to Christian D. Schunn.

7 recordsLinked to original sources

Questioning your brilliance in physics: Differential shifts in fixed mindsets by grade and gender

Students' domain-specific mindsets and their beliefs about their capacity to improve through effort play a crucial role in shaping their experiences and decisions to persist in STEM disciplines. Physics is generally seen as a field requiring innate brilliance, which can reinforce fixed mindsets, particularly after initial setbacks in performance that are common in introductory university courses. In this study, we examine changes in fixed mindsets and potential gender differences in an introductory calculus-based physics course. Our sample consisted of 508 students with an average age of 18, predominantly White, with men comprising the majority. Based upon survey response distributions, three distinct mindset categories were identified: Hesitant, Hopeful, and Confident, describing how strongly students rejected a fixed mindset in physics. The results suggested large gender differences in distributions at the high and low-end groups. We also found an overall decline toward fixed mindsets across the course, and logistic regressions controlling for initial mindsets showed that women were significantly more likely than men to shift away from the Confident category. While the majority of men tended to stay within the Confident category, the majority of women moved away from it. Particularly, this differential shift was seen among students receiving Bs or Cs, the most commonly awarded grades in this course. Furthermore, there were relatively small differences in the probability of change within men as a function of grades received, whereas women showed marked declines toward fixed beliefs with either a B or C. Our findings provide empirical evidence for the dynamic, grade-sensitive nature of students' mindsets in a calculus-based physics course.

physics.ed-ph↗

Temporal Stability and Few-Shot Prompting in Math Task Assessment

As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to prompt engineering techniques. This longitudinal study focused on different AI tools' ability to use the Task Analysis Guide (TAG; Stein \& Smith, 1998) to classify the cognitive demand of mathematics tasks. In particular, it examined whether this classification ability changed with (1) model version updates over time and (2) few-shot prompting using exemplar tasks. We tested a general-purpose AI tool (Gemini) and an education-specific AI tool (Coteach). The specific tools were selected because of their relatively high performance on relevant published benchmarks and prior task-specific tests. Models were tested at baseline, retested with model version updates, and then tested again using few-shot prompting (two exemplar tasks for each cognitive demand category). Results revealed that newer model versions alone produced mixed effects: Gemini's accuracy remained stable at 58\%, while Coteach's accuracy decreased from 75\% to 50\%. However, few-shot prompting improved both models' performance: Gemini increased to 67\% and Coteach recovered to 75\% accuracy. These findings demonstrate that prompt engineering techniques can have larger and more reliable effects than passive model improvements, and that version updates may not always improve performance on specialized educational tasks. The study has important implications for how educators and researchers should approach AI tool selection, evaluation, and implementation in educational contexts.

cs.AI↗

The Gendered Cost of Lower Grades: Women's Physics Perceived Recognition and Identity Suffer Disproportionately If They Earn Less Than A Grade

Perceptions of disciplinary recognition and identity can be shaped by various forms of feedback and experiences. Here we focus on the potential effects of course grades on the perceievd recognition and physics identity of students. We analyze patterns in changes in physics identity and perceived recognition from pre course to post course across three cohorts of university students enrolled in calculus-based Physics 1 (N=1,681). Students not receiving A grade, on average, showed declines in physics identity and perceived recognition. Even a B grade resulted in declines, and the declines were nonlinear across lower grades. Changes in perceived recognition fully mediated the changes in identity. Importantly, women showed significantly larger declines in identity and perceived recognition, compared to men, if they got less than A grade. The gender moderation was specifically localized to changes in perceived recognition, with no further gender effects on identity beyond the cascading effects on perceived recognition.

physics.ed-ph↗

Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities

While recent research has explored AI tools' ability to classify the quality of mathematical tasks (arXiv:2603.03512), little is known about their capacity to increase the quality of existing tasks. This study investigated whether AI tools could successfully upgrade low-cognitive-demand mathematics tasks. Eleven tools were tested, including six broadly available, general-purpose AI tools (e.g., ChatGPT and Claude) and five tools specialized for mathematics teachers (e.g., Khanmigo, coteach.ai). Using the Task Analysis Guide framework (Stein & Smith, 1998), we prompted AI tools to modify two different types of low-demand mathematical tasks. The prompting strategy aimed to represent likely approaches taken by knowledgeable teachers, rather than extensive optimization to find a more effective prompt (i.e., an optimistic typical outcome). On average, AI tools were only moderately successful: tasks were accurately upgraded only 64% of the time, with different AI tool performance ranging from quite weak (33%) to broadly successful (88%). Specialized tools were only moderately more successful than general-purpose tools. Failure modes included both "undershooting" (maintaining low cognitive demand) and "overshooting" (elevating tasks to an overly ambitious target category that likely would be rejected by teachers). Interestingly, there was a small negative correlation (r = -.35) between whether a given AI tool was able to correctly classify the cognitive demand of tasks and whether the AI was able to upgrade tasks, showing that the ability to modify tasks (i.e., a generative task) represents a distinct capability from the ability to classify them (i.e., judgement using a rubric). These findings have important implications for understanding AI's potential role in curriculum adaptation and highlight the need for specialized approaches to support teachers in modifying instructional materials.

cs.AI↗

Baseline Performance of AI Tools in Classifying Cognitive Demand of Mathematical Tasks

Teachers face increasing demands on their time, particularly in adapting mathematics curricula to meet individual student needs while maintaining cognitive rigor. This study evaluates whether AI tools can accurately classify the cognitive demand of mathematical tasks, which is important for creating or adapting tasks that support student learning. We tested eleven AI tools: six general-purpose (ChatGPT, Claude, DeepSeek, Gemini, Grok, Perplexity) and five education-specific (Brisk, Coteach AI, Khanmigo, Magic School, School$.$AI), on their ability to categorize mathematics tasks across four levels of cognitive demand using a research-based framework. The goal was to approximate the performance teachers will achieve with straightforward prompts. On average, AI tools accurately classified cognitive demand in only 63% of cases. Education-specific tools were not more accurate than general-purpose tools, and no tool exceeded 83% accuracy. All tools struggled with tasks at the extremes of cognitive demand (Memorization and Doing Mathematics), exhibiting a systematic bias toward middle-category levels (Procedures with/without Connections). The tools often gave plausible-sounding explanations likely to be persuasive to novice teachers. Error analysis of AI tools' misclassification of the broad level of cognitive demand (high vs. low) revealed that tools consistently overweighted surface textual features over underlying cognitive processes. Further, AI tools showed weaknesses in reasoning about factors that make tasks higher vs. lower cognitive demand. Errors stemmed not from ignoring relevant dimensions, but from incorrectly reasoning about multiple task aspects. These findings carry implications for AI integration into teacher planning workflows and highlight the need for improved prompt engineering and tool development for educational applications.

cs.CY↗

A mismatch between self-efficacy and performance: Undergraduate women in engineering tend to have lower self-efficacy despite earning higher grades than men

There is a significant underrepresentation of women in many Science, Technology, Engineering, and Mathematics (STEM) majors and careers. Prior research has shown that self-efficacy can be a critical factor in student learning, and that there is a tendency for women to have lower self-efficacy than men in STEM disciplines. This study investigates gender differences in the relationship between engineering students' self-efficacy and course grades in foundational courses. By focusing on engineering students, we examined these gender differences simultaneously in four STEM disciplines (mathematics, engineering, physics, and chemistry) among the same population. Using survey data collected longitudinally at three time points and course grade data from five cohorts of engineering students at a large US-based research university, effect sizes of gender differences are calculated using Cohen's d on two measures: responses to survey items on discipline-specific self-efficacy and course grades in all first-year foundational courses and second-year mathematics courses. In engineering, physics, and mathematics courses, we find sizeable discrepancies between self-efficacy and performance, with men appearing significantly more confident than women despite small or reverse direction differences in grades. In chemistry, women earn higher grades and have higher self-efficacy. The patterns are consistent across courses within each discipline. All self-efficacy gender differences close by the fourth year except physics self-efficacy. The disconnect between self-efficacy and course grades across subjects provides useful clues for targeted interventions to promote equitable learning environments. The most extreme disconnect occurs in physics and may help explain the severe underrepresentation of women in "physics-heavy" engineering disciplines, highlighting the importance of such interventions.

physics.ed-ph↗

Connecting Three Pivotal Concepts in K-12 Science State Standards and Maps of Conceptual Growth to Research in Physics Education

This paper describes three conceptual areas in physics that are particularly important targets for educational interventions in K-12 science. These conceptual areas are force and motion, conservation of energy, and geometrical optics, which were prominent in the US national and four US state standards that we examined. The four US state standards that were analyzed to explore the extent to which the K-12 science standards differ in different states were selected to include states in different geographic regions and of different sizes. The three conceptual areas that were common to all the four state standards are conceptual building blocks for other science concepts covered in the K-12 curriculum. Since these three areas have been found to be ripe with deep student misconceptions that are resilient to conventional physics instruction, the nature of difficulties in these areas is described in some depth, along with pointers towards approaches that have met with some success in each conceptual area.

physics.ed-ph↗