Searcharxiv⌕ Search

arXiv · 2609.38951

Measuring Student Self-Assessment against Viva-Demonstrated Mastery in a Large First-Year Programming Course

Abstract

Mastery-based education increasingly places the reporting of learning progress in students' hands, who record task completion on learning dashboards. The usefulness of such self-reports depends on how closely reported mastery corresponds to demonstrated competence. Most evidence on student self-assessment compares an overall self-rating with an overall examination score and therefore provides limited evidence about which tasks or which students account for the mismatch. This study examines first-year students' self-assessment against viva-demonstrated mastery at the level of individual tasks across a ladder of sixty programming tasks. The study draws on a large first-year programming course taught in 2023, involving 203 students and 12 examiners, in which every reported task was verified through an oral viva. Because a task entered the Viva only after it was reported, the design is one-sided and captures over-estimation but not under-estimation. Of 11,093 reported tasks, 10,885 (98.1%) were demonstrated, indicating a high degree of correspondence between self-report and demonstrated mastery. The remaining 208 overestimations were not evenly distributed. A small number of students accounted for most of the errors, and they occurred mainly on difficult tasks near the end of the task ladder rather than on higher-point tasks. This task-level analysis shows that high overall self-assessment accuracy can coexist with specific areas where reported and demonstrated mastery diverge. It also provides a practical basis for directing additional verification and formative feedback toward students and tasks where such divergence is more likely.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sakshi Sharma, Pavani Ayinampudi, Aditya B. M. V., Jinal Gupta, Prakash Hegade, Rohit Sharma, Meenakshi V, S. R. S. Iyengar. 2026-09-30. Measuring Student Self-Assessment against Viva-Demonstrated Mastery in a Large First-Year Programming Course. https://arxiv.org/abs/2609.38951

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

AI Behavioral Science: A Framework and Agenda

We discuss the challenges and opportunities present in the rapidly emerging area of ``AI Behavioral Science.'' We frame it via three subfields. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop models of AI and methods for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to model, assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, predicting, and analyzing human behaviors. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to analyze and model human-AI interactions including how human and AI behaviors depend on interactions at the individual level, how interacting systems of humans and AI behave, and ultimately how AI's integration into society affects economic and political outcomes. We discuss current research, questions, agendas, and goals in each of these three subfields and how they depend upon each other.

cs.HC↗

Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship

Assessing learning in large classrooms presents a significant challenge for individual instructors, who may have limited capacity to evaluate the understanding, participation, and assessment behaviour of every student. Peer assessments have been a way of distributing this responsibility among learners, allowing them to evaluate and provide feedback to one another while reducing dependence on instructor-led assessments. Building on this approach, we implemented a peer validation model within a large, multi-institutional internship programme in which students who demonstrated sufficient understanding were authorised to assess and validate their peers through short oral discussions. The assessment process began with the instructor validating a small group of students, who were then authorised to validate their peers, allowing the process to gradually expand across the cohort and operate at scale. This study examines how participants experienced the model and the extent to which assessment integrity was maintained, using an end-of-programme survey of 238 consenting respondents. Most participants regarded the activity as worthwhile, with 79.8% reporting that they solved problems they could not previously solve. However, 29.0% acknowledged at least one instance of reduced effort, a lowered validation standard, or reciprocal validation, while 88.7% believed that at least a little validation had occurred without proper examination. When asked how the process could be strengthened, participants selected post-validation discussion of solutions approximately twice as often as closer auditing or mentor-led validation. These findings provide descriptive evidence of both the potential and the integrity challenges of using peer validation as a scalable assessment approach in large learning environments.

cs.HC↗

Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI

Generative AI can support writing, but frictionless access may cause cognitive offloading before users develop their own ideas. We introduce Engage-to-Unlock, a productive-friction mechanism that unlocks generative capabilities after users meaningfully engage with the task. In a controlled experiment (N = 398), participants completed a writing task under one of four conditions: Human-Only, Standard Chatbot, Engage-to-Unlock, or Time-Matched Unlock, which matched unlock timing to Engage-to-Unlock participants but independent of users' engagement, then evaluated passages for evidence and inferential errors. Results show that Engage-to-Unlock redistributed effort across tasks: participants spent more time writing and less time evaluating, without increasing overall task duration. They also submitted more prompts than in other AI-assisted conditions and showed the highest accuracy-per-time evaluation efficiency across conditions. These findings suggest that designing GenAI access to encourage early human engagement may provide a productive form of friction, while retaining active AI use and efficient downstream evaluation.

cs.HC↗