Scalable and Personalized Oral Assessments Using Voice AI
Written work no longer certifies that a student understands it: a polished analysis now says little about who did the thinking. Oral examinations restore that evidentiary link, but they have never scaled, because conducting and grading them is expensive. We report on a system in which voice AI conducts a personalized oral exam and a council of three large language models (LLMs) grades the transcript, each model scoring independently and then revising after reading the others. Across two undergraduate cohorts at NYU Stern (36 students in Fall 2025, 37 in Spring 2026), a voice subscription covered all speaking time and grading stayed under one dollar per exam. The deployments yield five practical engineering lessons that should generalize wherever understanding must be tested under questioning, from job interviews to professional certification.