Searcharxiv⌕ Search

arXiv · 2610.09460

Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education

Abstract

As Large Language Models (LLMs) are increasingly adopted for automated grading and feedback in higher education, their structured outputs, including multi-dimensional rubric scores, detailed feedback comments, and improvement suggestions, create an appearance of thorough analytic evaluation. This study examines whether these outputs deliver what they appear to offer. Using the JorGPT dataset of 3,041 student responses to 50 open-ended computer science questions, scored by both human instructors and three commercial LLMs, we identify three systematic discrepancies between the apparent and actual quality of LLM-generated grading and feedback. The sub-dimension scores are highly correlated (r = 0.82-0.99, VIF up to 45), providing redundant rather than independent diagnostic information. The textual feedback rarely detects student misconceptions (5-7% vs. 15-31% for teachers), functioning as a coverage checklist rather than a diagnostic instrument. The feedback tone remains uniformly positive regardless of response quality, lacking the severity modulation observed in human feedback. Additionally, grading accuracy varies significantly by knowledge domain, with procedural topics most reliable. These findings provide empirically grounded guidance on which aspects of LLM-generated grading and feedback can be relied upon and which require continued human oversight.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xi Zhao, Xinyue Jiao, Zhen Xu. 2026-10-07. Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education. https://arxiv.org/abs/2610.09460

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Who Does Withholding Delay? A Welfare Model of Open-Weight AI Release Under Asymmetric Proliferation

Withholding a dual-use AI model delays only the actors that lack other routes to a comparable capability. If sophisticated adversaries obtain substitutes faster than distributed defenders, restriction can delay defenders more than the adversaries it targets. We compare controlled access, a defender-first window followed by public release, safeguarded open weights, and minimally restricted open weights in a discounted welfare model with actor-specific substitute acquisition. Under exponential acquisition, restriction gives adversaries a positive discounted access advantage exactly when they substitute faster than defenders, and, with equal usefulness, immediate release adds more expected capability at a fixed horizon to the slower-substituting group. Neither result implies that release is preferable, because opportunistic misuse, defensive reach, safeguard friction, and irreversible losses can reverse the ranking. In a linear benchmark, broad release overtakes control above a unique adversary-substitution threshold whenever such a threshold exists, and we derive the probability that selected defenders deploy before both adversary substitution and public release. In a nonlinear implementation, each of the four policies is optimal somewhere in the parameter space. Three nested 2,048-point designs over thirteen inputs show that policy shares depend strongly on the chosen parameter bounds. Release records and cybersecurity reports illustrate the quantities a release review would need to measure and are kept separate from the calibration.

cs.CY↗

What Personal Information Improves LLM-Based Next-Location Prediction?

Large language models (LLMs) are increasingly used for individual next-location prediction, with personal information easily added to prompts alongside mobility history. Yet the incremental predictive value of such information remains unclear. Using linked sociodemographic records and mobility traces from 5,000 Shenzhen residents, this study separates model responsiveness from predictive value. GPT-5 is the primary model, with GPT-5.5 and Claude Opus 4.6 used for replication. In 1,000 paired prediction instances, models rank 100 candidate destinations with and without age, gender, occupation and income while all other inputs are held fixed. Behavioural history raises top-1 accuracy from 5.6% to 18.5% as prior history increases from zero to six days. By contrast, sociodemographic attributes produce no detectable overall gain, although replacing correct attributes with those of another person reduces accuracy by 5.4 percentage points. Candidate construction also matters, removing distance raises accuracy by 7.7 points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These findings identify behavioural history as the clearest source of incremental value and show that personal information should be evaluated under matched, explicitly specified conditions before its privacy and governance costs are justified.

cs.CY↗

Political polarization and mental wellbeing: asymmetric evidence for bidirectionality

A growing body of evidence suggests that political polarization and mental health and wellbeing influence each other. Yet the two directions, how polarization shapes mental wellbeing and how mental wellbeing shapes polarization, remain siloed across disciplines. We review both literatures to evaluate the strength and alignment of evidence for bidirectional effects. We find that the literature varies widely in constructs, measures, and levels of analysis, and that the evidence for the two pathways is asymmetric: links from polarization to mental wellbeing are more direct, grounded in well-developed theoretical frameworks and social-relational mechanisms. In contrast, the reverse pathway is substantially less direct, with studies often examining cognitive and socio-emotional factors, or political outcomes adjacent to polarization rather than polarization itself. We identify clear gaps in the literature highlighting the need for future studies with conceptual clarity, unified frameworks, validated measures, and multilevel designs examining both pathways within the same empirical and theoretical setting.

cs.CY↗