arXiv · 2604.23347
Evaluating Large Language Models on Computer Science University Exams in Data Structures
Abstract
We present a comprehensive evaluation of Large Language Models (LLMs) on Computer Science (CS) Data Structure examination questions. Our work introduces a new benchmark dataset comprising exam questions from Tel Aviv University (TAU), curated to assess LLMs' abilities in handling closed and multiple-choice questions. We evaluated the performance of OpenAI's GPT 4o and Anthropic's Claude 3.5, popular LLMs, alongside two smaller LLMs, Mathstral 7B and LLaMA 3 8B, across the TAU exams benchmark. Our findings provide insight into the current capabilities of LLMs in CS education.
Explore related subjects
Keep this discovery
Edan Gabay, Yael Maoz, Jonathan Stahl, Naama Maoz, Abdo Amer, Orr Eilat, Hanoch Levy, Michal Kleinbort, Amir Rubinstein, Adi Haviv. 2026-04-25. Evaluating Large Language Models on Computer Science University Exams in Data Structures. https://arxiv.org/abs/2604.23347
Cite the original work for its findings. Save a collection to share your selection of sources.