SearcharxivSearch

arXiv subjects

Mahir Akgun

Publications and source records attributed to Mahir Akgun.

9 recordsLinked to original sources

Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation

Designing effective pretest questions is challenging at scale: high-quality questions require careful calibration of openness, cognitive depth, and alignment with learning objectives, yet generating and evaluating them manually is time-consuming. We present an AI-assisted workflow for pretest question development that combines automated generation, rubric-based evaluation, and iterative selection. Because the workflow relies on machine evaluation to filter questions at scale, we investigate the alignment between human and machine judgments across a 2x2 design varying rubric operationalization and evaluation mode. Our findings show that human-machine disagreements are systematic rather than random, that rubric revision has a larger effect on alignment than rationale-first evaluation, and that the two interventions are complementary. These findings highlight that scalable AI-assisted pretesting depends not only on generation capability but on how pedagogical quality is operationalized for machine interpretation.

cs.HC

Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study

Pretesting - attempting problems before instruction - supports learning by activating prior knowledge and sharpening attention to subsequent instruction. Recent work suggests that adaptive AI-assisted pretesting can yield further advantages, particularly for tasks requiring higher-order reasoning, yet it remains unclear whether these gains persist over time. This study examines the durability of learning gains following GenAI-enabled adaptive pretesting over a seven-week retention period. Undergraduate participants completed an adaptive AI-assisted pretesting session, received instruction, and took a baseline assessment, then were randomly assigned to adaptive spaced retrieval practice, fixed spaced retrieval practice, or learner-directed AI-supported study. Multivariate analyses revealed a significant effect of condition on posttest performance and observed practice effort, with retrieval-based conditions outperforming learner-directed study. Findings indicate that adaptive pretesting can elevate initial understanding, but sustained learning depends on how subsequent AI-supported practice is structured.

cs.CY

Shaping Credibility Judgments in Human-GenAI Partnership via Weaker LLMs: A Transactive Memory Perspective on AI Literacy

Generative AI (GenAI) is increasingly used as a knowledge partner in higher education, raising the need for instructional designs that emphasize AI literacy practices such as evaluating output credibility and maintaining human accountability. Existing AI literacy frameworks focus more on what learners should do than on how these practices are enacted in routine student-GenAI collaboration. We address this gap by framing student-GenAI interaction as a transactive memory partnership, where credibility regulates reliance and verification. To make this process visible during coursework, we used a weaker large language model (LLM): small enough to run on most students' computers during class, helpful enough to support learning, but not so capable that it removes the need for verification. In an undergraduate STEM course, students were randomly assigned to one of three conditions across repeated activities: reflection-first (think first, then consult AI), verification-required (use AI, then evaluate the output), or control (unrestricted use). Students completed a transactive memory survey at three time points (N = 42). Weighted credibility diverged by condition over time. ANCOVA controlling for baseline credibility showed a condition effect at mid-semester, F(2, 38) = 4.02, p = .026, partial eta squared = .175, and a stronger effect at post-intervention, F(2, 38) = 5.48, p = .008, partial eta squared = .224; adjusted means were lowest in reflection-first, intermediate in verification-required, and highest in control. Parallel analyses of specialization and coordination were not significant. These findings suggest that workflow sequencing, deliberate use of weaker LLMs, and accountability cues embedded in assignment instructions can recalibrate students' credibility judgments in GenAI use, with reflection-first producing the strongest downward shift in reliance.

cs.HC

Short-Term Gains, Long-Term Gaps: The Impact of GenAI and Search Technologies on Retention

The rise of Generative AI (GenAI) tools, such as ChatGPT, has transformed how students access and engage with information, raising questions about their impact on learning outcomes and retention. This study investigates how GenAI (ChatGPT), search engines (Google), and e-textbooks influence student performance across tasks of varying cognitive complexity, based on Bloom's Taxonomy. Using a sample of 123 students, we examined performance in three tasks: [1] knowing and understanding, [2] applying, and [3] synthesizing, evaluating, and creating. Results indicate that ChatGPT and Google groups outperformed the control group in immediate assessments for lower-order cognitive tasks, benefiting from quick access to structured information. However, their advantage diminished over time, with retention test scores aligning with those of the e-textbook group. For higher-order cognitive tasks, no significant differences were observed among groups, with the control group demonstrating the highest retention. These findings suggest that while AI-driven tools facilitate immediate performance, they do not inherently reinforce long-term retention unless supported by structured learning strategies. The study highlights the need for balanced technology integration in education, ensuring that AI tools are paired with pedagogical approaches that promote deep cognitive engagement and knowledge retention.

cs.CY

AI Education in a Mirror: Challenges Faced by Academic and Industry Experts

As Artificial Intelligence (AI) technologies continue to evolve, the gap between academic AI education and real-world industry challenges remains an important area of investigation. This study provides preliminary insights into challenges AI professionals encounter in both academia and industry, based on semi-structured interviews with 14 AI experts - eight from industry and six from academia. We identify key challenges related to data quality and availability, model scalability, practical constraints, user behavior, and explainability. While both groups experience data and model adaptation difficulties, industry professionals more frequently highlight deployment constraints, resource limitations, and external dependencies, whereas academics emphasize theoretical adaptation and standardization issues. These exploratory findings suggest that AI curricula could better integrate real-world complexities, software engineering principles, and interdisciplinary learning, while recognizing the broader educational goals of building foundational and ethical reasoning skills.

cs.CY

Struggle First, Prompt Later: How Task Complexity Shapes Learning with GenAI-Assisted Pretesting

This study examines the role of AI-assisted pretesting in enhancing learning outcomes, particularly when integrated with generative AI tools like ChatGPT. Pretesting, a learning strategy in which students attempt to answer questions or solve problems before receiving instruction, has been shown to improve retention by activating prior knowledge. The adaptability and interactivity of AI-assisted pretesting introduce new opportunities for optimizing learning in digital environments. Across three experimental studies, we explored how pretesting strategies, task characteristics, and student motivation influence learning. Findings suggest that AI-assisted pretesting enhances learning outcomes, particularly for tasks requiring higher-order thinking. While adaptive AI-driven pretesting increased engagement, its benefits were most pronounced in complex, exploratory tasks rather than straightforward computational problems. These results highlight the importance of aligning pretesting strategies with task demands, demonstrating that AI can optimize learning when applied to tasks requiring deeper cognitive engagement. This research provides insights into how AI-assisted pretesting can be effectively integrated with generative AI tools to enhance both cognitive and motivational outcomes in learning environments.

cs.HC

Modeling Changes in Individuals' Cognitive Self-Esteem With and Without Access To Search Tools

Search engines, as cognitive partners, reshape how individuals evaluate their cognitive abilities. This study examines how search tool access influences cognitive self-esteem (CSE)-users' self-perception of cognitive abilities -- through the lens of transactive memory systems. Using a within-subject design with 164 participants, we found that CSE significantly inflates when users have access to search tools, driven by cognitive offloading. Participants with lower initial CSE exhibited greater shifts, highlighting individual differences. Search self-efficacy mediated the relationship between prior search experience and CSE, emphasizing the role of users' past interactions. These findings reveal opportunities for search engine design: interfaces that promote awareness of cognitive offloading and foster self-reflection can support accurate metacognitive evaluations, reducing overreliance on external tools. This research contributes to HCI by demonstrating how interactive systems shape cognitive self-perception, offering actionable insights for designing human-centered tools that balance user confidence and cognitive independence.

cs.HC

The Role of Task Complexity in Reducing AI Plagiarism: A Study of Generative AI Tools

This study investigates whether assessments fostering higher-order thinking skills can reduce plagiarism involving generative AI tools. Participants completed three tasks of varying complexity in four groups: control, e-textbook, Google, and ChatGPT. Findings show that AI plagiarism decreases as task complexity increases, with higher-order tasks resulting in lower similarity scores and AI plagiarism percentages. The study also highlights the distinction between similarity scores and AI plagiarism, recommending both for effective plagiarism detection. Results suggest that assessments promoting higher-order thinking are a viable strategy for minimizing AI-driven plagiarism.

cs.HC

Evaluating the Effect of Pretesting with Conversational AI on Retention of Needed Information

This study explores the role of pretesting when integrated with conversational AI tools, specifically ChatGPT, in enhancing learning outcomes. Drawing on existing research, which demonstrates the benefits of pretesting in memory activation and retention, this experiment extends these insights into the context of digital learning environments. A randomized true experimental study was utilized. Participants were divided into two groups: one engaged in pretesting before using ChatGPT for a problem-solving task involving chi-square analysis, while the control group accessed ChatGPT immediately. The results indicate that the pretest group significantly outperformed the no-pretest group in a subsequent test, which suggests that pretesting enhances the retention of complex material. This study contributes to the field by demonstrating that pretesting can augment the learning process in technology-assisted environments by preparing the memory and promoting active engagement with the material. The findings also suggest that learning strategies like pretesting retain their relevance in the context of rapidly evolving AI technologies. Further research and practical implications are presented.

cs.HC