SearcharxivSearch

arXiv subjects

Don Hickerson

Publications and source records attributed to Don Hickerson.

3 recordsLinked to original sources

A Peek Behind the Curtain: Using Step-Around Prompt Engineering to Identify Bias and Misinformation in GenAI Models

Step-around prompting is a form of adversarial prompt engineering in which a user strategically reframes, sequences, or contextualises requests to test whether a generative AI model's safety guardrails, alignment mechanisms, or bias mitigations can be undermined, inconsistently applied, or bypassed outright. This study examines this technique through the lens of academic ethics, situating it as a tool that has a clear impact on academic integrity, responsible conduct of research, duty of care to students, and institutional oversight of GenAI use in higher education. We argue that step-around prompting is one tool within the wider practice of audit, red-teaming, and institutional evaluation, and that its main value lies in documenting how representational, cultural, linguistic, disciplinary, and misinformation-related biases may appear across student-facing and research-facing uses of GenAI. To show why the ethical governance of this practice is required, we provide two illustrative examples of the technique in action, demonstrating how easily guardrails can be circumvented and what is at stake when they are. We clarify which bias categories are in scope and identify who should use the method and for what purposes. We conclude with an operational ethics-and-governance framework for controlled academic application, organised as two pillars (technical safeguards and ethical governance) and enacted through a decision and audit cycle that scales oversight to potential risk, grounded in harm minimisation, duty of care, transparency, proportionality, responsible disclosure, legal and contractual compliance, and student protection.

cs.CY

GenAI Detection Tools, Adversarial Techniques and Implications for Inclusivity in Higher Education

This study investigates the efficacy of six major Generative AI (GenAI) text detectors when confronted with machine-generated content that has been modified using techniques designed to evade detection by these tools (n=805). The results demonstrate that the detectors' already low accuracy rates (39.5%) show major reductions in accuracy (17.4%) when faced with manipulated content, with some techniques proving more effective than others in evading detection. The accuracy limitations and the potential for false accusations demonstrate that these tools cannot currently be recommended for determining whether violations of academic integrity have occurred, underscoring the challenges educators face in maintaining inclusive and fair assessment practices. However, they may have a role in supporting student learning and maintaining academic integrity when used in a non-punitive manner. These results underscore the need for a combined approach to addressing the challenges posed by GenAI in academia to promote the responsible and equitable use of these emerging technologies. The study concludes that the current limitations of AI text detectors require a critical approach for any possible implementation in HE and highlight possible alternatives to AI assessment strategies.

cs.CY

Game of Tones: Faculty detection of GPT-4 generated content in university assessments

This study explores the robustness of university assessments against the use of Open AI's Generative Pre-Trained Transformer 4 (GPT-4) generated content and evaluates the ability of academic staff to detect its use when supported by the Turnitin Artificial Intelligence (AI) detection tool. The research involved twenty-two GPT-4 generated submissions being created and included in the assessment process to be marked by fifteen different faculty members. The study reveals that although the detection tool identified 91% of the experimental submissions as containing some AI-generated content, the total detected content was only 54.8%. This suggests that the use of adversarial techniques regarding prompt engineering is an effective method in evading AI detection tools and highlights that improvements to AI detection software are needed. Using the Turnitin AI detect tool, faculty reported 54.5% of the experimental submissions to the academic misconduct process, suggesting the need for increased awareness and training into these tools. Genuine submissions received a mean score of 54.4, whereas AI-generated content scored 52.3, indicating the comparable performance of GPT-4 in real-life situations. Recommendations include adjusting assessment strategies to make them more resistant to the use of AI tools, using AI-inclusive assessment where possible, and providing comprehensive training programs for faculty and students. This research contributes to understanding the relationship between AI-generated content and academic assessment, urging further investigation to preserve academic integrity.

cs.CY