Searcharxiv⌕ Search

arXiv subjects

Kweku Yamoah

Publications and source records attributed to Kweku Yamoah.

3 recordsLinked to original sources

Invisible Agents, Uninformed Patients: Towards Responsible Deployment Of Autonomous AI Diagnostic Agents In Sub-Saharan Africa

Autonomous AI diagnostic agents, systems that analyse patient-specific clinical data and produce diagnostic outputs or triage decisions without mandatory real-time human review, are increasingly deployed across eHealth platforms in sub-Saharan Africa at a pace that has outrun the governance infrastructure needed to oversee them. While significant bodies of work address AI accountability, transparency and explainability in healthcare, existing frameworks are largely clinician-centered and assume regulatory conditions that do not uniformly exist in low-resource settings. A patient-centered analysis of the disparity in patient awareness regarding autonomous agents, which results in a structural accountability gap, is mostly missing from the literature. This paper synthesizes existing research on informed consent, algorithmic accountability, and explainable AI to highlight three distinct challenges introduced by deploying AI agents in the sub-Saharan African context. Drawing on three documented deployment cases, including computer-aided tuberculosis detection in Tanzania, diabetic retinopathy and TB screening in Zambia, and mobile health chat-bot triage in Ghana, it demonstrates that these gaps are already present in active deployments across the region. In response, the paper proposes three foundational principles; agent-aware informed consent, human override as a structural requirement and contextually adapted explainability. This triad of principles lays a practical minimum standard for developers, health system administrators and policymakers in contexts where formal AI regulation remains nascent.

cs.CY↗

AI-Assisted Data Extraction for Systematic Reviews in Education

Systematic reviews are time-consuming endeavors that require knowledgeable human reviewers to screen studies for relevance and extract data following a specific coding scheme before any analysis or synthesis can occur. Large language models (LLMs) hold promise for substantially accelerating this process and reducing reviewer workload, yet their application within the context of systematic reviews in the field of education remains underexplored. We address this issue in two ways: through empirical studies and the iterative development of an open-source software tool. First, we conducted two empirical studies examining the efficacy of using LLMs for data extraction using data from a published review on pedagogical agents. We extracted a variety of data types from 112 studies and compared the results to data extracted by human coding. Results indicate that LLMs struggled with extracting data accurately and therefore are not ready to be used as primary data extraction tools without explicit human validation of the data extracted. These findings highlight the dire need for a human-in-the-loop (HIL) approach to AI-assisted data extraction. We then propose a HIL workflow and introduce and describe the development of a free, web-based, open-source tool designed to support user-friendly, human-validated data extraction with LLMs.

cs.HC↗

Fine-Tuning A Large Language Model for Systematic Review Screening

Systematic reviews traditionally have taken considerable amounts of human time and energy to complete, in part due to the extensive number of titles and abstracts that must be reviewed for potential inclusion. Recently, researchers have begun to explore how to use large language models (LLMs) to make this process more efficient. However, research to date has shown inconsistent results. We posit this is because prompting alone may not provide sufficient context for the model(s) to perform well. In this study, we fine-tune a small 1.2 billion parameter open-weight LLM specifically for study screening in the context of a systematic review in which humans rated more than 8500 titles and abstracts for potential inclusion. Our results showed strong performance improvements from the fine-tuned model, with the weighted F1 score improving 80.79% compared to the base model. When run on the full dataset of 8,277 studies, the fine-tuned model had 86.40% agreement with the human coder, a 91.18% true positive rate, a 86.38% true negative rate, and perfect agreement across multiple inference runs. Taken together, our results show that there is promise for fine-tuning LLMs for title and abstract screening in large-scale systematic reviews.

cs.CL↗