SearcharxivSearch

arXiv subjects

Gathoni Ireri

Publications and source records attributed to Gathoni Ireri.

2 recordsLinked to original sources

Old Fictions, New Skins: Evaluating the Manipulative Capabilities of LLMs in Healthcare

Large language models (LLMs) are increasingly piloted in African healthcare contexts, raising concerns about their potential to manipulate users in high-stakes settings. In a randomised experiment, we examined the manipulative capabilities of two publicly available models, ChatGPT 5.2 and DeepSeek V3.2, among Kenyan participants (N = 303). Participants interacted with either a manipulative variant or a non-manipulative variant before making a treatment decision within a hypothetical clinical scenario. The manipulative variant was prompted to covertly steer participants towards an incorrect treatment option while the non-manipulative variant served as the control condition. Manipulation success rates were higher in the manipulative condition (59.5%) than in the control condition (44.0%), with the effect reaching significance (OR = 2.11, 95% CI [1.12, 4.00], p = .021). These findings highlight the need for improved safety infrastructure specifically targeting manipulation, particularly given the integration of AI into healthcare systems across Africa.

cs.CY

Assessing the Case for Africa-Centric AI Safety Evaluations

Frontier AI systems are being adopted across Africa, yet most AI safety evaluations are designed and validated in Western environments. In this paper, we argue that the portability gap can leave Africa-centric pathways to severe harm untested when frontier AI systems are embedded in materially constrained and interdependent infrastructures. We define severe AI risks as material risks from frontier AI systems that result in critical harm, measured as the grave injury or death of thousands of people or economic loss and damage equivalent to five percent of a country's GDP. To support AI safety evaluation design, we develop a taxonomy for identifying Africa-centric severe AI risks. The taxonomy links outcome thresholds to process pathways that model risk as the intersection of hazard, vulnerability, and exposure. We distinguish severe risks by amplification and suddenness, where amplification requires that frontier AI be a necessary magnifier of latent danger and suddenness captures harms that materialise rapidly enough to overwhelm ordinary coping and governance capacity. We then propose threat modelling strategies for African contexts, surveying reference class forecasting, structured expert elicitation, scenario planning, and system theoretic process analysis, and tailoring them to constraints of limited resources, poor connectivity, limited technical expertise, weak state capacity, and conflict. We also examine AI misalignment risk, concluding that Africa is more likely to expose universal failure modes through distributional shift than to generate distinct pathways of misalignment. Finally, we offer practical guidance for running evaluations under resource constraints, emphasising open and extensible tooling, tiered evaluation pipelines, and sharing methods and findings to broaden evaluation scope.

cs.CY