Searcharxiv⌕ Search

arXiv · 2609.26865

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Abstract

Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety. The code for our Safety Nudges extension is publicly available at https://github.com/jtbwedgwood/safety-nudges.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith. 2026-09-24. Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness. https://arxiv.org/abs/2609.26865

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

"If You're Not Doing It, Somebody Else Is": Active Negotiation and the Invisible Labor of Sustained LLM Use

Large language models (LLMs) have become fixtures of academic work even as their users describe them as degrading their writing, thinking, and skills. Dominant adoption frameworks read continued use as evidence of satisfaction, and cannot explain continued use of a distrusted tool. We interviewed 36 graduate student workers, balanced between English-as-a-foreign-language (EFL) and non-EFL speakers, and introduce the Active Negotiation framework: a model of sustained LLM use as a recurring cycle of risk, mitigation, and justification. A failure surfaces a risk, mitigation labor addresses it, and a justification renders the residual risk tolerable until the next failure reopens the cycle. The cycle runs across three dimensions: practical, auditing output; internal, auditing one's own cognition and identity; and social, managing how peers and institutions perceive use. EFL participants invoke linguistic parity as a further justification. We reframe continued adoption as compliance sustained by invisible labor.

cs.HC↗

Skill Profiling with Attributable Reasoning (SPAR): A Wearable Analysis System for Boxing

A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deployable systems only classify which punch was thrown rather than assess how well it was thrown. We present Skill Profiling with Attributable Reasoning (SPAR), an eight-IMU garment and pressure-insole system that classifies each punch as expert or novice and treats an explanation of that prediction as feedback. Feedback is only useful if the person receiving it can act on it, so SPAR explains the prediction at three tiers, a per-joint attribution for the analyst, a counterfactual over kinetic-chain layers for the coach, and a plain-language narrative of the two for the athlete. Across 17 participants and 4,713 punches, SPAR reaches a leave-one-participant-out AUC of 0.842 (95% CI [0.769, 0.907] over participants). A frozen time-series foundation model encodes the joint-angle and plantar-force series, and a small transformer trained on the cohort classifies the encoding. We audit the two quantitative tiers and report six themes from a thematic analysis of interviews with six practicing boxing coaches.

cs.HC↗

Sampling Safe Futures: Multimodal Trajectory Planning for Personalized Safety in Anthropomorphic AI

Anthropomorphic artificial intelligence systems increasingly remember personal details, display empathy, and are engaged with as social counterparts, creating forms of risk that emerge from the evolution of the user-system relationship over time. Existing safeguards largely operate at the level of individual conversational turns and cannot determine whether a sequence of seemingly acceptable interactions is cumulatively moving a particular user toward harm. This paper introduces personalized trajectory-level safety, a framework that treats relational safety as a sequential decision problem over a latent escalation state inferred from the user's messages and influenced by the system's responses. At each turn, a screening step first discards any response strategy that does not preserve at least one safe continuation of the interaction under every plausible model of the user. Among the remaining strategies, we formulate action selection as multimodal trajectory sampling, and use a Generative Flow Network to generate diverse future evolutions in proportion to their plausibility, safety, and utility. The system then selects the strategy that preserves the largest fraction of safe and useful continuations. We evaluate the framework in simulation, calibrated on statistics reported for real human-chatbot interactions, and using response strategies derived from public benchmarks. Results show that trajectory-aware decision making substantially reduces the frequency of harmful states while keeping helpful interaction. This work reframes safety for anthropomorphic AI from response-level filtering to personalized control over the future evolution of human-AI relationships. The source code is available at https://github.com/benedettapicano/ANTHROPOMORPHIC_SAFETY_TRAJ.

cs.HC↗