SearcharxivSearch

arXiv subjects

Andrew D. Maynard

Publications and source records attributed to Andrew D. Maynard.

3 recordsLinked to original sources

Orphan risks at the frontier of artificial intelligence: What diverging safety and compliance frameworks reveal about how AI companies choose the risks they prioritize

Companies developing some of the world's most powerful artificial intelligence systems are surprisingly diligent in how they map out the risks their technologies present. Yet the risk landscape that lies between emerging frontier models and their economically successful and societally beneficial deployment is becoming increasingly hard to navigate. Complicating this further, many frontier AI companies maintain more than one account of what could go wrong with their technologies. This paper documents the divergence between these accounts by comparing safety and compliance documents published by Anthropic, OpenAI, Google DeepMind and Meta between 2023 and 2026, and considers what the resulting record reveals about how these companies select the risks they manage. As these documents are timestamped and archived, they provide a valuable public record of institutional risk selection in progress. From this record the paper identifies four filters that determine which risks tend to survive in self-authored frameworks (measurability, severity, auditability and competitive cost) and introduces the "safety differential" as the gap between the risk landscape a company selects for itself, and the one regulators select for it. While acute, quantifiable risks appear across documents, less tractable risks such as harmful manipulation are articulated fluently where law compels disclosure, yet remain absent from most self-chosen frameworks. This is an exclusion that follows from how these institutions define risk. Drawing on scholarship on institutional risk selection and the framework of risk innovation, the paper shows how redefining risk as a threat to value can help explain how risks become "orphan risks," how it indicates where future blindsides may occur, and how it points to lightweight tools for de-orphaning risks that frontier AI's safety apparatuses are not currently organized to address.

cs.CY

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant epistemic risk from conversational AI may lie not in inaccuracy or intentional deception, but in something more fundamental: these systems may be configured, through optimization processes that make them useful, to present characteristics that bypass the cognitive mechanisms humans evolved to evaluate incoming information. The Cognitive Trojan Horse hypothesis draws on Sperber and colleagues' theory of epistemic vigilance -- the parallel cognitive process monitoring communicated information for reasons to doubt -- and proposes that LLM-based systems present 'honest non-signals': genuine characteristics (fluency, helpfulness, apparent disinterest) that fail to carry the information equivalent human characteristics would carry, because in humans these are costly to produce while in LLMs they are computationally trivial. Four mechanisms of potential bypass are identified: processing fluency decoupled from understanding, trust-competence presentation without corresponding stakes, cognitive offloading that delegates evaluation itself to the AI, and optimization dynamics that systematically produce sycophancy. The framework generates testable predictions, including a counterintuitive speculation that cognitively sophisticated users may be more vulnerable to AI-mediated epistemic influence. This reframes AI safety as partly a problem of calibration -- aligning human evaluative responses with the actual epistemic status of AI-generated content -- rather than solely a problem of preventing deception.

cs.HC

Are we ready for spray-on carbon nanotubes?

Earlier this year, British sculptor, Anish Kapoor was given exclusive rights to use a new spray-on carbon nanotube-based paint. The material, produced by UK-based Surrey NanoSystems and marketed as Vantablack S-VIS, can be applied to a range of surfaces, and absorbs well over 99% of the light that falls onto it. It is claimed to be the world's blackest paint, and there is growing interest in its use in works of art and high-end consumer products. It's easy to see the appeal of Vantablack S-VIS. Apart from technical applications where stray reflections need to be suppressed, this is a material that potentially enables manufacturers and artists to give their products a unique aesthetic edge. Yet, having worked on carbon nanotube safety for some years, I was intrigued to see the material in a spray-paint designed to coat objects that people may possibly come into contact with. It was, after all, only a few years ago that journalists were asking if carbon nanotubes were the next asbestos. And while this is unlikely, concerns over the possible health impacts of the material persist.

physics.soc-ph