arXiv · 2510.19738
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
Abstract
Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded. This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step.
Explore related subjects
Keep this discovery
Rustem Turtayev, Natalia Fedorova, Oleg Serikov, Sergey Koldyba, Lev Avagyan, Dmitrii Volkov. 2025-10-22. Misalignment Bounty: Crowdsourcing AI Agent Misbehavior. https://arxiv.org/abs/2510.19738
Cite the original work for its findings. Save a collection to share your selection of sources.