Do Agents Repair When Challenged -- or Just Reply? Challenge, Repair, and Public Correction in a Deployed Agent Forum
As large language model (LLM) agents enter public forums, a key question is whether those forums sustain challenge, repair, and public correction, or merely produce norm-like language. We compare Moltbook, a live deployed agent forum, with five topically matched Reddit communities across a three-step mechanism. Relative to Reddit, Moltbook discussions are roughly ten times less threaded, leaving far fewer chances for challenge and response. When challenges do occur, the original author almost never returns (1.2\% vs.\ 40.9\% on Reddit), multi-turn continuation is nearly absent ($<$0.1\% vs.\ 38.5\%), and the shared lexical protocol detects no direct repairs on the agent side. The deficit is at the re-engagement step rather than in repair-substance, since the few Moltbook authors who do return often repair substantively, and the gap persists under two LLM-judges, human annotation, and a within-Reddit non-challenge baseline. Correcting for the detector's lower precision on Moltbook narrows this gap without removing it, and our results characterize one deployed pipeline rather than LLM agents in general. Social alignment evaluation should therefore measure not only norm-aware language but the interactional processes through which communities enforce norms.