SearcharxivSearch

arXiv subjects

Ben M. Tappin

Publications and source records attributed to Ben M. Tappin.

7 recordsLinked to original sources

When Large Language Models are More PersuasiveThan Incentivized Humans, and Why

Large Language Models (LLMs) have been shown to be highly persuasive, but when and why they outperform humans is still an open question. We compare the persuasiveness of two LLMs (Claude 3.5 Sonnet and DeepSeek v3) against humans who had incentives to persuade, using an interactive, real-time conversational setting. We demonstrate that LLMs persuasive superiority is context-dependent: it depends on whether the persuasion attempt is truthful (towards the right answer) or deceptive (towards the wrong answer) and on the LLM model, and wanes over repeated interactions (unlike human persuasiveness). In our first large-scale experiment, humans vs LLMs (Claude 3.5 Sonnet) interacted with other humans who were completing an online quiz for a reward, attempting to persuade them toward a given (either correct or incorrect) answer. Claude was more persuasive than incentivized human persuaders both in truthful and deceptive contexts and it significantly increased accuracy if persuasion was truthful, but decreased it if persuasion was deceptive. In a follow-up experiment with Deepseek v3, we replicated the findings about accuracy but found greater LLM persuasiveness only if the persuasion was deceptive. Linguistic analyses of the persuaders texts suggest that these effects may be due to LLMs expressing higher conviction than humans.

cs.CL

AI systems out-persuade expert humans

Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether it can out-persuade skilled and highly incentivized humans has remained unclear. Here, in a series of four preregistered experiments (n = 18,978 conversations from 6,923 people), we pitted AI systems against a range of human persuaders, including laypeople, winners of a separately preregistered four-round online persuasion tournament, professional canvassers, and world championship debaters. We found that AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses. In a follow-up study, AI's advantage persisted after experts received a coaching tool that let them practice against the AI that beat them, review their performance history, and see what AI would have said at key moments. We found converging evidence that AI's advantage stemmed from rapidly deploying larger quantities of information: after coaching, expert humans could tie an AI constrained to respond at human speeds and with human-length messages. In a final study, we show that AI's advantage extends to consequential real-world behavior: AI was nearly 3x more effective than professional canvassers from a UK fundraising firm at raising real-money donations to Save the Children. Together, these results establish that frontier AI systems out-persuade expert humans in conversation, with significant implications for political communication.

cs.CY

Artificial intelligence can persuade people to take political actions

There is substantial concern about the ability of advanced artificial intelligence to influence people's behaviour. A rapidly growing body of research has found that AI can produce large persuasive effects on people's attitudes, but whether AI can persuade people to take consequential real-world actions has remained unclear. In two large preregistered experiments N=17,950 responses from 14,779 people), we used conversational AI models to persuade participants on a range of attitudinal and behavioural outcomes, including signing real petitions and donating money to charity. We found sizable AI persuasion effects on these behavioural outcomes (e.g. +19.7 percentage points on petition signing). However, we observed no evidence of a correlation between AI persuasion effects on attitudes and behaviour. Moreover, we replicated prior findings that information provision drove effects on attitudes, but found no such evidence for our behavioural outcomes. In a test of eight behavioural persuasion strategies, all outperformed the most effective attitudinal persuasion strategy, but differences among the eight were small. Taken together, these results suggest that previous findings relying on attitudinal outcomes may generalize poorly to behaviour, and therefore risk substantially mischaracterizing the real-world behavioural impact of AI persuasion.

cs.CY

The Levers of Political Persuasion with Conversational AI

There are widespread fears that conversational AI could soon exert unprecedented influence over human beliefs. Here, in three large-scale experiments (N=76,977), we deployed 19 LLMs-including some post-trained explicitly for persuasion-to evaluate their persuasiveness on 707 political issues. We then checked the factual accuracy of 466,769 resulting LLM claims. Contrary to popular concerns, we show that the persuasive power of current and near-future AI is likely to stem more from post-training and prompting methods-which boosted persuasiveness by as much as 51% and 27% respectively-than from personalization or increasing model scale. We further show that these methods increased persuasion by exploiting LLMs' unique ability to rapidly access and strategically deploy information and that, strikingly, where they increased AI persuasiveness they also systematically decreased factual accuracy.

cs.CL

Evidence of a log scaling law for political persuasion with large language models

Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10 U.S. political issues from 24 language models spanning several orders of magnitude in size. We then deploy these messages in a large-scale randomized survey experiment (N = 25,982) to estimate the persuasive capability of each model. Our findings are twofold. First, we find evidence of a log scaling law: model persuasiveness is characterized by sharply diminishing returns, such that current frontier models are barely more persuasive than models smaller in size by an order of magnitude or more. Second, mere task completion (coherence, staying on topic) appears to account for larger models' persuasive advantage. These findings suggest that further scaling model size will not much increase the persuasiveness of static LLM-generated messages.

cs.CL

Does observability amplify sensitivity to moral frames? Evaluating a reputation-based account of moral preferences

A growing body of work suggests that people are sensitive to moral framing in economic games involving prosociality, suggesting that people hold moral preferences for doing the "right thing". What gives rise to these preferences? Here, we evaluate the explanatory power of a reputation-based account, which proposes that people respond to moral frames because they are motivated to look good in the eyes of others. Across two pre-registered experiments (total N = 3,610), we investigated whether reputational incentives amplify sensitivity to framing effects. Both experiments manipulated (i) whether moral or neutral framing was used to describe a Trade-Off Game (in which participants chose between prioritizing equality or efficiency) and (ii) whether Trade-Off Game choices were observable to a social partner in a subsequent Trust Game. We find that framing effects are relatively insensitive to reputational incentives: observability did not significantly amplify sensitivity to moral framing. However, our results are not inconsistent with the possibility that observability has some amplification effect; quantitatively, the observed framing effect was 74% as large when decisions were private as when they were observable. These results suggest that moral frames may tap into moral preferences that are relatively deeply internalized, and that power of moral frames to promote prosociality may not be strongly enhanced by making the morally-framed behavior observable to others.

physics.soc-ph

Doing good vs. avoiding bad in prosocial choice: A refined test and extension of the morality preference hypothesis

Prosociality is fundamental to human social life, and, accordingly, much research has attempted to explain human prosocial behavior. Capraro and Rand (Judgment and Decision Making, 13, 99-111, 2018) recently provided experimental evidence that prosociality in anonymous, one-shot interactions (such as Prisoner's Dilemma and Dictator Game experiments) is not driven by outcome-based social preferences - as classically assumed - but by a generalized morality preference for "doing the right thing". Here we argue that the key experiments reported in Capraro and Rand (2018) comprise prominent methodological confounds and open questions that bear on influential psychological theory. Specifically, their design confounds: (i) preferences for efficiency with self-interest; and (ii) preferences for action with preferences for morality. Furthermore, their design fails to dissociate the preference to do "good" from the preference to avoid doing "bad". We thus designed and conducted a preregistered, refined and extended test of the morality preference hypothesis (N=801). Consistent with this hypothesis, our findings indicate that prosociality in the anonymous, one-shot Dictator Game is driven by preferences for doing the morally right thing. Inconsistent with influential psychological theory, however, our results suggest the preference to do "good" was as potent as the preference to avoid doing "bad" in this case.

physics.soc-ph