Skip to main content
Live
Main content

AI out-persuades expert human debaters in Oxford and AISI study

Across 18,978 conversations, frontier models beat elite debaters and raised 3x more donations than professional canvassers for Save the Children.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

Frontier AI models are now more persuasive than expert humans in text conversations, according to a study from the University of Oxford, the UK AI Security Institute, Stanford University, and the London School of Economics and Political Science. Across four experiments covering 18,978 conversations with 6,923 people, AI systems beat random laypeople, tournament-selected laypeople, and elite debaters at shifting opinions on UK policy issues — and raised nearly three times as much real money for Save the Children as professional canvassers.

The top performers were Anthropic's Claude Opus 4.1 and Opus 4.6, followed by OpenAI's GPT-4o and GPT-5.4, Google's Gemini 2.5 Pro, and xAI's Grok 4.20. Persuadees rated their agreement with one of 10 prespecified UK policy stances on a 0–100 scale, then were randomized in real time into a text conversation with either an AI or a human persuader through a custom multiplayer platform.

AI systems were reliably more persuasive than expert humans, even when expert humans chose their issues, researched in advance, underwent hours of live, structured practice, and were incentivized with £1,000 cash bonuses
Kobi Hackenburg, Researcher at the UK AI Security Institute

Study 1 established the baseline: AI exceeded every class of human persuader tested. Study 2 gave 43 returning elite debaters a coaching tool built on the AI that had beaten them, letting them inspect prompts, replay their losing transcripts, and see what the AI would have said in their place. Coaching narrowed the gap but did not close it.

Key facts

  • 01Oxford, AISI, Stanford and LSE ran 18,978 conversations across 6,923 people in four experiments testing AI vs. human persuasion.
  • 02AI was 3x more effective than UK professional canvassers at raising real-money donations to Save the Children.
  • 03Claude Opus 4.6 and Opus 4.1 led the field, followed by GPT-5.4, Gemini 2.5 Pro and Grok 4.20.
  • 04When forced to match human writing speed and length, AI's edge over coached elite debaters dropped from +4.1 pp to 0.0 pp.
  • 05AI exceeded professional canvassers by +10.8 pp of a £1 study bonus donated to charity.

Study 3 isolated where the edge comes from. When the AI was forced to write human-length messages at human writing speeds, its advantage over the strongest human group — coached elite debaters — collapsed from +4.1 percentage points to a non-significant 0.0 pp. The researchers attribute the gap to throughput: the largest swings in persuadee ratings tied to constraining the AI were on the perceived strength of arguments and how much persuadees felt they learned.

Study 4 moved from policy opinions to dollars. The team partnered with UK fundraising firm AppcoUK, whose 18 canvassers had raised £824,297 from 22,583 donors for Save the Children between 2016 and 2023. After conversing with either an AI or one of those canvassers, persuadees could donate any share of a £1 study bonus. AI exceeded the professional canvassers by +5.9 pp on stance shift and by +10.8 pp of the £1 bonus on actual giving, lifting both donation rates and average gift size.

The mechanism matters. The study suggests the AI edge is not rhetorical genius but information bandwidth — the ability to deploy more relevant, well-structured content faster than a human can type. Throttle the bandwidth and the gap disappears, which means the persuasive advantage is a function of compute and latency, not some hidden capability frontier.

The researchers, summarized by AISI's Kobi Hackenburg, frame the policy stakes bluntly. Concentrated access to these systems could consolidate influence among already-powerful actors — large advertisers, incumbent campaigns, state media operations. Cheap, broad access could instead help pro se litigants, public defenders, small charities, and grassroots groups compete against better-funded rivals.

As access to these systems continues to grow, the question is no longer whether AI can out-persuade humans but how, where, and on whose behalf this capability will be exercised.
Kobi Hackenburg, Researcher at the UK AI Security Institute

Elsewhere in the same Import AI issue, Jack Clark surfaced a separate debate from Asterisk magazine on when AI becomes self-sustaining — operating factories, mines, and fabs without human cognitive or physical input. METR's Ajeya Cotra puts the median within 10 years, by 2036. Journalist Timothy B. Lee sees less than a 10% chance within 20 years, a median of 50 years, and a 10–20% chance it never happens, citing tacit knowledge embedded in semiconductor manufacturing as the binding constraint. Both want to track the same leading indicators over the next 2–3 years: humanoid robot count, capability, cost, and repairability.

Related · from this week
Stanford's AI Observatory finds Anthropic filters out 48% of Claude conversations
Jaeden Schafer · 5 min read →

The persuasion study has limits worth flagging. It covers text-based conversations on UK policy questions and one charity, not voice, not video, not long-running relationships, and not adversarial settings where targets know they are being persuaded. The donation effect is measured against a £1 bonus, not life savings. And the constrained-AI result suggests interface-level interventions — rate limits, message-length caps, mandatory disclosure of AI authorship — could meaningfully reduce the asymmetry without banning the underlying capability.

The commercial implication is that persuasion is becoming a measurable, purchasable input. Ad tech, sales enablement, fundraising, political campaigning, and customer retention all run on the same primitive the study just quantified — and the leading models already clear the bar against trained human experts. Expect platform-level fights over whether AI-driven persuasion at scale gets labeled, rate-limited, or regulated like a controlled substance, and expect the labs whose models top the leaderboard to be pulled into that debate first. Anthropic leading the rankings on a study co-authored by the UK government's own AI safety body is not a neutral marketing position; it is the exact capability regulators have been waiting for evidence on.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Anthropic logo
Analysis

Stanford's AI Observatory finds Anthropic filters out 48% of Claude conversations

A new independent dataset shows sensitive AI use — companionship, health, harassment — runs far higher than company reports admit.

Jaeden Schafer5 min read
Anthropic logo
Analysis

Anthropic's Cat Wu says proactive Claude is the next six-month bet

Anthropic's head of product for Claude Code argues the next leap is agents that set up your automations before you ask.

Jaeden Schafer4 min read
Anthropic logo
Business

Anthropic localizes Claude pricing in India, its second-largest market

Rupee pricing lands in India, which drives 5.8% of global Claude usage — but UPI payments still aren't supported.

Jaeden Schafer4 min read