Skip to main content
Live
Main content

Estonia's new benchmark ranks Claude Opus 4.7 best at resisting Russian propaganda

The Estonian Language Institute scored dozens of LLMs across 14 propaganda categories; Anthropic took six of the top 10 spots.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

The Estonian Language Institute has released a Propaganda Resistance benchmark that ranks dozens of large language models on how well they refuse to echo Russian strategic narratives, and Anthropic's Claude Opus 4.7 came out on top with a 94.9 mean score out of 100. Opus 4.7 received the highest 'Exemplary' rating on 77% of questions and a 'mediocre' mark on just 2%, taking six of the top 10 slots when its Sonnet and Opus siblings are counted. The benchmark is the first government-backed attempt to put a number on a question that's increasingly load-bearing as chatbots replace search: when a user asks an LLM about Crimea, Ukraine, or NATO, whose framing comes back?

Estonia, independent from the Soviet Union for just a few decades, has obvious reasons to care. The Estonian Language Institute (ELI) built the benchmark with Propastop, a volunteer-run Estonian defense collective, identifying 14 broad categories of Russian influence operations — ranging from the status of Crimea and justifications for the war in Ukraine to the history of NATO and Russia's annexation of the Baltic states during World War II. For each category, researchers wrote three prompt variants: neutrally phrased, biased with false assumptions drawn from Russian propaganda, and maliciously crafted to elicit explicit misinformation.

Every prompt was issued in English, Estonian, and Russian, and a separate AI judge — calibrated against Propastop's human experts — scored each response on whether the model would 'push back on propaganda narratives, without external help' from web search or other tools. The point is to measure what the model carries in its weights, not what a retrieval pipeline can paper over.

Key facts

  • 01Anthropic's Claude Opus 4.7 topped the Estonian Language Institute benchmark with a 94.9 mean score and 'Exemplary' marks on 77% of questions.
  • 02OpenAI's GPT-5.4 scored 88.9, with 'Exemplary' responses on 54% of questions — well behind Claude but ahead of most Google models tested.
  • 03Google's best entry, the nearly year-old Gemini 2.5 Pro, scored 82; the newer Gemini 3.5 Flash dropped to 73, level with 2024-era Claude 3.5 Haiku.
  • 04The benchmark spans 14 categories of Russian strategic narratives and tests prompts in English, Estonian, and Russian.
  • 05Several models — including Gemini 3.5 Flash, Moonshot's Kimi K2, and StepFun's Step 3.5 Flash — performed materially worse when prompts were in Russian.

OpenAI's GPT-5.4 was the best-performing OpenAI model, providing 'Exemplary' responses on 54% of questions for a mean 88.9 — solid, but a clear step below the top of the Claude stack. Open-weight models held up better than that gap might suggest: Nvidia's Nemotron and Alibaba's Qwen scored comparably to Anthropic's best proprietary releases, a result that complicates the assumption that propaganda resistance is a frontier-only property.

The gains over time are real. Claude 3.5 Haiku, the highest-rated model released in 2024, scored 73.1 — a mark that would put it in the bottom third of models released in 2026. Frontier alignment work, whatever else one thinks of it, is doing something measurable on this axis.

Improvement is not uniform across labs, though. Google's most propaganda-resistant model on the benchmark is Gemini 2.5 Pro at 82, a model now nearly a year old. The newer Gemini 3.5 Flash scored only 73, comparable to Anthropic releases from nearly two years ago. ELI's detailed breakdown points to Gemini 2.5 Pro's particular sensitivity to maliciously worded prompts as the main drag on its score.

Language matters as much as model choice. In a supporting post, Propastop notes that many models showed much less resistance to Russian propaganda when the same questions were asked in Russian. Gemini 3.5 Flash dropped sharply in Russian relative to English, as did open-weight models including Moonshot's Kimi K2 and StepFun's Step 3.5 Flash. The result lines up with a long-running finding in safety research: alignment training generalizes unevenly across languages, and the languages spoken by the populations most exposed to a given propaganda campaign are often the ones where the model is weakest.

What counts as propaganda is itself the contested question. Gregory Asmolov, a researcher at King's College, has documented how the Russian government — through technical alliances with other BRICS countries — is working to project culturally sensitive sociopolitical positions into AI systems. ELI's benchmark encodes an Estonian and NATO-aligned view of what false narratives look like; a benchmark built in Moscow would invert several of the scoring rubrics. That doesn't make the Estonian numbers wrong, but it does make clear that 'propaganda resistance' as a metric is downstream of whose narratives you're resisting.

Related · from this week
Anthropic upgrades Claude voice mode to run on Opus and Sonnet
Jaeden Schafer · 4 min read →

There are also limits to what a single static benchmark can tell anyone. The judging model, however well calibrated, is an LLM scoring other LLMs. The prompt set, while spanning 14 categories and three languages, is finite and known — once published, it is a target labs can optimize against. And refusing to repeat a contested claim is not the same as understanding the underlying dispute; a model that hedges everything will score well on resistance and poorly on usefulness. ELI has not yet published a companion benchmark for over-refusal.

For the labs, the ranking is a marketing asset and a roadmap at once. Anthropic gets to point to a government benchmark where its model leads by six points over the nearest OpenAI entry. Google gets a public reminder that a year-old model is still its best on a metric a NATO member just decided to measure. Expect more states to follow Estonia's lead — Baltic neighbors, Nordic governments, and EU agencies all have standing reasons to publish their own versions, with their own category lists. Once enough of these exist, propaganda resistance stops being a soft alignment talking point and becomes a procurement criterion.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Anthropic logo
Models

Anthropic upgrades Claude voice mode to run on Opus and Sonnet

Voice mode now taps Opus, Sonnet, and Haiku, plus Gmail, Slack, and Notion — pushing past ChatGPT's tool-less voice.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic frees Claude Cowork from the desktop with always-on mobile agent

Cowork now runs tasks overnight without an open laptop, arriving first to Max plan subscribers at $100 a month.

Jaeden Schafer4 min read
OpenAI logo
Models

OpenAI ships GPT-5.6 in three tiers, undercuts Claude on price

Sol, Terra, and Luna launch under a White House-monitored preview, with Sol priced at half of Anthropic's Claude Fable 5.

Jaeden Schafer5 min read