Skip to main content
Live
Main content
Review · Platforms
LMArena logo

LMArena

Editor rating
4.5/ 5
Starting price
$0
Free tier
Yes
Platforms
Web
Developer
LMArena (originally LMSYS / UC Berkeley)
Launched
2023

LMArena review

4.5 / 5By LMArena (originally LMSYS / UC Berkeley)Researched overview by AI Chat DailyUpdated Visit official site ↗
The verdict

LMArena is the most-watched AI model leaderboard in 2026 and one of the best free ways to compare frontier models. Not a perfect benchmark — human voting has biases — but remains the industry's best crowdsourced signal for model quality.

Try LMArenaOpens lmarena.ai

How this was put together. This is a researched overview, not a hands-on review — compiled by the AI Chat Daily desk from LMArena's own documentation, pricing pages and release notes, plus how the product has been received. The score reflects documented capability and market position rather than our own testing. Last checked Apr 24, 2026. No sponsorship, no affiliate relationship. Read our editorial standards and corrections policy.

LMArena (formerly Chatbot Arena, originally from LMSYS at UC Berkeley) launched in May 2023 as a research project to crowdsource AI model comparisons. Three years later, it has become the most-watched AI model leaderboard in the world — referenced by researchers, cited in blog posts, and optimized for by AI labs. In 2026, LMArena is both an indispensable tool for comparing frontier models and a subject of serious debate about what it actually measures.

This review covers LMArena in 2026: how it works, why it matters, the limitations, and how to use it effectively.

The good
  • Free access to compare nearly every major AI model
  • Blind voting reduces brand bias in comparisons
  • Most-watched leaderboard by AI labs and researchers
  • Covers text, vision, and code-specific models
  • Continuously updated with new models
  • Useful for individual users deciding which model to use
Watch out
  • Human vote biases influence results (formatting, length, style)
  • Doesn't directly measure truthfulness or accuracy
  • Models optimize specifically for Arena, potentially gaming it
  • Heavy weight on subjective preferences
  • Not a substitute for task-specific benchmarks
Best for
  • AI researchers comparing model capabilities
  • Developers choosing between LLMs for their applications
  • Users wanting to test multiple AI models side-by-side
  • AI enthusiasts tracking model progress over time
  • Writers and content creators comparing output quality
Avoid if
  • You need rigorous scientific benchmarks (use MMLU, HumanEval, etc.)
  • You need task-specific evaluation for a narrow use case
  • You prefer structured testing over crowdsourced opinion
  • You're sensitive to model quality variance in a single interaction

Pricing

Free
$0

Full access to compare models and vote. Community-supported service.

What LMArena is

LMArena is a web-based platform with two core features:

Battle Mode. You type a prompt. Two anonymous AI models (you don't know which) respond. You vote for the better response. Your vote contributes to the leaderboard Elo score of whichever model was which.

Side-by-Side Mode. You pick two specific models and compare them directly.

Direct Chat. Chat with a specific model (more limited).

Leaderboard. Rankings of models by Elo score derived from millions of user votes. Filters for overall, coding, math, vision, multi-turn, style control, and more.

The concept is simple but powerful. Aggregating millions of user preference votes produces a ranking that's more robust than any individual human judge could provide. The blind voting element reduces brand bias — you vote for quality of response, not because you love OpenAI or Anthropic.

LMArena is free to use. The operators are researchers; costs are covered by cloud credits from AI labs, research grants, and donations. There are no ads, subscriptions, or paywalls.

How the Elo rating works

LMArena uses an Elo-style rating system (similar to chess):

  • Each model starts at a baseline rating
  • When users vote, the winning model gains points and the losing model loses points
  • Points exchanged depend on the rating difference (upsets gain more)
  • Over many votes, ratings stabilize around actual user preference

This approach is more robust than percentage-based rankings because it handles unbalanced matchups. If GPT-4 plays against a weaker model and wins, it gains few points (expected win). If it loses to a weaker model, it loses many points (upset).

Over millions of votes, the Elo ratings converge on a stable ordering that reflects user preference reliably.

The leaderboard

The LMArena leaderboard in 2026 shows dozens of models with Elo scores. Filters let you see:

Overall. General performance across all prompt types.

Coding. Performance on coding-specific prompts.

Math. Math and reasoning performance.

Vision. Multimodal performance on image inputs.

Long Query. Performance on complex, long prompts.

Multi-turn. Performance in multi-turn conversations.

Style Control. Performance when controlling for style biases (newer feature addressing length/formatting concerns).

Models at the top typically include the latest GPT, Claude, Gemini, and their competitors. The top 5-10 positions change regularly as new model versions release.

Why AI labs care

LMArena has become influential enough that AI labs optimize specifically for it:

Marketing impact. "#1 on LMArena" generates coverage and drives user adoption. Labs celebrate when they top the leaderboard.

Benchmark for progress. Arena scores are a real signal that a new model is actually better than the previous version (or not).

Early signal. Labs often release preview versions to Arena before broader launch to gauge reception.

Benchmark hacking concerns. Some argue models are being trained specifically to win Arena votes rather than to be actually better, which creates measurement-vs-reality concerns.

The influence has grown to the point where Arena performance is sometimes prioritized over other benchmarks (MMLU, HumanEval, etc.) in marketing materials.

Known biases and limitations

LMArena isn't a perfect measure. Researchers (including Arena's own operators) have published analyses of biases:

Length bias. Longer responses tend to win. Not always because they're better — sometimes just because they feel more thorough.

Formatting bias. Markdown with bullet points, headers, and code blocks wins more often than prose responses. Some responses are better in prose but lose to formatted alternatives.

Style bias. Confident, assertive responses win more than hedged or uncertain ones. Even when the uncertain response is more accurate.

Recency bias. New models get a boost initially that may not persist.

User population bias. Arena voters are tech-forward, AI-enthusiastic users. Not representative of general populations.

Prompt population bias. People voting on Arena ask certain types of questions. Model performance on underrepresented prompts may differ.

Safety versus preference tension. Users sometimes prefer responses that are less safe (more willing to write creative content, more specific detail on sensitive topics). Labs must balance preference against safety guidelines.

To address some of these, LMArena has added "Style Control" filters that try to normalize length and formatting. The leaderboard with style control sometimes differs significantly from the default.

Using LMArena effectively

For individual users choosing an AI model, LMArena is useful but shouldn't be the only input:

Use the leaderboard to narrow the field. Top 10 models are all very capable.

Use task-specific filters (coding, math) if you care about specific capabilities.

Try Battle Mode on your actual use cases. What wins on average matters less than what wins on the prompts you'll actually send.

Check style-controlled rankings if you care about substantive quality over presentation.

Read recent Arena discussions to understand current biases and debates.

Cross-reference with other benchmarks (Artificial Analysis, Hugging Face, provider benchmarks).

How LMArena compares

Against MMLU, HumanEval, etc.: Academic benchmarks measure specific capabilities in controlled ways. More rigorous, but may not reflect real-world performance. LMArena measures actual user preference. Complementary.

Against Hugging Face Model Hub: HF has broader model coverage with community reviews and various benchmarks. Different approach. LMArena is more focused and widely referenced.

Against Artificial Analysis: Artificial Analysis provides technical benchmarks, speed, and pricing comparisons. More quantitative, less preference-based. Complements LMArena.

Against provider benchmarks: OpenAI, Anthropic, Google publish their own benchmark results. Marketing-influenced but useful for specific capabilities. LMArena is more neutral.

Against personal testing: Nothing beats testing models on your actual use cases. LMArena is a starting point, not the final answer.

Who should use LMArena

AI researchers tracking model development over time.

Developers choosing LLMs for their applications.

Users comparing models before subscribing to paid AI services.

AI enthusiasts curious about the current state of the field.

Content creators and writers comparing output quality.

Anyone wanting to try frontier models for free via Battle Mode.

Who should skip LMArena

Users needing rigorous benchmarks — use academic evaluations.

Users with narrow use cases — task-specific testing is more informative.

Users who dislike subjective rankings — Arena is inherently preference-based.

Users who already know their model preference — use what you like.

The verdict

LMArena is the most-watched AI model leaderboard in 2026 for good reason. The crowdsourced approach captures real user preference in a way other benchmarks don't. The free access makes it a useful tool for anyone comparing models. The transparency about methodology and biases is better than most proprietary benchmarks.

The limitations are real. Length bias, formatting bias, and benchmark hacking are genuine concerns. Arena scores aren't the same as objective quality — they measure what users prefer, which is correlated with but not identical to what's actually best.

For individual users and developers, LMArena is a valuable starting point. Check the overall leaderboard to see which models are considered top-tier. Use task-specific filters if you care about coding or math. Test models on your actual use cases before committing. Cross-reference with other benchmarks.

For the AI industry, LMArena has become genuinely influential — enough that labs optimize for it and marketing materials reference Arena rank. That influence has both positive effects (accountability and public transparency about model quality) and negative ones (benchmark hacking concerns). It's the best we have in terms of crowdsourced model comparison, and it deserves to be used thoughtfully.

<!-- p2-p1: best crowdsourced ai benchmarking platform -->

Is LMArena the best crowdsourced AI benchmarking platform?

By active reach, LMArena is the most-used crowdsourced AI benchmarking platform in 2026. Over a million blind votes have been cast since launch; every frontier lab from OpenAI and Anthropic to xAI and Google submits models to it; the leaderboard is the de-facto industry signal that a new release is "real."

That said, the term "best" depends on what you're measuring. Arena scores capture user preference, not raw capability. A model that produces longer, more confident-sounding answers can outscore a model that's actually more accurate but blunter — the well-documented length bias. Format bias compounds it: nicely-formatted markdown wins over correct-but-plain answers. And the major labs now optimize specifically for Arena performance, which means recent score jumps reflect targeted training as much as underlying capability gains.

The honest answer: LMArena is the best single signal of how regular users perceive AI quality. It's not the best signal of pure capability, where benchmarks like GPQA Diamond and SWE-bench Verified are stricter measures. The right move is to triangulate — read the Arena leaderboard for vibe, but cross-reference Artificial Analysis or HELM for objective measurements before betting a workflow on a model. Most serious AI evaluators in 2026 do exactly this.

Frequently asked questions

What is LMArena?
LMArena (formerly Chatbot Arena) is a platform where users can chat with two anonymous AI models side-by-side and vote for the better response. Votes feed into an Elo-style leaderboard that ranks models by user preference. Started by UC Berkeley researchers, now a standard industry reference.
Is LMArena a reliable benchmark?
For capturing user preference, yes. For objective accuracy or safety, no. LMArena measures what users like, which correlates with quality but also with style, length, and formatting preferences. Use it alongside other benchmarks.
Why do AI labs care about LMArena?
The leaderboard drives perception of model quality. A top Arena rank generates positive coverage and attracts users. Labs optimize their models to perform well on Arena. This has led to concerns about benchmark hacking.
Which model is #1 on LMArena?
Rankings change regularly. Historically the top has rotated between GPT-4, Claude 3.5+, and Gemini Pro with various competitor models moving in and out. Check the current leaderboard at lmarena.ai for up-to-date results.
Is LMArena free?
Yes — completely free. Community-supported by research institutions and occasional grants. You can chat with any model, compare any two models, and vote without subscription.
What biases affect LMArena voting?
Length bias (longer responses often win), formatting bias (well-formatted Markdown wins), style bias (confident tone wins), recency bias (newer models get initial boost), and user experience biases. Researchers are aware of these and publish analyses.
Can I use LMArena to pick the right AI for my use case?
As one input, yes. Top-ranked models are generally high quality. But for specific tasks (coding, math, creative writing), task-specific benchmarks matter more than the overall Arena score. Use Arena to narrow the field, then test specifically.
How do models get added to LMArena?
LMArena curators add new models as they're released, typically major releases from AI labs. Smaller models or fine-tunes may be added through community contributions. Not all models are represented.
Explore further
AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More on LMArena

Arena hits $100M run-rate eight months after launching paid evaluations
Business

Arena hits $100M run-rate eight months after launching paid evaluations

The UC Berkeley-born AI leaderboard tripled revenue from $30M in January, competing for post-training dollars against Scale AI, Mercor, and Surge.

Jaeden Schafer5 min read