Vellum is an AI evaluation platform — companies use it to run structured tests against their LLM-powered features. Their public leaderboard is, in effect, a free preview of that product: the same kind of structured comparison they sell to enterprise customers, run on the frontier commercial models, published openly.
That framing is useful because it explains both the strengths and the limits.
- Curated model selection — only frontier models, no clutter
- Documented methodology and clear scoring criteria
- Pairs benchmark scores with editorial commentary
- Updated reasonably promptly after major model releases
- Good entry point for buyers who don't want to wade through MMLU minutiae
- Smaller catalog — open-source models often missing
- Vellum is a vendor; methodology choices reflect their product priorities
- Updates lag the major benchmark sites by days to weeks
- Doesn't break out hosting-provider differences for open models
- Less data behind it than Artificial Analysis or LMArena
- Buyers wanting an editorial take on "which model right now"
- Teams already using Vellum's evaluation product
- Quick cross-check against LMArena and AA
- Reading commentary on what the numbers mean
- You need broad model coverage including open-source
- You want hosting-provider price/speed comparisons
- You want vendor-neutral data without an editorial overlay
- You need real-time scores immediately after a model launch
Pricing
Full leaderboard access, methodology docs, and historical comparisons. No signup required.
What Vellum's leaderboard is
A short list of frontier models — typically GPT-4 / GPT-5 family, Claude Opus / Sonnet, Gemini Pro / Ultra, Llama 3.1 / Llama 4 — compared on a handful of curated benchmarks. The benchmark selection is opinionated: Vellum picks the evals they think predict real-world product utility (multi-turn instruction following, structured output reliability, function-calling accuracy, long-context comprehension) and skips the more academic ones.
Each model row carries a score per benchmark plus a short editorial line explaining how the model performed in practice. The methodology page is open: you can read which prompts were used, which scoring rubric was applied, and how often each test was repeated.
Where it wins
Editorial framing. Most benchmark sites give you numbers and leave the interpretation to you. Vellum tells you what their team thinks the numbers mean — which is useful for buyers who don't want to spend an afternoon learning what GPQA Diamond actually measures. The commentary is opinionated but transparent about the reasoning.
Curated model list. Walking onto LMArena and seeing 200+ models can be paralysing if you're trying to decide between, say, GPT-5.5 and Claude 4.7. Vellum starts from "you're probably picking between these eight" and goes from there. Less time, less optionality, more direct.
Methodology transparency. Most "we ran our own evals" comparison posts are one-shot blog posts. Vellum keeps a living methodology page, dates each run, and re-runs benchmarks when they update. That makes the leaderboard auditable in a way that "AI Engineer X tested these models" Twitter threads aren't.
Where it loses
Vendor curation. Vellum sells an evaluation product. The benchmarks they choose to highlight are the ones that overlap most with what their product helps customers measure. That's not dishonest — it's their stated methodology — but it does mean the model rankings are framed around Vellum's view of "what's important." A buyer with different priorities (cheap inference, on-prem deployment, fine-tuning ergonomics) will get less from Vellum than from Artificial Analysis.
Smaller catalog. Open-source models, smaller fine-tunes, and self-hosted variants generally don't appear. If your shortlist is "GPT-5.5 vs Claude Opus," the catalog is fine. If your shortlist is "best 8B model for $X budget," Vellum is the wrong reference.
Update cadence. Major launches (GPT-5.5 GA, Claude Opus 4.7) typically appear within a couple of weeks. Smaller releases or hosting-provider variations often don't appear at all. LMArena and Artificial Analysis are both faster.
How to use it
Treat Vellum as your second or third opinion, not your primary reference. The flow that works:
1. Form a hypothesis from LMArena ("users seem to prefer Claude over GPT on coding chat"). 2. Validate the deployment economics on Artificial Analysis ("Claude is 2× the price but actually faster — okay"). 3. Cross-check the editorial take on Vellum ("Vellum's commentary calls out instruction-following gaps in GPT-5.5 — relevant if your product depends on structured outputs").
When all three agree, the case is strong. When Vellum disagrees with the consensus, it's worth reading the commentary to understand why — sometimes Vellum's curated benchmark picked up a regression the broad sites missed; sometimes Vellum's methodology gave one model an edge it doesn't have in the wild.
Verdict
A useful supplement to LMArena and Artificial Analysis, not a replacement for either. The curated framing and editorial commentary are genuinely valuable for buyers who don't want to read three benchmark papers to make a model choice. The vendor-curation caveat is real — Vellum's methodology choices reflect Vellum's product priorities — but the methodology page is open enough that you can adjust for it.
Bookmark it, check it once a week, use it as a sanity check. Don't make it your only reference.

