Arena, the crowdsourced AI model leaderboard born out of UC Berkeley in 2023, has hit $100 million in annualized run-rate revenue just eight months after launching its first paid product. The company's AI Evaluations service, which sells deep-dive model performance analytics to labs and enterprises, went live in September and has since become one of the fastest-growing revenue lines in the AI post-training market. Arena's annualized revenue stood at $30 million as recently as January, when it raised a $150 million Series A at a $1.7 billion post-money valuation.
The public-facing Arena product remains free: users type a prompt, two anonymized models answer, and the user picks the better response. Over 10 million such evaluations have fed the leaderboard, which has become the de facto scoreboard model labs check after every release. Arena now ranks models across text, coding, vision, image generation, and multi-step workflows through a newly introduced Agent Mode.
What changed in September was monetization. AI Evaluations packages the community's preference data into structured analytics that labs use to diagnose where their models lose to competitors and to guide post-training. CEO Anastasios Angelopoulos told TechCrunch that the speed of the revenue ramp has caught even Arena's own users off guard.
Key facts
- 01Arena reached $100M in annualized run-rate revenue eight months after launching its AI Evaluations service in September.
- 02Revenue tripled from $30M annualized in January, when Arena closed a $150M Series A at a $1.7B post-money valuation.
- 03The crowdsourced leaderboard is built on over 10M user evaluations comparing model outputs head-to-head.
- 04Arena has raised $250M total from Felicis, Andreessen Horowitz, Kleiner Perkins, Lightspeed, and UC Investments.
- 05Competitors for post-training dollars include Mercor at $1B annualized and Handshake's AI training arm at nearly $1B.
Angelopoulos clarified that the $100M figure is annualized run-rate, not strictly recurring — customers pay for consumption, so revenue can move with usage. Even with that caveat, the trajectory from $30M in January to $100M eight months after launching the product puts Arena in the same revenue tier as the human-labeling firms it now competes with for model-lab budgets.
Those competitors include Mercor, Surge, and Scale AI, the post-training data shops that have ridden the same wave. Angelopoulos said Arena is competing "for the same dollar" as those firms, even though the product shape is different — community preference data versus contracted human raters. Mercor's annualized revenue topped $1 billion this year, up from $500 million in September. Handshake's AI training arm grew gross annualized revenue from $550 million in January to nearly $1 billion by April, according to The Information.
The category is being pulled forward by the post-training arms race. As pretraining gains taper and labs lean harder on reinforcement learning from human and AI feedback, evaluation and preference data have moved from a research input to a core operating cost. A leaderboard sitting on 10 million human comparisons is, in that context, a structural asset rather than a side project.
Arena was co-founded by Angelopoulos, fellow UC Berkeley postdoctoral researcher Wei-Lin Chiang, who serves as CTO, and Ion Stoica, the Berkeley professor and Databricks co-founder who advised the project before it incorporated as a company in April 2025. The team has now raised $250 million total from Felicis, Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners, Laude Ventures, and UC Investments.
Arena's only direct crowdsourced competitor, Yupp, shut down in March, leaving the head-to-head model-comparison format effectively uncontested. That matters because Arena's moat is not the analytics product itself — labs could in theory build comparable internal tooling — but the volume and breadth of community evaluations feeding it. Recreating 10 million ranked comparisons across every frontier model is not something a new entrant can do quickly.
The risks are structural. Consumption pricing means revenue can dip as fast as it climbed if a major lab pulls spend or builds internal alternatives. Crowdsourced leaderboards have also faced scrutiny over whether community preferences correlate with real-world utility, and over gaming by labs that tune for the leaderboard rather than for users. Arena has not disclosed customer concentration, which in a market dominated by a handful of frontier labs is the number that matters most.
Arena's run from research project to $100M annualized in under a year is the cleanest evidence yet that the value in AI is migrating downstream from training compute to the data layer that shapes model behavior. The labs paying Arena are the same labs paying Mercor, Surge, and Scale — and the combined spend across that category is now a multi-billion-dollar line item that did not meaningfully exist three years ago. Whoever owns the preference data that decides which model wins owns a quiet but increasingly expensive bottleneck in the stack.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




