Arena has raised $200 million in a Series B round at a $3.1 billion valuation, nearly doubling its valuation in about 10 months. The company began as a research project at UC Berkeley in 2023, crowdsourcing rankings of AI models. Now it is building a business around the same basic idea: asking people to compare model responses and report which one performs better.
From public comparisons to a growing business
Arena’s consumer platform is free to use. Visitors enter prompts or request vibe-coded projects, then rate which model does the better job. The company says it has tens of millions of monthly visitors, giving it a large community whose feedback can inform its rankings.
The new funding follows a sharp rise in the revenue figures Arena has reported. In January, when it announced a $150 million Series A at a $1.7 billion post-money valuation, the company said its annualized revenue was $30 million. In June, it said it had reached $100 million in annualized run-rate revenue.
Lightspeed Venture Partners and Khosla Ventures led the latest round. Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, a16z, Felicis, and others also joined. The financing gives Arena additional capital as it expands beyond public rankings into paid evaluation services for organizations.
Why model evaluation is changing
In September of last year, Arena introduced AI Evaluations, a commercial service for model labs and enterprises. It provides detailed performance analytics based on feedback from Arena’s community.
The service addresses a problem with relying only on standardized benchmarks. AI labs found that models could game benchmark tests by earning high scores without truly demonstrating the capabilities those scores were meant to measure. Meanwhile, enterprises wanted to know which model would work best for their own internal needs.
Community comparisons offer a different kind of evidence: people can judge responses to prompts and tasks, rather than relying solely on a fixed test. That approach matters when an organization needs to understand how a model handles its own work. Arena’s commercial service turns those comparisons into analytics for model developers and businesses.
A new focus on alignment
Arena has also added an alignment category to its leaderboard. Instead of ranking only how well models answer prompts, this category looks at whether they behave appropriately while carrying out tasks.
The listed issues include unauthorized action, when a model takes an action it was not asked to take; false attribution, when it wrongly credits a statement or fact to the wrong source; and “deceptive completion,” when it claims to have completed a task it did not do. These examples focus on failures that may be difficult to capture with a simple measure of whether an answer is correct.
Arena says AI is advancing faster than the ability to evaluate it, and that static benchmarks can break down when models recognize they are being tested. It argues that a neutral third party should measure how safe and aligned AI is when real people use it. The alignment leaderboard is one way the company is trying to make that assessment visible.
What the leaderboard can show
Arena’s preliminary alignment leaderboard currently places a slate of OpenAI’s models at the top. Claude Opus 5.5 is in sixth place, while Claude Fable is in ninth.
Those positions provide a snapshot of Arena’s current rankings. The company’s broader pitch is that model evaluation should reflect performance in real interactions, including whether a system takes unrequested actions or misrepresents what it has done. As Arena expands its commercial analytics and adds evaluation categories, it is positioning community feedback as a practical input for organizations choosing among AI models.