Arena Leaderboard Reaches $100M Annualized Run-Rate, Eight Months After Commercial Launch
UC Berkeley's crowdsourced AI evaluation platform hits $100M ARR on consumption-based pricing, up from $30M in January 2026.
UC Berkeley's crowdsourced AI evaluation platform hits $100M ARR on consumption-based pricing, up from $30M in January 2026.
Researchers use a classic mystery game to test how well AI agents reason through multi-step deduction and suspect elimination.
Research shows that imperfect LLM-based evaluators can still meaningfully improve AI agent performance, challenging the assumption that evaluation noise is prohibitively harmful.