Arena.ai launches redesigned Leaderboard Overview
The new overview gathers live signals on newly added models, category performance, capability notes, and Arena news, and can show cross-arena scores for models such as GPT-6 Sol and Claude Opus 5.5.
The new overview gathers live signals on newly added models, category performance, capability notes, and Arena news, and can show cross-arena scores for models such as GPT-6 Sol and Claude Opus 5.5.
Tenure-track faculty at U.S. universities can submit Fall 2026 proposals on the scientific foundations of AI evaluation. Each project may receive up to $50,000, and the deadline is October 30, 2026.
Arena.ai says it has generated side-by-side outputs from Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol, and that scores for Claude Opus 5.5 are coming soon.
Arena.ai reports that OpenAI’s GPT-6 Sol (Max) scored 1,689 points and placed fourth in Code Arena: WebDev, at a blended price of $8 per million tokens. It is 72 points above GPT-5.6 Sol (xHigh), while Arena.ai says GPT-6 Luna’s score is still pending.
On Arena.ai, both models can be tried in Battle Mode and Agent Mode.
Arena.ai has added Claude Opus 5.5 to Agent Arena and to Battle Mode for WebDev, Text, Vision, and Document. Claude says it is the first Claude 5.5 model, performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.