Arena.ai launches redesigned Leaderboard Overview
The new overview gathers live signals on newly added models, category performance, capability notes, and Arena news, and can show cross-arena scores for models such as GPT-6 Sol and Claude Opus 5.5.
The new overview gathers live signals on newly added models, category performance, capability notes, and Arena news, and can show cross-arena scores for models such as GPT-6 Sol and Claude Opus 5.5.
Arena.ai says it has generated side-by-side outputs from Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol, and that scores for Claude Opus 5.5 are coming soon.
With the same /design prompt, Command Code gave DeepSeek-V4.1-Flash 9/10 at $0.024, Mimo-V2.6-Flash 8/10 at $0.018, and Grok-4.7 3/10 at $0.35.
Artificial Analysis has extended the Controlled Voice Arena with separate text-to-speech leaderboards for Japanese, Mandarin Chinese, Hindi, Spanish, German, French, Portuguese, Vietnamese, and Arabic. Rankings use native-language prompts, standardized cloned voices, and preference votes from first-language speakers.
The Controlled Voice Arena now ranks text-to-speech models in Japanese, Mandarin Chinese, Hindi, Spanish, German, French, Portuguese, Vietnamese, and Arabic. Cartesia’s Sonic family leads eight of the nine boards, while Inworld AI’s Realtime TTS-2 ranks first in Mandarin.
OpenAI’s two models can be tested on Arena, where votes on real-world agentic tasks feed the leaderboard. Arena.ai says petergostev has also compared GPT-6 Sol with GPT-5.6 Sol using the same prompts at max reasoning.
OpenAI says it will support independent assessments with deep access across training, evaluation, and deployment, and is outlining four priority areas and principles for rigorous, secure, and independent work.