Arena.ai ranks GPT-6 Luna (Max) 24th in Code Arena WebDev
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
Arena.ai says Claude Opus 5.5 (Max) leads its Code Arena WebDev ranking at a blended $16 per million tokens, 60% below the price of the next-ranked model, GPT-6 Astra (Max). Scores are based on live use by Arena.ai’s global community.
Arena.ai says Claude Opus 5.5 (Max) delivers top performance at a blended $16 per million tokens, and describes that price-performance point as reshaping the Pareto frontier.
Arena.ai places Claude Opus 5.5 (Max) first in Code Arena WebDev at 1,818 points, 26 ahead of GPT-6 Astra (Max) and 126 above Opus 5 (Max). Claude says the model matches Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.
WebCraftBench has agents use a live app, then scores aesthetics, usability, and whether the original request was met. Tencent Hy reports an 85.3% match with human preference on 197 validated pairs.
Arena.ai places Qwen-Image-2.1 first among open-source models in both arenas. On Image Edit, it scores 1,367 points, ranks 16th overall, and trails GPT-Image-1.5-high-fidelity by three points.
OpenAI’s two models can be tested on Arena, where votes on real-world agentic tasks feed the leaderboard. Arena.ai says petergostev has also compared GPT-6 Sol with GPT-5.6 Sol using the same prompts at max reasoning.
OpenAI’s GPT-6 Sol and GPT-6 Luna can now be tested in Arena.ai’s Agent Arena, with scores still to come, and in Code Arena for WebDev, Text, Vision, Search, and Document. OpenAI says the models build on GPT-6 Astra and that their API prices are 50% lower than GPT-5.6 promotional pricing.
Arena.ai gives Qwen-Image-2.1 1,228 points and first place among open models, 17th overall, four points behind Gemini-3-pro-image-preview and 11 behind GPT-Image-1.5-high-fidelity.