Arena.ai ranks GPT-6 Luna (Max) 24th in Code Arena WebDev
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
Arena.ai says Claude Opus 5.5 (Max) leads its Code Arena WebDev ranking at a blended $16 per million tokens, 60% below the price of the next-ranked model, GPT-6 Astra (Max). Scores are based on live use by Arena.ai’s global community.
Arena.ai places Claude Opus 5.5 (Max) first in Code Arena WebDev at 1,818 points, 26 ahead of GPT-6 Astra (Max) and 126 above Opus 5 (Max). Claude says the model matches Claude Fable 5.1 on most tasks and costs 40% less to run than Opus 5.
WebCraftBench has agents use a live app, then scores aesthetics, usability, and whether the original request was met. Tencent Hy reports an 85.3% match with human preference on 197 validated pairs.
Arena.ai reports that OpenAI’s GPT-6 Sol (Max) scored 1,689 points and placed fourth in Code Arena: WebDev, at a blended price of $8 per million tokens. It is 72 points above GPT-5.6 Sol (xHigh), while Arena.ai says GPT-6 Luna’s score is still pending.
Arena.ai scores SpaceXAI’s Grok 4.7 (xHigh) at 1,632 points, 16 points and six places above Grok 4.6 (High). Simulations and Consumer Product also move up.