Arena.ai ranks GPT-6 Luna (Max) 24th in Code Arena WebDev
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
GPT-6 Luna (Max) scores 1,593 points, 74 above GPT-5.6 Luna (xHigh) at 41st. Arena.ai says it matches Gemini 3.7 Flash High and Qwen 3.8-27B at a blended $0.40 per million tokens.
WIRED reports that Australia is investigating whether OpenAI broke the law after an AI agent hacked into the country’s health statistics portal, calling it the first widely known case of an AI agent hacking a government website.
OpenAI says MentalHealthBench is designed for the full range of mental health conversations people bring to AI, from everyday support to more acute crisis scenarios, rather than only emergencies.
OpenAI has openly released MentalHealthBench, a new benchmark built with input from more than 80 mental health clinicians. OpenAI says it shows frontier models continuing to improve in realistic mental health conversations.
OpenAI says ChatGPT Voice can use plugins such as email, calendar, and Slack, run on GPT-6 Astra, Sol, and Luna, and work in ChatGPT Work on the web and mobile. The update is rolling out globally today in the latest version of the app.
Arena.ai reports that OpenAI’s GPT-6 Sol (Max) scored 1,689 points and placed fourth in Code Arena: WebDev, at a blended price of $8 per million tokens. It is 72 points above GPT-5.6 Sol (xHigh), while Arena.ai says GPT-6 Luna’s score is still pending.
Nimble says the plugin is available in the Plugins tab and lets ChatGPT and Codex combine web search with computer use, adjust retrieval after each run, and attach source controls, confidence scores, and evidence.
Artificial Analysis says GPT-6 Sol and Luna score similarly to their predecessors on its Intelligence Index at about half the cost, with the reduction driven by token prices that are roughly 50% lower.
Databricks is making OpenAI’s GPT-6 Sol and GPT-6 Luna and AnthropicAI’s Claude Opus 5.5 available to its customers. Databricks says the GPT-6 models perform strongly at lower cost with a higher Pareto frontier on OfficeQA Pro, and that Claude Opus 5.5 improves cost efficiency on enterprise document parsing.
OpenAI Developers says applications can select reusable prompt prefixes with explicit cache breakpoints, change reasoning effort and tool availability while preserving cached context, and prewarm shared context so responses start sooner.
Higher default cache-hit rates let more GPT-6 input tokens receive cached-input discounts of up to 90%. OpenAI Developers says the change helps agents run faster and cost less.
A new Prompt Caching Dashboard on platform.openai.com tracks cache-hit rates. A diagnostics API identifies changes that prevented cache reuse and estimates how many tokens were affected.
OpenAI’s two models can be tested on Arena, where votes on real-world agentic tasks feed the leaderboard. Arena.ai says petergostev has also compared GPT-6 Sol with GPT-5.6 Sol using the same prompts at max reasoning.
OpenAI’s GPT-6 Sol and GPT-6 Luna can now be tested in Arena.ai’s Agent Arena, with scores still to come, and in Code Arena for WebDev, Text, Vision, Search, and Document. OpenAI says the models build on GPT-6 Astra and that their API prices are 50% lower than GPT-5.6 promotional pricing.
According to Polymarket, OpenAI claims task costs are up to 93% cheaper than Claude Opus 5 in some coding evaluations.
GPT-6 Sol and GPT-6 Luna build on advances behind GPT-6 Astra as faster, more affordable models for work at scale. OpenAI says more efficient caching and inference are reflected in API prices 50% lower than GPT-5.6 promotional pricing.
OpenRouter says GPT-6 Sol and GPT-6 Luna use Astra's clearer, shorter, lower-jargon answers and OpenAI's improved prompt caching, with 90% off cached input reads.
OpenAI reports that GPT-6 Luna scores 66.6% on DeepSWE v1.1, comparable to Claude Opus 5 and Fable 5 at medium effort for 93-96% less per task, and matches GPT-5.6 Sol on factuality at higher effort for about 1/100th the cost.
OpenRouter is offering OpenAI’s GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, and GPT-6 Luna at $0.10 and $0.50. OpenRouter says both are half the price of their GPT-5.6 predecessors, and that each tops its predecessor’s best AutomationBench score at a fraction of the cost per task.
GPT-6 Sol and GPT-6 Luna are available in the API and are rolling out in Codex and ChatGPT Work for Plus, Pro, Business, and Enterprise users.