Did Codex Reset
GitHub

Databricks tests find Opus 5.5 and GPT-6 Luna expand coding cost and quality

Patrick Wendell

Patrick Wendell says Databricks’ analysis of online workloads from 2,400 engineers, together with offline evaluations, finds that two of the three models released the prior week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna.

In that comparison, Opus 5.5 is the highest-quality mid-tier model. It beat all prior Opus models, GPT-6 Sol, and GPT-5.6 Sol, and reduced the cost of the same tasks by 20% versus Opus 4.8 in both offline and online analysis. Databricks is encouraging Opus 5.5 as an everyday default model for coding.

GPT-6 Luna was at least 20 times cheaper per task than Opus 5.5 in every offline benchmark Databricks tested and in observed online use. On one of its most difficult evaluation suites, Luna roughly matched Opus 4.6 while costing 99.3% less per task than Opus 4.6 did at that time, which Wendell describes as about a 100-fold cost reduction in roughly nine months. He says that Luna quality finding is preliminary and that Databricks is still evaluating it across a broader set of offline and online tests.

The production setup used Unity Gateway to route workloads across models and trace agentic interactions, with end-user harnesses including Omingent, Claude Code, Codex, and Cursor.