Sonnet 5.5 matches Opus 5.5 High on Ibragim’s my-mini-bench at lower time and cost
Ibragim compared Claude Sonnet 5.5 with Opus 5.5 High on my-mini-bench, using tasks from his own work: data conversion, frontend, refactoring, and running task evaluations. The setup was 10 tasks run three times each, with High reasoning, in pi-agent and Claude Code through harbor.
In pi, both Sonnet 5.5 and Opus 5.5 High passed 30 of 30 runs. Sonnet 5.5 averaged 46 seconds and $0.10 per task, compared with 103 seconds and $0.30 for Opus 5.5 High—about 2.2 times faster and about 3 times cheaper.
Against Sonnet 5 in pi, Sonnet 5.5 moved from 25 of 30 passes to 30 of 30, turns from 12.7 to 4.3, and output tokens from 17.7k to 6.5k. Ibragim says the trajectories show a different approach to reading files, writing code, and testing. He says the benchmark is already saturated and uses it to find a cheaper, faster setup that still completes his tasks.