Did Codex Reset
GitHub

Self-organizing team of three models averages 66.7% in Stanford and Together AI paper

DAIR.AI

DAIR.AI reports that a paper from Stanford and Together AI tested o3-mini, Claude Sonnet 4, and DeepSeek-V3 as a self-organizing team. Across five math and physics benchmarks, the team averaged 66.7%. The strongest member alone scored 48.8%, and a perfect router choosing among the members' independent answers scored 59.0%.

On AIME 2026, the team reached 71.2%, 13.4 points above that router. One member reviews the team's earlier exchanges and rewrites the teamwork strategy, covering roles, the order of discussion phases, who participates, and how partial answers are combined. Those strategies were learned from 15 AIME 2024 problems and then applied unchanged to held-out AIME 2024 problems and four new benchmarks.