Self-organizing team of three models averages 66.7% in Stanford and Together AI paper
DAIR.AI reports that a Stanford and Together AI paper found o3-mini, Claude Sonnet 4, and DeepSeek-V3 averaged 66.7% across five math and physics benchmarks as a self-organizing team, above 48.8% for the strongest member alone and 59.0% for a perfect router.