Did Codex Reset
GitHub

Arena.ai finds AI judges favor their own answers over human votes

Arena.ai

Arena.ai compared 34,580 verdicts from 12 models with human votes across 1,460 battles on Arena. On average, a model picked its own answer 58% of the time; people picked that same answer 34% of the time. GPT-6 Astra picked itself 88% of the time.

The models also rarely called a draw. People marked a tie or “both bad” in 32% of battles, while GPT-5.6 Sol picked a winner 96% of the time. The AI judges agreed with other models 79% of the time and with people 57% of the time, and every judge sided with other AIs over people by 18 to 27 percentage points.

Full results are in DawidGalarowicz’s Arena article.