Arena.ai finds AI judges favor their own answers over human votes
Arena.ai compared 34,580 verdicts from 12 models with human votes across 1,460 battles on Arena. On average, a model picked its own answer 58% of the time; people picked that same answer 34% of the time. GPT-6 Astra picked itself 88% of the time.
The models also rarely called a draw. People marked a tie or “both bad” in 32% of battles, while GPT-5.6 Sol picked a winner 96% of the time. The AI judges agreed with other models 79% of the time and with people 57% of the time, and every judge sided with other AIs over people by 18 to 27 percentage points.
Full results are in DawidGalarowicz’s Arena article.