# Hassan releases GeoGuess Bench to evaluate AI models on GeoGuessr

GeoGuess Bench evaluates AI models on 210 worldwide geolocation photo tasks. Claude Opus 5.5 ranked first overall, while open models delivered competitive accuracy at lower inference costs.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/66e118cb2507476e89779c33.html

Content language: en
Localization state: sameLanguage

Source: Hassan · Published 10/6/2026, 09:27:53

Hassan has released GeoGuess Bench, a benchmark designed to evaluate how well AI models play the geolocation game GeoGuessr. The benchmark provides each model with 210 photos from across the globe, asks it to predict where each photo was taken, and scores performance based on distance from the true location.

In the benchmark results, Claude Opus 5.5 took the top spot, scoring above Fable 5.1. Open models also demonstrated strong cost efficiency: Muse Glimmer 30B outperformed GPT 6 Astra at roughly 45 times lower cost, while GLM 5.3 Flash matched GPT 6.1 Sol at approximately one-fifth the cost. The full leaderboard and guesses are available on [GeoGuess Bench](http://geoguessbench.com), with code hosted on [GitHub](http://github.com/Nutlope/geoguessbench).

Tags: Benchmark, GeoGuessr, Multimodal AI, Open Source

[View original post](https://x.com/nutlope/status/2107502264101331150)
