Did Codex Reset
GitHub

BRIDGE ASR 2.0 tests 23 speech recognition models on long multilingual conversations

elvis

The benchmark tests 23 models. Its code-switch F1 score checks whether English words mixed into an Indic sentence stay in English; a transcription that writes "data backup" in Devanagari script scores zero.

The leaderboard can be filtered by language and by metric. Methodology and evaluation data are public on the BRIDGE ASR 2.0 page, and elvis says the team wants researchers to test the benchmark and find where it breaks.