Gemini 3.8 Flash TTS ranks first on the Artificial Analysis Pronunciation Robustness Benchmark
Artificial Analysis says Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Both are Google DeepMind text-to-speech models that support preset voices and voice replication. Gemini 3.1 Flash TTS led the Pronunciation Robustness Benchmark at launch; Artificial Analysis now places Gemini 3.8 Flash TTS first and 64 Elo points above Gemini 3.1 Flash TTS on the Provider Voice Arena.
On the Provider Voice Arena, Gemini 3.8 Flash TTS debuts at No. 2 with an Elo of 1,263, behind Sonic 3.6 at 1,272 and ahead of Qwen-Audio-3.0-TTS-Plus at 1,260. Gemini 3.8 Flash-Lite TTS debuts at No. 6 with an Elo of 1,236, one point behind Simba 3.2. Both rank above Gemini 3.1 Flash TTS at 1,199. On pronunciation robustness, Gemini 3.8 Flash TTS scores 89.5%, ahead of Gemini 3.1 Flash TTS at 88.2%, SpaceXAI TTS at 87.6%, and Gemini 3.8 Flash-Lite TTS at 87.4%. Artificial Analysis says that places three Google models in the top four.
Across its nine multilingual Controlled Voice Arenas, Gemini 3.8 Flash-Lite TTS debuts at No. 1 in Japanese, No. 2 in Arabic, and No. 3 in German, while Gemini 3.8 Flash TTS ranks No. 2 in Japanese and Portuguese. Artificial Analysis measures Gemini 3.8 Flash TTS at 44.1 characters per second, about 2.7 times realtime, and Gemini 3.8 Flash-Lite TTS at 40.2 characters per second, about 2.4 times realtime. It lists prices of $32.98 per 1 million characters for Gemini 3.8 Flash TTS and $22.07 per 1 million characters for Gemini 3.8 Flash-Lite TTS, compared with $18.31 for Gemini 3.1 Flash TTS and $100 for Eleven v3.