Qwen introduces Qwen-Audio-3.1 and cuts TTS, Realtime, and ASR prices
Qwen has introduced Qwen-Audio-3.1 as five models covering audio understanding, generation, interaction, and creation. ASR, TTS, and Realtime are upgraded, and the lineup adds ASR-Next and TTS-Next. Qwen says prices are lower across the lineup: about 70% off for TTS, about 85% off for Realtime, and up to 95% off for ASR.
Qwen says the upgraded ASR is stronger at multilingual and dialect recognition and includes polishing that removes fillers and repetitions for cleaner, more logical transcripts. ASR-Next supports multi-speaker transcription with speaker labels, timestamps, and aligned transcripts, and is presented as recognizing emotion, ambient sound, and machine sound for captioning, event localization, audio question answering, and reasoning. TTS covers multilingual and dialect synthesis, cross-lingual voice transfer, and instruction-based control of emotion, speed, and style. TTS-Next uses a combined language-model and diffusion setup to generate voice, sound effects, and background audio in one pass for audiobooks, podcasts, games, and ads. Realtime is described as speaking and listening at the same time, allowing interruption at any point, and slowing down with a more empathetic response when it detects a low mood.
Qwen-Audio-3.1-ASR and Qwen-Audio-3.1-Realtime are listed on qwencloud.com. Qwen says more APIs are coming soon.