StepFun releases StepAudio 3 ASR at 1.7% WER on the AA-WER Index
Artificial Analysis says StepFun has released StepAudio 3 ASR, a speech-to-text model available through the StepFun API for non-streaming transcription. It is the first StepFun model to reach the top of Artificial Analysis’s non-streaming speech-to-text leaderboard, with a 1.7% word error rate on the AA-WER Index, effectively tied with Alibaba’s Fun-Realtime-ASR-preview at 1.7%. Artificial Analysis recorded 4.7% WER for StepAudio 2.5 ASR.
On AA-AgentTalk, Artificial Analysis’s held-out voice-agent dataset, StepAudio 3 ASR scores 1.4% WER. On long-form Earnings22 calls it scores 2.8%, behind Fun-Realtime-ASR-preview at 1.8%.
Artificial Analysis measures its transcription speed at 88x real time, behind MAI-Transcribe-2 at 374x, Smallest AI Pulse Pro at 285x and Grok Voice Transcribe 2.0 at 154x. That is roughly level with ElevenLabs Scribe v2 at 84x and Gemini 3.5 Transcribe at 91x, and ahead of Fun-Realtime-ASR-preview at 20x. StepAudio 3 ASR costs $0.40 per hour, or $6.67 per 1,000 minutes. Artificial Analysis calls it the most expensive of the five most accurate models; MAI-Transcribe-2 and Grok Voice Transcribe 2.0 cost $1.67 per 1,000 minutes, and ElevenLabs Scribe v2 costs $3.67 per 1,000 minutes.