Why GPT Transcribe still trails rivals on error rates

OpenAI has introduced GPT Transcribe and GPT Live Transcribe through its API, with one model focused on pre-recorded audio and the other on real-time streaming. GPT Transcribe improves over GPT-4o Transcribe and costs less, but AA-WER benchmark results still place ElevenLabs, Google, and Mistral ahead on error rates.

WTF Index NEUTRAL
◄ Terminator 0 Idiocracy 1 ►

This is a routine transcription model release and benchmark comparison with only a mild dependency/automation angle.

Why GPT Transcribe still trails rivals on error rates

OpenAI’s newest speech recognition release shows clear progress, but the competitive picture is still mixed. GPT Transcribe improves on GPT-4o Transcribe, lowers the cost per audio minute, and adds practical context options, yet it does not lead the AA-WER ranking reported by Artificial Analysis.

What OpenAI released

OpenAI has released two speech recognition models through its API: GPT Transcribe and GPT Live Transcribe. The distinction between them is straightforward. GPT Transcribe is designed for pre-recorded audio files, while GPT Live Transcribe is built for real-time streaming with low latency.

That split matters because transcription workloads are not all the same. Some users need to process audio after a recording is complete. Others need words to appear while speech is still happening, where latency becomes central to the experience.

For pre-recorded audio, GPT Transcribe processes files about 34 times faster than real time. That makes the model relevant for workloads where speed at scale is important, such as turning stored audio into text quickly through an API pipeline.

Accuracy improved, but not enough to lead

According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe reaches a word error rate of 3.31 percent. That is a 0.7 percentage point improvement over GPT-4o Transcribe, its year-old predecessor.

The improvement gives OpenAI a stronger transcription model than before. Still, the benchmark ranking shows that several competitors remain ahead on error rates.

  • ElevenLabs Scribe v2 leads with a 2.3 percent error rate.
  • Google's Gemini 3 Pro follows at 2.9 percent.
  • Mistral's Voxtral Small is listed at 3 percent.
  • GPT Transcribe is reported at 3.31 percent.

For buyers comparing speech recognition models, the difference is not just whether a model works. The benchmark numbers frame the tradeoff between better accuracy, cost, speed, and the specific needs of a product or workflow.

Pricing is part of the story

OpenAI also reduced pricing with the new model. GPT Transcribe costs $0.0045 per minute of audio, which the source describes as a 25 percent drop.

That price cut improves the case for GPT Transcribe, especially for API users who process large volumes of audio. But the competitive pressure is visible here too. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute.

This creates a practical comparison for developers and companies. GPT Transcribe brings faster-than-real-time processing and a lower OpenAI price than before, but rivals still have stronger reported error rates in the AA-WER ranking, and Mistral has a lower starting price for Voxtral Transcribe V2.

Context and language support

Both GPT Transcribe and GPT Live Transcribe accept text as transcription context, keywords, and multiple input languages. Those options can be important because speech recognition is often used in settings where names, terminology, or domain-specific language matter.

Text context and keywords give users a way to guide the transcription model without changing the audio itself. The source does not provide examples of how much those inputs affect results, so the clearest takeaway is that OpenAI is exposing these controls as part of the API offering.

The models also sit alongside OpenAI’s recently announced Realtime model generation. That broader release includes the real-time transcription model GPT-Realtime-Whisper, placing GPT Live Transcribe within a larger set of real-time audio capabilities.

What this means for speech recognition choices

The release makes OpenAI more competitive in speech-to-text, but not dominant on the reported benchmark. GPT Transcribe improves accuracy compared with GPT-4o Transcribe, lowers price, and handles pre-recorded audio quickly. GPT Live Transcribe addresses real-time streaming, where low latency is the defining requirement.

At the same time, the AA-WER ranking gives buyers a reason to compare carefully. ElevenLabs Scribe v2, Google's Gemini 3 Pro, and Mistral's Voxtral Small all post lower error rates than GPT Transcribe in the figures reported by Artificial Analysis.

The result is a market where OpenAI’s new transcription models are stronger than their predecessor, but the best choice still depends on the job. For some use cases, API access, context support, speed, and integration may matter most. For others, the lowest reported word error rate or the lowest per-minute price may carry more weight.