OpenAI Releases GPT Transcribe and GPT Live Transcribe
OpenAI has introduced two new speech recognition models: GPT Transcribe for batch audio processing (up to 34x faster than real-time) and GPT Live Transcribe for low-latency streaming transcription. Both are available via API, support multilingual input, and take text context and keywords into account. According to Artificial Analysis benchmarks, GPT Transcribe achieves a 3.3% word error rate (WER)—0.7 percentage points better than its predecessor GPT-4o Transcribe, though still behind the category leader Fun-Realtime-ASR-Preview from Alibaba Group at 1.7%. OpenAI has also cut pricing to $0.0045 per minute for GPT Transcribe and $0.017 per minute for the live version.
Related: OpenAI Releases GPT-Realtime 2 Voice Models, OpenAI Develops Bidirectional Audio Model
| GPT Transcribe API docs | GPT Live Transcribe API docs | AA-WER benchmark |