Whisper (OpenAI)
OpenAI's own newer transcription models are now cheaper and more accurate than Whisper itself. It's not obsolete — free, open-source, and still the safe default — but no longer the only or best OpenAI option.
What is Whisper?
Whisper is OpenAI's automatic speech recognition system, originally released September 21, 2022, and trained on 680,000 hours of multilingual and multitask supervised audio scraped from the web. It's a Transformer-based model supporting multilingual transcription, language identification, timestamp prediction, and speech-to-English translation across 50-99+ languages depending on the source cited. Genuinely important and often overlooked: Whisper is open-source under the MIT license, meaning anyone can download it and self-host unlimited transcription for free, without ever touching OpenAI's hosted API — a real, distinctive advantage over most other speech-to-text tools in this category that require an ongoing subscription or per-minute fee no matter what.
The genuinely important, current thing worth understanding clearly in 2026: Whisper itself is now effectively a legacy model within OpenAI's own lineup. The hosted API (model ID whisper-1) still costs $0.006 per minute of audio, but OpenAI has since introduced newer, GPT-4o-based transcription models that are, in several real cases, both cheaper and more accurate — GPT-4o Mini Transcribe runs about $0.003/minute (half the cost of Whisper), GPT-4o Transcribe matches Whisper's $0.006/minute price with meaningfully better accuracy and built-in speaker diarization, and GPT-Realtime-Whisper handles live streaming transcription at around $0.017/minute. The honest, practical guidance from one detailed 2026 pricing analysis: "the mistake is assuming whisper-1 is the only OpenAI transcription choice" — for a new build, it's worth testing GPT-4o Mini Transcribe and GPT-4o Transcribe against your actual audio conditions (good and bad microphones, background noise, overlapping speakers, accents, specialist terminology) rather than assuming Whisper is still the default, though staying on whisper-1 remains sensible for existing integrations that already work well.
Whisper's origin, licensing, and technical specifications are drawn from an independent 2026 technical breakdown (Gate.AI) and OpenRouter's model documentation. The "legacy" positioning and newer OpenAI model comparison are drawn from a detailed independent pricing guide (DIY AI). The self-hosting break-even point is drawn from an independent cost analysis (BrassTranscripts). Competitive benchmarking against Groq, Google, and AssemblyAI is drawn from TokenMix.ai's April 2026 dataset.
Key features
Multilingual transcription
Supports 50-99+ languages depending on model variant and source.
Speech-to-English translation
Translates non-English audio directly into English text.
Fully open-source (MIT)
Self-host and run unlimited transcription without any API cost.
Timestamp prediction
Generates time-aligned transcripts suitable for captions and subtitles.
Simple REST API
Accessible via OpenAI's /v1/audio/transcriptions and /v1/audio/translations endpoints.
Direct video file support
Automatically extracts and transcribes audio tracks from video uploads.
Pricing
Self-hosted
- Free forever, but you manage your own infrastructure
- Real DevOps overhead starts around $276+/month
- Only cost-effective at roughly 2,400+ hours/month of audio
Hosted API (whisper-1)
- GPT-4o Mini Transcribe is cheaper at ~$0.003/minute
- GPT-4o Transcribe matches this price with better accuracy
- No volume discounts on standard OpenAI API pricing
GPT-Realtime-Whisper
- Purpose-built for real-time use cases Whisper wasn't designed for
- Higher per-minute cost reflects the streaming/live constraint
- Speaker diarization available on gpt-4o-transcribe-diarize
At high volume (10K hours/month), independent benchmarking shows Groq at roughly $400 (cheapest), Google around $2,800 (with volume discounts), OpenAI around $3,600 (no discount), and AssemblyAI around $7,500 (richest bundled features). Prices reflect OpenAI's published API pricing as of July 2026.
Available models
Integrations & platforms
Pros, cons & best for
Pros
- Genuinely free and open-source, unlike most category competitors
- Broad language support with a simple, well-documented API
- Remains a safe, reliable default for existing integrations
Cons
- No longer OpenAI's cheapest or most accurate transcription option
- Synchronous endpoint with a 25MB file limit needs careful workload design
- Neither cost nor accuracy leader against dedicated competitors like Groq or Google
Best for
- Existing integrations already built and working well on whisper-1
- Developers wanting genuinely free, self-hosted transcription infrastructure
- Not the pick for new builds — test GPT-4o Transcribe or Mini Transcribe first
Take a look inside
Alternatives
For meeting-specific transcription or human-reviewed accuracy instead:
Our verdict
Whisper's genuine, lasting contribution is real: a free, open-source speech recognition model that made high-quality transcription accessible to anyone willing to self-host it, and it remains a perfectly reliable, well-documented choice for existing integrations that already work. The honest thing worth internalizing clearly heading into the rest of 2026: it's no longer OpenAI's best or cheapest transcription option even within OpenAI's own lineup — GPT-4o Mini Transcribe is genuinely cheaper, and GPT-4o Transcribe is genuinely more accurate at the same price, and neither Whisper nor its newer siblings lead the broader market on cost (Groq) or accuracy (Google) compared to dedicated speech-to-text specialists. That's not a criticism so much as a reflection of a genuinely fast-moving category — for new builds, it's worth testing the newer OpenAI models or a specialist competitor against your specific audio conditions rather than defaulting to Whisper purely out of habit.
FAQ
Is Whisper still worth using in 2026?
For existing integrations, yes — it remains reliable and well-documented. For new builds, OpenAI's own newer GPT-4o-based models are often both cheaper and more accurate, so it's worth testing those first rather than defaulting to Whisper.
Is Whisper free to use?
Yes, if self-hosted — it's fully open-source under the MIT license with no per-minute fee at all. OpenAI's hosted API version costs $0.006 per minute of audio if you'd rather not manage your own infrastructure.
When does self-hosting Whisper actually save money over the API?
Roughly above 2,400-3,000 hours of audio per month, once real DevOps and infrastructure overhead (starting around $276+/month) is factored in — below that volume, the hosted API is generally cheaper.
What are OpenAI's newer transcription models compared to Whisper?
GPT-4o Mini Transcribe (~$0.003/minute, cheaper than Whisper), GPT-4o Transcribe (~$0.006/minute, same price with better accuracy and speaker diarization), and GPT-Realtime-Whisper (~$0.017/minute) for live streaming transcription.
How does Whisper compare to Groq, Google, and AssemblyAI?
Per independent 2026 benchmarking, Groq leads on cost and speed, Google leads on accuracy (especially on noisy or accented audio), and AssemblyAI offers the richest bundled feature set — Whisper remains a safe, simple default rather than the leader on any single dimension.
What file size and format limits does Whisper's API have?
A 25MB file-upload limit and a synchronous endpoint, meaning long or high-volume workloads need request timeouts, queuing, and retry logic designed around those constraints.