Skip to main content
OpenAI

Whisper V3 Turbo

Whisper V3 Turbo is OpenAI's speech-to-text model, a decoder-reduced variant of Whisper large-v3 built for faster multilingual transcription and English translation.

From
$0.0034 / minute
across 4 providers

API Pricing

Cheapest on Scaleway — 84% below avg
ProviderPrice /minAlt /min
$0.0034/min-
$0.020/min-
$0.040/min-
$0.270$0.850

Prices updated daily. Last check: Oct 3, 2026

Whisper V3 Turbo pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Model Details

General

Creator
OpenAI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No
Aliases
Whisper Large V3 Turbo, whisper-large-v3-turbo, openai/whisper-large-v3-turbo

Strengths & Limitations

Strengths

  • Decoder-reduced architecture decodes faster than Whisper large-v3 while reusing the same audio encoder
  • Weights are publicly released, enabling self-hosting and on-device deployment alongside hosted API use
  • Multilingual transcription plus speech translation into English from the same checkpoint
  • Automatic language detection, or explicit language pinning via a decoding parameter
  • Segment-level timestamps suitable for subtitle and caption generation
  • Broad ecosystem support: Transformers, faster-whisper/CTranslate2, whisper.cpp, and multiple inference APIs
  • Initial-prompt conditioning can bias domain vocabulary, names, and punctuation style

Limitations

  • Accuracy on difficult audio and lower-resource languages is generally below the full Whisper large-v3 checkpoint
  • Speech translation quality is weaker than transcription; large-v3 is the safer choice for X→English translation
  • No built-in speaker diarization — separate tooling is required to label who spoke when
  • 30-second processing windows mean long-form audio depends on the chunking or sequential strategy of the serving library, which can affect boundary accuracy
  • Like other Whisper models, it can hallucinate text during silence, music, or heavy background noise

Key Features

•Automatic speech recognition with multilingual support
•Speech-to-English translation mode
•Reduced decoder-layer architecture derived from Whisper large-v3
•Automatic language identification with manual override
•Segment-level timestamp output
•Initial prompt conditioning for vocabulary and formatting bias
•Openly released weights for self-hosted and local inference
•Support across Transformers, faster-whisper/CTranslate2, and whisper.cpp runtimes

About Whisper V3 Turbo

Whisper V3 Turbo (also distributed as whisper-large-v3-turbo) is an automatic speech recognition model from OpenAI and part of the Whisper family of encoder-decoder transcription models. It is derived from Whisper large-v3 by substantially reducing the number of decoder layers while keeping the same audio encoder, which lowers decoding time per audio segment. OpenAI released the weights publicly under a permissive license, so the model is served both through hosted inference APIs and run locally via frameworks such as Hugging Face Transformers, faster-whisper/CTranslate2, and whisper.cpp. The model takes audio input and produces text output. Like other Whisper checkpoints, it processes audio in 30-second windows with log-Mel spectrogram features, and longer recordings are handled by chunking or sequential long-form decoding in the serving library. It supports multilingual transcription across the same broad language set as large-v3, plus speech translation into English, and can emit segment-level timestamps. Language can be auto-detected or pinned with a decoding parameter, and an initial prompt can be supplied to bias vocabulary and formatting. In practice Turbo is chosen when throughput and latency matter more than squeezing out the last increment of accuracy: batch transcription of media archives, captioning pipelines, voice-note ingestion, and near-real-time streaming front-ends. Compared with the full large-v3 checkpoint it trades some accuracy — most noticeably on harder or lower-resource languages and on speech translation — for faster decoding, while remaining considerably more capable than the small and base Whisper checkpoints. Hosted providers typically bill audio transcription per minute or per hour of audio rather than per token, so the cost model differs from chat LLMs; see the pricing table on this page for current provider rates.

Common Use Cases

Whisper V3 Turbo suits high-volume and latency-sensitive transcription work: bulk processing of podcast, lecture, and video archives; subtitle and caption generation using its segment timestamps; voice-memo and meeting-note ingestion in consumer apps; and the transcription stage of voice agent pipelines where an LLM consumes the text downstream. Its multilingual coverage makes it usable for mixed-language media libraries and for producing English translations of foreign-language audio, though teams whose priority is maximum accuracy on noisy audio, accented speech, low-resource languages, or translation should benchmark it against the full Whisper large-v3 checkpoint first. Because the weights are openly available, it is also a common choice for on-premise or edge deployments where audio cannot leave a controlled environment.

Frequently Asked Questions

How much does Whisper V3 Turbo cost to use?

Pricing depends on the provider and the billing model — speech-to-text is usually charged per minute or per hour of audio processed rather than per token, and self-hosting shifts the cost to GPU time instead. Rates vary widely between hosted APIs. Check the pricing table on this page for current provider pricing.

What is Whisper V3 Turbo best used for?

It is best suited to fast, high-volume transcription: media archives, captioning and subtitling, voice notes, and the ASR stage of voice assistant pipelines. It handles multilingual audio and can translate speech into English, making it a general-purpose transcription workhorse where decoding speed matters.

How does it differ from Whisper large-v3?

Turbo keeps the same audio encoder but uses far fewer decoder layers, which reduces decoding time per segment. The trade-off is somewhat lower accuracy on harder audio and lower-resource languages, and noticeably weaker speech translation. Use large-v3 when transcription accuracy or translation quality is the priority, and Turbo when throughput and latency dominate.

Can I run Whisper V3 Turbo myself instead of using an API?

Yes. OpenAI released the weights publicly, and the model is supported by Hugging Face Transformers, faster-whisper/CTranslate2, whisper.cpp, and similar runtimes, including quantized builds for consumer hardware. Self-hosting is common for privacy-sensitive audio or for very large batch workloads.

Does it identify different speakers in a recording?

No — Whisper models transcribe speech but do not perform speaker diarization on their own. Producing speaker-labelled transcripts requires pairing the model with a separate diarization system such as a pyannote-based pipeline.