Skip to main content
OpenAI

o3

o3 is a reasoning-focused chat model from OpenAI in the company's o-series, designed to produce internal reasoning before answering.

Input from
$1.00 / 1M tokens
across 4 providers

API Pricing

Cheapest on OpenAI — 40% below avg
ProviderInput / 1MOutput / 1MCached / 1M
OpenAI logo
OpenAIBatch
$1.00$4.00-
$1.00$4.00$0.250
$2.00$8.00$0.500
$2.00$8.00$0.500
$2.00$8.00$0.500
$2.00$8.00$0.500

Prices updated daily. Last check: Oct 10, 2026

Compare API pricing for every OpenAI model →

o3 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
20.2 / 100
Math
88.3 / 100
Output Speed
109 t/s
Latency (TTFT)
5.7s

Reasoning & Knowledge

  • MMLU-Pro85.3%
  • GPQA Diamond82.7%
  • Humanity's Last Exam20.1%

Coding

  • LiveCodeBench80.8%

Math

  • AIME 202588.3%
  • AIME90.3%
  • MATH-50099.2%

Agentic & Tool Use

  • Terminal-Bench Hard37.1%
  • τ²-bench80.7%

Instruction & Long Context

  • IFBench71.4%
  • Long-Context Reasoning74.7%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
OpenAI
Modalities
Text

Capabilities

Open Source
No

Strengths & Limitations

Strengths

  • Reasoning-oriented design: the model works through problems internally before producing a final answer, which suits multi-step tasks
  • Measured output speed of approximately 138.8 tokens per second once generation begins (Artificial Analysis)
  • Part of OpenAI's o-series, so it shares the reasoning-model request pattern and tooling used across that line
  • Streaming generation pace is comparable to standard chat models despite the reasoning step
  • Available through multiple providers, allowing price and latency comparison in the table on this page
  • Well suited to workloads where a single correct answer is worth more than a fast one, such as math, analysis, and debugging

Limitations

  • Time to first token measured at roughly 4,788 ms, which is slow for interactive chat UIs that expect immediate output
  • Reasoning tokens consumed during deliberation are typically billed as output, making per-request cost less predictable than non-reasoning models
  • Context window size is not tracked in our database — verify against OpenAI or provider documentation before designing long-context workloads
  • Model weights are not distributed by OpenAI for self-hosting
  • Higher deliberation effort is wasted on simple, high-volume tasks that a lighter model would handle at lower latency

Key Features

•Internal reasoning pass before final answer generation (o-series behavior)
•Measured throughput of ~138.8 output tokens/second
•Measured time to first token of ~4,788 ms
•Chat completion interface compatible with OpenAI-style API clients
•Streaming output support once the reasoning phase completes
•Available through more than one inference provider, listed in the pricing table above

About o3

o3 is a chat model developed by OpenAI as part of its o-series line of reasoning models. Unlike conventional chat completion models that respond immediately, o-series models allocate additional computation to an internal reasoning process before emitting a final answer, which shifts where latency appears in a request and changes how the model is billed and benchmarked. Independent measurements collected by Artificial Analysis put o3's output speed at roughly 138.8 tokens per second, with a time to first token of about 4,788 ms. That latency profile is characteristic of reasoning models: the initial wait covers deliberation before any visible output, after which generation proceeds at a normal streaming pace. Applications that display partial output to users should account for the multi-second gap before the first token arrives. Other specifications for o3 — including context window size and supported input modalities — are not tracked in our database, so consult OpenAI's model documentation or the individual provider's API reference for those details. In practice, o3 is used where accuracy on multi-step problems matters more than response latency: mathematical work, scientific analysis, debugging and code review, and planning steps inside agentic pipelines. Compared with OpenAI's non-reasoning chat models, it trades faster first-token response for additional deliberation, and compared with smaller o-series variants it sits at the higher-effort end of that trade-off. Providers and pricing for o3 vary; see the pricing table on this page for current options.

Common Use Cases

o3 fits workloads where correctness on multi-step problems justifies a slower first response: mathematical derivations, scientific and quantitative analysis, algorithm design, code review and bug diagnosis, and the planning or verification stages of agentic systems where a single reasoning call decides subsequent actions. Its roughly 4.8-second time to first token makes it a poor match for latency-sensitive conversational front ends, autocomplete, or high-volume classification and extraction jobs — those are better served by a lighter, non-reasoning model, with o3 reserved for the harder subset of requests. A common pattern is routing: handle the bulk of traffic with a fast model and escalate only difficult cases to o3.

Frequently Asked Questions

How much does o3 cost to use?

Pricing for o3 varies by provider and by pricing type (input tokens, output tokens, cached input, and in some cases batch rates). Because reasoning models bill their internal deliberation as output tokens, effective cost per request can be higher than the headline output rate implies. Check the pricing table on this page for current per-provider rates.

What is o3 best used for?

Tasks that reward deliberation: multi-step math, scientific and quantitative reasoning, code debugging and review, and planning steps inside agentic workflows. It is less appropriate for latency-sensitive chat interfaces or high-volume, low-complexity jobs.

Why does o3 take several seconds before responding?

o3 is an OpenAI o-series reasoning model, so it performs an internal reasoning pass before producing visible output. Artificial Analysis measured its time to first token at about 4,788 ms. Once generation starts, output streams at roughly 138.8 tokens per second.

How does o3 compare to OpenAI's non-reasoning chat models?

The main difference is where time and tokens are spent. Non-reasoning chat models begin emitting output almost immediately; o3 deliberates first, which raises time to first token and adds reasoning tokens to the bill, in exchange for stronger handling of problems that require several logical steps.

What context window does o3 support?

Our database does not track a confirmed context window for o3. Consult OpenAI's model documentation or the specific provider's API reference before building applications that depend on long inputs.

Can I self-host o3?

OpenAI does not distribute o3 weights for self-hosting. Access is through API providers, which are listed with their current rates in the pricing table above.