o3
o3 is a reasoning-focused chat model from OpenAI in the company's o-series, designed to produce internal reasoning before answering.
API Pricing
Cheapest on OpenAI — 40% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.00 | $4.00 | - | |
| $1.00 | $4.00 | $0.250 | |
| $2.00 | $8.00 | $0.500 | |
| $2.00 | $8.00 | $0.500 | |
| $2.00 | $8.00 | $0.500 | |
| $2.00 | $8.00 | $0.500 |
Prices updated daily. Last check: Oct 10, 2026
Compare API pricing for every OpenAI model →o3 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro85.3%
- GPQA Diamond82.7%
- Humanity's Last Exam20.1%
Coding
- LiveCodeBench80.8%
Math
- AIME 202588.3%
- AIME90.3%
- MATH-50099.2%
Agentic & Tool Use
- Terminal-Bench Hard37.1%
- τ²-bench80.7%
Instruction & Long Context
- IFBench71.4%
- Long-Context Reasoning74.7%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- OpenAI
- Modalities
- Text
Capabilities
- Open Source
- No
Strengths & Limitations
Strengths
- Reasoning-oriented design: the model works through problems internally before producing a final answer, which suits multi-step tasks
- Measured output speed of approximately 138.8 tokens per second once generation begins (Artificial Analysis)
- Part of OpenAI's o-series, so it shares the reasoning-model request pattern and tooling used across that line
- Streaming generation pace is comparable to standard chat models despite the reasoning step
- Available through multiple providers, allowing price and latency comparison in the table on this page
- Well suited to workloads where a single correct answer is worth more than a fast one, such as math, analysis, and debugging
Limitations
- Time to first token measured at roughly 4,788 ms, which is slow for interactive chat UIs that expect immediate output
- Reasoning tokens consumed during deliberation are typically billed as output, making per-request cost less predictable than non-reasoning models
- Context window size is not tracked in our database — verify against OpenAI or provider documentation before designing long-context workloads
- Model weights are not distributed by OpenAI for self-hosting
- Higher deliberation effort is wasted on simple, high-volume tasks that a lighter model would handle at lower latency
Key Features
About o3
Common Use Cases
o3 fits workloads where correctness on multi-step problems justifies a slower first response: mathematical derivations, scientific and quantitative analysis, algorithm design, code review and bug diagnosis, and the planning or verification stages of agentic systems where a single reasoning call decides subsequent actions. Its roughly 4.8-second time to first token makes it a poor match for latency-sensitive conversational front ends, autocomplete, or high-volume classification and extraction jobs — those are better served by a lighter, non-reasoning model, with o3 reserved for the harder subset of requests. A common pattern is routing: handle the bulk of traffic with a fast model and escalate only difficult cases to o3.
Frequently Asked Questions
How much does o3 cost to use?
Pricing for o3 varies by provider and by pricing type (input tokens, output tokens, cached input, and in some cases batch rates). Because reasoning models bill their internal deliberation as output tokens, effective cost per request can be higher than the headline output rate implies. Check the pricing table on this page for current per-provider rates.
What is o3 best used for?
Tasks that reward deliberation: multi-step math, scientific and quantitative reasoning, code debugging and review, and planning steps inside agentic workflows. It is less appropriate for latency-sensitive chat interfaces or high-volume, low-complexity jobs.
Why does o3 take several seconds before responding?
o3 is an OpenAI o-series reasoning model, so it performs an internal reasoning pass before producing visible output. Artificial Analysis measured its time to first token at about 4,788 ms. Once generation starts, output streams at roughly 138.8 tokens per second.
How does o3 compare to OpenAI's non-reasoning chat models?
The main difference is where time and tokens are spent. Non-reasoning chat models begin emitting output almost immediately; o3 deliberates first, which raises time to first token and adds reasoning tokens to the bill, in exchange for stronger handling of problems that require several logical steps.
What context window does o3 support?
Our database does not track a confirmed context window for o3. Consult OpenAI's model documentation or the specific provider's API reference before building applications that depend on long inputs.
Can I self-host o3?
OpenAI does not distribute o3 weights for self-hosting. Access is through API providers, which are listed with their current rates in the pricing table above.