GPT-5.3 Codex
GPT-5.3 Codex is a coding-oriented chat model from OpenAI in the GPT-5.x Codex line, measured at roughly 121 output tokens per second in third-party throughput testing.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.75 | $14.00 | $0.180 | |
| $1.75 | $14.00 | - | |
| $1.75 | $14.00 | $0.175 | |
| $1.75 | $14.00 | $0.175 |
Prices updated daily. Last check: Oct 10, 2026
Compare API pricing for every OpenAI model →GPT-5.3 Codex pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond91.5%
- Humanity's Last Exam42.5%
Agentic & Tool Use
- Terminal-Bench Hard53.0%
- τ²-bench86.0%
Instruction & Long Context
- IFBench75.4%
- Long-Context Reasoning83.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- OpenAI
- Modalities
- Text
Capabilities
- Open Source
- No
Strengths & Limitations
Strengths
- Positioned in OpenAI's Codex line, which is oriented toward software engineering and code-editing workloads rather than general chat
- Sustained output throughput measured at about 121 tokens per second by Artificial Analysis, which is competitive for a deliberation-heavy model
- Served through a standard chat interface, so integration mirrors other OpenAI chat models
- Long pre-response deliberation time suggests the model is tuned for multi-step problem solving rather than single-shot autocompletion
- Suited to agentic coding harnesses where a task is dispatched and the result collected asynchronously
- Part of an actively iterated family, so prompts and tooling built for earlier Codex-line models generally transfer
Limitations
- Time to first token measured at roughly 37.6 seconds, which is unsuitable for inline editor autocomplete or real-time conversational use
- Context window is not among the specifications we have confirmed, so long-repository workflows should be validated before committing
- Model weights are not distributed for self-hosting by OpenAI's usual practice, so deployment depends on hosted APIs
- We do not track confirmed coding benchmark scores for this model, making direct capability comparisons with peers difficult from our data alone
- Coding-tuned models can underperform general-purpose siblings on non-technical writing and open-domain chat
Key Features
About GPT-5.3 Codex
Common Use Cases
GPT-5.3 Codex fits software engineering workloads where correctness matters more than response latency: implementing a feature across several files, diagnosing a failing test suite, refactoring a module, reviewing a diff, or acting as the reasoning engine inside an autonomous coding agent that runs tasks in the background. The measured ~37.6 second time to first token makes it a poor match for inline IDE autocomplete or live pair-programming chat, where developers expect sub-second responses; those workloads are better served by lower-latency models, with Codex-line models reserved for the harder tasks queued behind them. Because it generates at roughly 121 tokens per second once it starts, it can produce substantial patches and explanations in a reasonable wall-clock window after the initial wait. Teams should verify context window limits against their repository size before adopting it for whole-codebase analysis.
Frequently Asked Questions
What is GPT-5.3 Codex best used for?
Software engineering tasks that benefit from deliberation: multi-file feature implementation, debugging, refactoring, code review, and serving as the reasoning model inside an agentic coding harness. Its long time to first token makes it less appropriate for inline autocomplete or real-time chat.
How much does GPT-5.3 Codex cost?
Pricing varies by provider and by pricing type — input versus output tokens, cached input, and any batch or committed-use discounts all differ between hosts. Check the pricing table on this page for current per-provider rates.
How fast is GPT-5.3 Codex?
Third-party testing by Artificial Analysis measured about 121.4 output tokens per second with a time to first token of roughly 37.6 seconds. In practice, expect a noticeable pause before the response begins, followed by steady generation.
How does GPT-5.3 Codex differ from OpenAI's general GPT-5.x models?
The Codex label marks the branch of OpenAI's lineup tuned for code and software engineering agents rather than general assistant use. General GPT-5.x models are typically the better default for mixed workloads such as writing, summarization, and open-domain question answering.
What context window does GPT-5.3 Codex support?
We do not have a confirmed context window figure for this model in our database. Consult OpenAI's model documentation or your serving provider's specifications, and validate against your own repository sizes before relying on it for large-context work.