GPT-5.6 Terra is a chat model from OpenAI, tracked here with third-party throughput and latency measurements alongside provider pricing.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.00 | $6.00 | $0.100 | |
| $1.00 | $6.00 | $0.100 | |
| $2.00 | $12.00 | $0.200 | |
| $2.00 | $12.00 | $0.200 | |
| $2.00 | $12.00 | $0.200 | |
| $2.00 | $12.00 | $0.200 |
Prices updated daily. Last check: Sep 5, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
GPT-5.6 Terra fits general chat and instruction-following workloads where sustained generation speed matters more than instant first-token response: drafting and rewriting text, summarization, question answering over supplied context, customer-support assistants, and internal knowledge tools. The near-100 tokens per second output rate is fast enough that long answers render at a natural reading pace, while the roughly two-second startup delay makes it a weaker fit for autocomplete, voice-turn-taking, or other interactions where users notice sub-second lag. Batch and asynchronous jobs — bulk summarization, content generation queues, evaluation harnesses — are unaffected by the first-token delay and benefit from the throughput. Because we do not track a confirmed context window or modality set for this model, teams planning long-context retrieval pipelines or image-input workflows should confirm those specifications with OpenAI before committing.
Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are usually billed separately, and providers change their rates over time. Check the pricing table on this page for current tracked figures rather than relying on a fixed number.
It is a chat model, so it suits conversational assistants, drafting and editing, summarization, and question answering. Its measured throughput of about 98.8 tokens per second makes it comfortable for streaming long responses, while the roughly 2.2-second time to first token makes it less suited to autocomplete or voice interfaces where startup latency is noticeable.
Artificial Analysis measured approximately 98.8 output tokens per second with a time to first token of about 2,156 milliseconds. Actual figures depend on the provider, prompt length, region, and load at the time of the request.
We do not currently have a confirmed context window recorded for this model in our database. Consult OpenAI's model documentation or your serving provider for the authoritative limit before designing long-context or retrieval-heavy applications.
Our entry for GPT-5.6 Terra carries throughput and latency measurements but no verified quality benchmarks, so a capability comparison cannot be made from this page alone. On speed, you can compare its ~98.8 tokens per second and ~2,156 ms first-token latency directly against the figures listed on other model pages in this catalog, and compare cost using the pricing table above.