Skip to main content
OpenAI

GPT-5.6 Terra

GPT-5.6 Terra is a chat model from OpenAI, tracked here with third-party throughput and latency measurements alongside provider pricing.

Input from
$1.00 / 1M tokens
across 2 providers

API Pricing

Cheapest on OpenAI 33% below avg
ProviderInput / 1MOutput / 1MCached / 1M
OpenAI logo
OpenAIBatch
$1.00$6.00$0.100
$1.00$6.00$0.100
$2.00$12.00$0.200
$2.00$12.00$0.200

Prices updated daily. Last check: Aug 26, 2026

GPT-5.6 Terra pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
56.6 / 100
Coding
76.7 / 100
Output Speed
109 t/s
Latency (TTFT)
63.8s

Reasoning & Knowledge

  • GPQA Diamond92.5%
  • Humanity's Last Exam42.9%

Coding

  • SciCode53.9%

Agentic & Tool Use

  • Terminal-Bench Hard57.6%
  • Terminal-Bench v2.188.0%
  • τ²-bench86.3%
  • τ-bench Banking40.2%

Instruction & Long Context

  • IFBench71.2%
  • Long-Context Reasoning79.7%

Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
OpenAI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured output throughput of approximately 98.8 tokens per second in Artificial Analysis testing, suitable for streaming chat interfaces
  • Built by OpenAI, so it is generally reachable through the widely supported OpenAI-compatible chat completions request format
  • Independent third-party performance data is available rather than only vendor-published figures
  • Positioned in the GPT-5.x generation, keeping it aligned with OpenAI's current API conventions and SDKs
  • Chat-tuned, so it can be dropped into existing assistant and instruction-following pipelines without a separate prompt format
  • Pricing across available providers is tracked live on this page for direct comparison against peer models

Limitations

  • Time to first token of roughly 2,156 ms adds a visible pause before responses begin streaming, which is a drawback for latency-sensitive UIs
  • We do not currently track a confirmed context window for this model, so long-document workloads need verification against OpenAI's documentation
  • No verified quality benchmark scores are recorded in our database, making capability comparisons against siblings difficult from this page alone
  • Supported input and output modalities are not confirmed in our data
  • Availability may be limited to a small number of providers, reducing routing and failover options

Key Features

Chat completions interface for multi-turn conversational use
Measured output speed of ~98.8 tokens per second (Artificial Analysis)
Measured time to first token of ~2,156 ms (Artificial Analysis)
Token streaming for incremental response rendering
Part of OpenAI's GPT-5.x model generation
Accessible through OpenAI-compatible client libraries and tooling
Live multi-provider pricing tracked on this page

About GPT-5.6 Terra

GPT-5.6 Terra is a chat-oriented large language model from OpenAI, listed in our catalog as part of the GPT-5.x line. Our database entry for this model is intentionally sparse: we record the creator, the model type, and independently measured performance figures, and we do not currently track a confirmed context window, modality list, or published benchmark scores for it. Where those details are absent below, that reflects gaps in our data rather than an absence of the capability in the model itself. The performance data we do carry comes from Artificial Analysis, which measured GPT-5.6 Terra at roughly 98.8 output tokens per second with a time to first token of about 2,156 milliseconds. That combination — a healthy sustained generation rate paired with a first-token delay of just over two seconds — is characteristic of models that do some amount of work before the first token appears. In practice it means the model streams responses at a comfortable reading pace once it starts, but interfaces that feel snappy on very short prompts will show a noticeable pause at the beginning of each turn. As a chat model, GPT-5.6 Terra is intended for conversational and instruction-following workloads: assistants, drafting, question answering, summarization, and similar text-in/text-out tasks. Buyers comparing it against other OpenAI entries or third-party alternatives on this site should weigh the measured throughput and latency against provider pricing shown in the table on this page, and consult OpenAI's own documentation for the authoritative specification of context length, supported modalities, and API features.

Common Use Cases

GPT-5.6 Terra fits general chat and instruction-following workloads where sustained generation speed matters more than instant first-token response: drafting and rewriting text, summarization, question answering over supplied context, customer-support assistants, and internal knowledge tools. The near-100 tokens per second output rate is fast enough that long answers render at a natural reading pace, while the roughly two-second startup delay makes it a weaker fit for autocomplete, voice-turn-taking, or other interactions where users notice sub-second lag. Batch and asynchronous jobs — bulk summarization, content generation queues, evaluation harnesses — are unaffected by the first-token delay and benefit from the throughput. Because we do not track a confirmed context window or modality set for this model, teams planning long-context retrieval pipelines or image-input workflows should confirm those specifications with OpenAI before committing.

Frequently Asked Questions

How much does GPT-5.6 Terra cost?

Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are usually billed separately, and providers change their rates over time. Check the pricing table on this page for current tracked figures rather than relying on a fixed number.

What is GPT-5.6 Terra best used for?

It is a chat model, so it suits conversational assistants, drafting and editing, summarization, and question answering. Its measured throughput of about 98.8 tokens per second makes it comfortable for streaming long responses, while the roughly 2.2-second time to first token makes it less suited to autocomplete or voice interfaces where startup latency is noticeable.

How fast is GPT-5.6 Terra?

Artificial Analysis measured approximately 98.8 output tokens per second with a time to first token of about 2,156 milliseconds. Actual figures depend on the provider, prompt length, region, and load at the time of the request.

What context window does GPT-5.6 Terra support?

We do not currently have a confirmed context window recorded for this model in our database. Consult OpenAI's model documentation or your serving provider for the authoritative limit before designing long-context or retrieval-heavy applications.

How does GPT-5.6 Terra compare to other GPT-5.x models?

Our entry for GPT-5.6 Terra carries throughput and latency measurements but no verified quality benchmarks, so a capability comparison cannot be made from this page alone. On speed, you can compare its ~98.8 tokens per second and ~2,156 ms first-token latency directly against the figures listed on other model pages in this catalog, and compare cost using the pricing table above.