Skip to main content
SpaceXAI

Grok 4.5

Grok 4.5 is a chat model in the Grok family, listed in our catalog under creator SpaceXAI, tracked here for inference API pricing and throughput.

Input from
$0.800 / 1M tokens
across 4 providers

API Pricing

Cheapest on Velokey — 53% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.800$2.40$0.200
$2.00$6.00$0.500
$2.00$6.00$0.300
$2.00$6.00$0.300

Prices updated daily. Last check: Oct 11, 2026

Compare API pricing for every Grok model →

Grok 4.5 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
38.8 / 100
Coding
72.4 / 100

Reasoning & Knowledge

  • GPQA Diamond93.1%
  • Humanity's Last Exam42.7%

Coding

  • SciCode55.0%

Agentic & Tool Use

  • Terminal-Bench v2.181.6%
  • τ-bench Banking42.1%

Instruction & Long Context

  • Long-Context Reasoning79.3%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
SpaceXAI
Modalities
Text

Capabilities

Open Source
No

Strengths & Limitations

Strengths

  • Measured output throughput of approximately 53.8 tokens per second in independent Artificial Analysis benchmarking
  • Throughput and latency figures come from a third-party benchmark rather than vendor-published claims
  • Part of the Grok family, so prompts and workflows built for earlier Grok releases are a reasonable starting point
  • Served through inference APIs, so usage is pay-per-token with no GPU provisioning required
  • Provider-level pricing for this model is tracked and compared directly in the table on this page
  • Suited to asynchronous generation workloads where sustained output speed matters more than first-token latency

Limitations

  • Time to first token of roughly 13.1 seconds is high, making it a poor fit for latency-sensitive interactive interfaces
  • Context window size is not tracked in our database, so long-document limits must be confirmed with the provider
  • Input and output modality support is not tracked in our data — do not assume image or audio input
  • Tool calling and structured output support are not confirmed in our metadata
  • No task-level benchmark scores (reasoning, coding, math) are recorded for this entry in our database

Key Features

•Text chat and instruction-following interface
•Grok family model generation 4.5
•Measured output speed of about 53.8 tokens per second (Artificial Analysis)
•Measured time to first token of about 13.1 seconds (Artificial Analysis)
•Streaming token output via hosted inference APIs
•Per-token API billing across tracked providers
•Cross-provider price comparison on this page

About Grok 4.5

Grok 4.5 is a text chat model in the Grok model family, catalogued on this page under the creator entry SpaceXAI. It sits alongside earlier Grok releases as a general-purpose conversational and instruction-following model served through inference APIs, and this page tracks the providers that host it along with their pricing. Our measured throughput data for Grok 4.5, sourced from Artificial Analysis, shows roughly 53.8 output tokens per second and a time to first token of about 13.1 seconds. The long first-token latency combined with moderate sustained output speed is a pattern typically seen with models that perform extended internal processing before emitting a response, which matters more for interactive chat than for batch workloads. Other specifications for this entry — context window, input and output modalities, and tool-calling support — are not currently tracked in our database, so we do not make claims about them here; check the provider's own documentation for those details before building against the model. In practice, a model with this latency profile is most comfortably used for asynchronous or non-real-time work: drafting, analysis, summarization, and code generation where a multi-second wait before the first token is acceptable. Compared with lighter-weight chat models in our catalog that begin streaming in well under a second, Grok 4.5 trades responsiveness for whatever additional processing happens before generation begins. Readers comparing options should weigh the throughput numbers on this page against their own latency budget.

Common Use Cases

Grok 4.5 fits text generation and analysis workloads where a multi-second delay before the first token is tolerable: long-form drafting, document summarization, code generation and review, research synthesis, and batch processing of prompts through a queue or background job. Its sustained output rate of roughly 54 tokens per second is adequate for producing multi-paragraph responses, so content pipelines and offline evaluation harnesses are a natural match. It is a weaker choice for voice interfaces, autocomplete, live chat widgets, or agent loops that chain many short calls, since the ~13 second time to first token compounds across each step — for those patterns, compare against lower-latency models in the catalog. Because context window and modality support are not tracked in our data for this entry, confirm those specifications with your chosen provider before committing to workloads involving very long inputs or non-text media.

Frequently Asked Questions

How much does Grok 4.5 cost to use?

Pricing depends on which provider is serving the model, whether you are billed for input or output tokens, and the pricing type on offer. Because those rates change frequently, we do not quote figures in this description — see the pricing table on this page for the current per-provider comparison.

What is Grok 4.5 best used for?

Text workloads where throughput matters more than responsiveness: long-form writing, summarization, code generation, research synthesis, and batch or background processing. Its measured time to first token of about 13 seconds makes it less suitable for real-time chat, voice, or autocomplete experiences.

How fast is Grok 4.5?

Artificial Analysis measurements recorded in our database show roughly 53.8 output tokens per second with a time to first token of about 13.1 seconds. That means responses stream at a usable pace once they start, but there is a noticeable wait before the first token appears.

What context window does Grok 4.5 support?

Our database does not currently track a context window figure for this entry, so we do not state one. Check the documentation of the provider you plan to use — context limits can also differ between providers serving the same model.

Does Grok 4.5 accept images or other non-text input?

We do not track modality details for this entry, so we cannot confirm what input types are supported. Treat it as a text chat model unless the provider's documentation states otherwise.

How does Grok 4.5 compare to other models in the catalog?

On the data we hold, its distinguishing characteristic is the latency profile: moderate sustained output speed paired with a long time to first token. Many general-purpose chat models begin streaming in under a second, so if interactive responsiveness is a requirement, compare those alternatives; if you are generating long outputs asynchronously, the difference matters less.