Grok 4.5
Grok 4.5 is a chat model in the Grok family, listed in our catalog under creator SpaceXAI, tracked here for inference API pricing and throughput.
API Pricing
Cheapest on Velokey — 53% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.800 | $2.40 | $0.200 | |
| $2.00 | $6.00 | $0.500 | |
| $2.00 | $6.00 | $0.300 | |
| $2.00 | $6.00 | $0.300 |
Prices updated daily. Last check: Oct 11, 2026
Compare API pricing for every Grok model →Grok 4.5 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond93.1%
- Humanity's Last Exam42.7%
Coding
- SciCode55.0%
Agentic & Tool Use
- Terminal-Bench v2.181.6%
- τ-bench Banking42.1%
Instruction & Long Context
- Long-Context Reasoning79.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- SpaceXAI
- Modalities
- Text
Capabilities
- Open Source
- No
Strengths & Limitations
Strengths
- Measured output throughput of approximately 53.8 tokens per second in independent Artificial Analysis benchmarking
- Throughput and latency figures come from a third-party benchmark rather than vendor-published claims
- Part of the Grok family, so prompts and workflows built for earlier Grok releases are a reasonable starting point
- Served through inference APIs, so usage is pay-per-token with no GPU provisioning required
- Provider-level pricing for this model is tracked and compared directly in the table on this page
- Suited to asynchronous generation workloads where sustained output speed matters more than first-token latency
Limitations
- Time to first token of roughly 13.1 seconds is high, making it a poor fit for latency-sensitive interactive interfaces
- Context window size is not tracked in our database, so long-document limits must be confirmed with the provider
- Input and output modality support is not tracked in our data — do not assume image or audio input
- Tool calling and structured output support are not confirmed in our metadata
- No task-level benchmark scores (reasoning, coding, math) are recorded for this entry in our database
Key Features
About Grok 4.5
Common Use Cases
Grok 4.5 fits text generation and analysis workloads where a multi-second delay before the first token is tolerable: long-form drafting, document summarization, code generation and review, research synthesis, and batch processing of prompts through a queue or background job. Its sustained output rate of roughly 54 tokens per second is adequate for producing multi-paragraph responses, so content pipelines and offline evaluation harnesses are a natural match. It is a weaker choice for voice interfaces, autocomplete, live chat widgets, or agent loops that chain many short calls, since the ~13 second time to first token compounds across each step — for those patterns, compare against lower-latency models in the catalog. Because context window and modality support are not tracked in our data for this entry, confirm those specifications with your chosen provider before committing to workloads involving very long inputs or non-text media.
Frequently Asked Questions
How much does Grok 4.5 cost to use?
Pricing depends on which provider is serving the model, whether you are billed for input or output tokens, and the pricing type on offer. Because those rates change frequently, we do not quote figures in this description — see the pricing table on this page for the current per-provider comparison.
What is Grok 4.5 best used for?
Text workloads where throughput matters more than responsiveness: long-form writing, summarization, code generation, research synthesis, and batch or background processing. Its measured time to first token of about 13 seconds makes it less suitable for real-time chat, voice, or autocomplete experiences.
How fast is Grok 4.5?
Artificial Analysis measurements recorded in our database show roughly 53.8 output tokens per second with a time to first token of about 13.1 seconds. That means responses stream at a usable pace once they start, but there is a noticeable wait before the first token appears.
What context window does Grok 4.5 support?
Our database does not currently track a context window figure for this entry, so we do not state one. Check the documentation of the provider you plan to use — context limits can also differ between providers serving the same model.
Does Grok 4.5 accept images or other non-text input?
We do not track modality details for this entry, so we cannot confirm what input types are supported. Treat it as a text chat model unless the provider's documentation states otherwise.
How does Grok 4.5 compare to other models in the catalog?
On the data we hold, its distinguishing characteristic is the latency profile: moderate sustained output speed paired with a long time to first token. Many general-purpose chat models begin streaming in under a second, so if interactive responsiveness is a requirement, compare those alternatives; if you are generating long outputs asynchronously, the difference matters less.