Grok 4.7 is a chat model in the Grok family from xAI (listed in our catalog as SpaceXAI), served through hosted inference APIs and tracked here for cross-provider pricing.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.60 | $4.80 | $0.400 |
Prices updated daily. Last check: Sep 22, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Grok 4.7 fits general-purpose assistant workloads where one model handles a wide mix of tasks: conversational support agents, drafting and rewriting text, summarizing documents or threads, answering questions over supplied context, and coding assistance such as explanation, review, and small refactors. Its sub-second measured time to first token makes it reasonable for interactive front-ends where users see a response begin quickly, while its mid-range output throughput means very long generated outputs will take noticeably longer than on faster-serving small models. For agent pipelines, batch classification, or anything that depends on a specific context length, tool-calling format, or image input, confirm those capabilities directly with the provider you plan to use, since our metadata does not record them for this release.
Pricing depends on the provider and the pricing type — input versus output tokens, cached input, and any batch or committed-use discounts all differ between hosts, and rates change frequently. Check the pricing table on this page for the current per-provider figures rather than relying on a number quoted in an article.
It is a general-purpose chat model, so it suits conversational assistants, drafting and summarization, question answering over provided context, and coding help. Its measured ~712 ms time to first token makes it workable for interactive interfaces where response start time matters.
The Grok series comes from xAI; our database records the creator for this entry as SpaceXAI. It is a point release within the Grok 4.x generation.
Artificial Analysis measurements tracked in our database show roughly 38.35 output tokens per second with a time to first token of about 712 ms. Both figures vary by provider, region, load, and prompt length, so treat them as a reference point rather than a guarantee.
We do not currently track a confirmed context window for this release. Check the model card of the provider you intend to use, since context limits can also differ between hosts serving the same model.