Grok 4.6
Grok 4.6 is a chat-oriented large language model in the Grok family, listed in our catalog under creator SpaceXAI and tracked across inference API providers.
API Pricing
Cheapest on Velokey — 55% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.800 | $2.40 | $0.080 | |
| $2.00 | $6.00 | $0.500 | |
| $2.00 | $6.00 | - | |
| $2.00 | $6.00 | $0.500 | |
| $2.00 | $6.00 | $0.500 |
Prices updated daily. Last check: Oct 11, 2026
Compare API pricing for every Grok model →Grok 4.6 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond94.9%
- Humanity's Last Exam42.9%
Coding
- SciCode56.5%
Agentic & Tool Use
- Terminal-Bench v2.188.4%
- τ-bench Banking50.7%
Instruction & Long Context
- Long-Context Reasoning80.3%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- SpaceXAI
- Modalities
- Text
Capabilities
- Open Source
- No
Strengths & Limitations
Strengths
- Part of the Grok 4 generation lineage, so prompting patterns and integration work carry over from earlier Grok 4.x deployments
- Independently measured by Artificial Analysis, giving a third-party throughput and latency reference rather than vendor-only claims
- Measured output speed of roughly 53.7 tokens per second, adequate for streaming chat interfaces where text is rendered as it arrives
- Available through API endpoints tracked in the pricing table on this page, allowing per-token cost comparison across hosts
- Point-release positioning (4.6) means it can often be swapped in for an earlier 4.x model with minimal code changes
- Chat-formatted interface, so it fits standard message-role request patterns used by most LLM SDKs
Limitations
- Time to first token measured at roughly 3,958 ms, which is slow for latency-sensitive interactive or voice applications
- Output speed of about 53.7 tokens per second is moderate, so long generations take noticeably longer than on high-throughput models
- We do not track a confirmed context window size for Grok 4.6 — verify the limit with your chosen provider before designing long-document workflows
- We do not have confirmed modality data, so image or audio input support should not be assumed without checking provider documentation
- Provider availability for Grok-family models is narrower than for the most widely mirrored open-weight models, which limits fallback routing options
Key Features
About Grok 4.6
Common Use Cases
Grok 4.6 suits asynchronous or semi-interactive chat workloads where a few seconds of initial thinking time is acceptable: drafting and rewriting text, question answering over pasted context, code explanation and review, research summarization, and back-office automation where responses are consumed by a queue or a human reviewer rather than a real-time UI. Its measured time to first token of roughly four seconds makes it a poorer fit for voice agents, autocomplete, or typeahead features where sub-second responsiveness matters, and its moderate output rate means very long generations should be chunked or streamed rather than awaited in a single blocking call. Teams already running an earlier Grok 4.x model are the most natural audience, since Grok 4.6 can typically be evaluated as a direct substitute in an existing pipeline. Because we do not track confirmed context window or modality figures for this model, any workload that depends on very long inputs or on image understanding should be validated against the specific provider's documentation first.
Frequently Asked Questions
How much does Grok 4.6 cost to use?
Pricing for Grok 4.6 varies by provider and by pricing type — input and output tokens are usually billed at different rates, and some hosts offer batch or cached-input tiers. Because rates change frequently, we do not quote figures in this write-up; see the pricing table on this page for current per-provider rates and compare them against other models in the catalog.
What is Grok 4.6 best used for?
It fits chat and text-generation workloads that tolerate a few seconds of startup latency: drafting and editing content, question answering over supplied context, code explanation, summarization, and background automation tasks. It is less suited to latency-critical interfaces such as voice assistants or inline code completion, given a measured time to first token of around four seconds.
How fast is Grok 4.6?
According to Artificial Analysis measurements in our database, Grok 4.6 generates roughly 53.7 output tokens per second with a time to first token of about 3,958 ms. Real-world numbers depend on the provider, prompt length, region, and current load, so treat these as a reference baseline and benchmark your own traffic.
How does Grok 4.6 differ from earlier Grok 4.x models?
The 4.6 designation marks it as a later revision within the same Grok 4 generation rather than a new architecture family, so it is generally evaluated as a substitute for an earlier 4.x deployment. We do not track a detailed feature-level changelog between the versions; consult the provider's release notes for specifics and run a side-by-side evaluation on your own prompts.
Does Grok 4.6 support image input or tool calling?
Our catalog does not carry confirmed modality or tool-calling data for Grok 4.6, so we cannot state either way. Capability exposure also differs between first-party and third-party hosts, so check the documentation of the specific provider listed in the pricing table before building against those features.
What context window does Grok 4.6 have?
We do not track a confirmed context window figure for this model. Providers sometimes serve the same model with different maximum context limits, so verify the supported input length with your chosen host if your workload involves long documents or extended conversation history.