Grok 4.20 is xAI's flagship multimodal model with text and image capabilities, featuring an extensive 2 million token context window for large-scale processing tasks.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.25 | $2.50 | $0.200 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Grok 4.20 is designed for applications requiring extensive context processing and multimodal understanding. Its 2 million token context window makes it particularly valuable for analyzing large documents, entire codebases, research papers, or legal documents in a single session. The model excels at tasks involving lengthy conversations, document summarization, code analysis across multiple files, and complex reasoning tasks that benefit from maintaining extensive context. Organizations dealing with large-scale text analysis, document processing, or applications requiring comprehensive understanding of substantial information volumes will find the model's context capabilities advantageous. The multimodal support also enables use cases involving document analysis with embedded images or visual content understanding alongside text processing.
Grok 4.20 pricing varies by provider and may include different rates for input and output tokens. Check the pricing table above for current rates across all available providers offering this model.
Grok 4.20 excels at tasks requiring extensive context processing, such as analyzing large documents, codebases, or maintaining coherent long conversations. Its 2 million token context window and multimodal capabilities make it ideal for document analysis, research paper processing, code review across multiple files, and complex reasoning tasks that benefit from substantial context retention.
No, Grok 4.20 does not currently support function calling or tool execution capabilities. The model focuses on text generation, multimodal understanding, and large context processing rather than external tool integration or structured function execution.