Llama 4 Maverick 17B is Meta's lightweight multimodal model supporting text and image inputs with a 128K token context window and tool calling capabilities.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.120 | $0.485 | - | |
| $0.200 | $0.800 | - | |
| $0.200 | $0.696 | - | |
| $0.200 | $0.696 | - | |
| $0.240 | $0.970 | - | |
| $0.270 | $0.850 | - | |
| $0.274 | $0.899 | $0.137 |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Llama 4 Maverick 17B is designed for applications requiring multimodal processing at scale, such as content moderation systems that analyze both text and images, customer support chatbots handling visual queries, and document processing workflows involving charts and diagrams. Its lightweight architecture makes it suitable for high-volume deployments where cost efficiency matters, while the open-source licensing enables custom fine-tuning for domain-specific applications. The fast inference speed and tool calling capabilities support interactive applications and automated workflows that need to process visual content alongside text with minimal latency.
Llama 4 Maverick 17B pricing varies by provider and pricing type (standard vs batch). Check the pricing table above for current rates across all providers.
Llama 4 Maverick 17B excels at high-volume applications requiring multimodal processing, such as content moderation, customer support with visual elements, and document analysis. Its lightweight design and fast inference make it ideal for interactive applications where response speed matters more than maximum reasoning capability.
Yes, Llama 4 Maverick 17B is open source with full model weights available, allowing complete customization through fine-tuning. This makes it suitable for domain-specific applications where proprietary models cannot be modified, though you'll need appropriate computational resources for the fine-tuning process.