Pixtral Large is Mistral's flagship multimodal model supporting text and image inputs with a 131K token context window.
Prices updated daily. Last check: Sep 6, 2026
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Pixtral Large serves applications requiring sophisticated multimodal understanding, particularly where both visual and textual analysis are critical. Its flagship-tier capabilities make it suitable for complex document processing involving charts, diagrams, and text, visual content moderation and analysis, multimodal research applications, and detailed image captioning or visual question answering. The 131K context window enables analysis of lengthy reports with embedded images or processing multiple images alongside extensive text. Organizations typically deploy it for high-value use cases where the superior multimodal reasoning capabilities justify the cost of a flagship model, rather than high-volume applications better suited for lighter alternatives.
Pixtral Large pricing varies by provider and may differ for input versus output tokens, as well as text versus image processing. Check the pricing table above for current rates across all providers offering this model.
Pixtral Large excels at complex multimodal tasks requiring analysis of both text and images, such as processing documents with charts and diagrams, visual content analysis, detailed image captioning, and applications where sophisticated cross-modal reasoning is needed. Its flagship-tier capabilities and 131K context window make it suitable for high-value applications rather than high-volume use cases.
No, Pixtral Large does not support tool calling or function execution capabilities. It focuses on multimodal understanding and generation tasks with text and image inputs, but cannot interact with external tools or APIs through structured function calls.