Qwen 2.5 VL 72B is Alibaba's flagship multimodal model supporting text and image inputs with a 128K token context window and tool calling capabilities.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.800 | $1.00 | $0.400 | |
| $1.95 | $8.00 | - |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Qwen 2.5 VL 72B is well-suited for complex multimodal applications requiring both visual and textual understanding. Its capabilities make it ideal for document analysis involving charts, graphs, and mixed media content, visual question answering systems, and educational applications that need to process textbook pages or technical diagrams. The model's tool calling features enable integration into agentic workflows for tasks like automated report generation from visual data or multimodal content creation. Organizations prioritizing data privacy or requiring customization benefit from its open-source nature, allowing for on-premises deployment and fine-tuning for specific domains like medical imaging analysis, technical documentation processing, or multimodal customer service applications.
Qwen 2.5 VL 72B pricing varies by provider and deployment method, with different rates for hosted API access versus self-hosting the open-source model. Check the pricing table above for current rates across all providers offering this model.
Qwen 2.5 VL 72B excels at multimodal tasks requiring analysis of both text and images, such as document understanding with visual elements, chart and graph interpretation, visual question answering, and educational content processing. Its tool calling capabilities make it suitable for building agents that can interact with external services while processing multimodal inputs.
Yes, Qwen 2.5 VL 72B is open source with publicly available model weights, allowing for self-hosting and on-premises deployment. However, the 72B parameter size requires substantial computational resources including high-memory GPUs and significant storage capacity for optimal performance.