Qwen 3.5 122B is Alibaba's flagship multimodal model supporting text, image, and video inputs with a 262K token context window.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.290 | $2.40 | |
| $0.290 | $2.40 | |
| $0.400 | $3.20 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Qwen 3.5 122B is designed for complex multimodal applications requiring analysis of text, images, and video content within a single workflow. Its large context window makes it suitable for document analysis combined with visual elements, content moderation across multiple media types, educational applications involving multimedia materials, and research tasks requiring comprehensive understanding of mixed content formats. The model's flagship positioning and video capabilities make it appropriate for media analysis, content creation workflows, and enterprise applications where multimodal understanding is essential for business processes.
Qwen 3.5 122B pricing varies by provider and may include different rates for text and image tokens. Check the pricing table above for current rates across all available providers.
Qwen 3.5 122B excels at multimodal tasks involving text, image, and video analysis. Its large context window and video understanding capabilities make it well-suited for content analysis, document processing with visual elements, educational applications, and enterprise workflows requiring comprehensive multimedia understanding.
No, Qwen 3.5 122B does not support function calling or tool use capabilities. It focuses on multimodal understanding and generation tasks rather than agentic workflows that require external tool integration.