Qwen 3.5 9B is Alibaba's lightweight multimodal model supporting text, image, and video inputs with a 256K token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.100 | $0.150 | - | |
| $0.100 | $0.150 | - | |
| $0.150 | $0.200 | $0.040 | |
| $0.170 | $0.250 | - | |
| $0.170 | $0.250 | - |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Qwen 3.5 9B is well-suited for applications requiring multimodal content processing at scale, including document analysis with embedded images, video content summarization, and educational content creation. Its lightweight architecture and fast inference speeds make it appropriate for real-time applications like customer service chatbots that need to handle mixed media inputs, content moderation systems processing images and videos, and automated transcription services. The large context window supports processing of lengthy documents with multimedia elements, while the efficient performance characteristics enable deployment in cost-sensitive environments where high throughput is prioritized over maximum model capability.
Qwen 3.5 9B pricing varies by provider and pricing type (standard vs batch). Check the pricing table above for current rates across all providers.
Qwen 3.5 9B excels at multimodal content processing tasks including document analysis with images, video summarization, and real-time applications requiring fast inference speeds. Its lightweight architecture makes it suitable for high-throughput scenarios where efficiency is prioritized.
No, Qwen 3.5 9B does not include tool calling or function execution capabilities. It focuses on multimodal understanding and generation tasks across text, image, and video inputs without external tool integration.