Qwen 3.6 27B is a 27-billion-parameter multimodal model from Alibaba's Qwen 3.6 family, accepting text, image, and video input with a 262K token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.300 | $2.00 | $0.030 | |
| $0.320 | $3.20 | - | |
| $0.343 | $3.01 | $0.172 | |
| $0.600 | $3.00 | - | |
| $0.600 | $3.60 | - |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Qwen 3.6 27B suits workloads that need multimodal input paired with a large context window: analyzing long PDFs alongside their figures and tables, summarizing or answering questions about video content, processing UI screenshots for automation and QA, and building assistants that reason over extended conversation history. Its mid-tier position in the Qwen 3.6 family makes it a reasonable default for general-purpose text tasks — summarization, structured extraction, translation, and coding assistance — where a compact model is not accurate enough but the largest family members are more capacity than the task requires. For very high-volume, latency-critical classification, a smaller Qwen 3.6 variant may be a better fit; for deep multi-step reasoning or long agentic workflows, the larger models in the family are worth evaluating first.
Pricing depends on which inference provider you use and whether you are paying per token, per hour of dedicated capacity, or hosting the model yourself on rented GPUs. Because Qwen models are served by many hosts, rates can differ significantly between them. Check the pricing table on this page for current per-provider figures.
It works well for long-context multimodal tasks — document and screenshot analysis, video summarization, and question answering over large inputs — plus general text work like summarization, extraction, and coding help. The 262K context window and image/video input are its main differentiators at this size.
Yes. Our metadata confirms text, image, and video input modalities, so it can handle image-based tasks such as chart and document reading as well as video understanding, subject to the limits of the specific provider serving the model.
At 27B parameters it is a mid-sized member of the family. Larger Qwen 3.6 models generally offer more headroom on complex reasoning and coding, while smaller ones are cheaper and faster for high-volume simple tasks. The 27B variant is a middle point between the two.
Artificial Analysis measured roughly 57 output tokens per second with a time to first token of about 1,526 ms. These are provider-dependent measurements — actual speed varies with the host, hardware, batch size, and prompt length, so test on your chosen endpoint.