Qwen 3.6 27B
Qwen 3.6 27B is a 27-billion-parameter multimodal model from Alibaba's Qwen 3.6 family, accepting text, image, and video input with a 262K token context window.
API Pricing
Cheapest on Deep Infra — 25% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.320 | $3.20 | - | |
| $0.320 | $2.70 | $0.150 | |
| $0.343 | $3.01 | $0.172 | |
| $0.560 | $3.70 | - | |
| $0.600 | $3.60 | - |
Prices updated daily. Last check: Sep 25, 2026
Qwen 3.6 27B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond82.9%
- Humanity's Last Exam15.1%
Agentic & Tool Use
- Terminal-Bench Hard21.2%
- Terminal-Bench v2.151.3%
- τ²-bench93.6%
- τ-bench Banking9.3%
Instruction & Long Context
- IFBench45.7%
- Long-Context Reasoning66.7%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Alibaba
- Family
- Qwen 3.6
- Context Window
- 262K
- Modalities
- Text, Image, Video
Capabilities
- Tool Calling
- No
- Open Source
- No
- Aliases
- qwen3.6-27b, Qwen3.6-27B, Qwen/Qwen3.6-27B
Strengths & Limitations
Strengths
- 262,144 token context window supports long documents, large codebases, and extended conversation histories in one request
- Accepts image input in addition to text, enabling document, chart, and screenshot analysis
- Accepts video input, which is less common among models in this parameter range
- 27B parameter size sits between the small and large tiers of the Qwen 3.6 family, giving a middle capability/cost point
- Measured at roughly 57 output tokens per second by Artificial Analysis, a workable rate for interactive applications
- Available under multiple provider aliases (Qwen/Qwen3.6-27B), so buyers can compare hosts rather than depend on a single endpoint
Limitations
- Smaller parameter count than the larger models in the Qwen 3.6 family, so complex reasoning and long agentic chains may be weaker
- Time to first token measured around 1,526 ms, which is noticeable in latency-sensitive chat or voice use
- Throughput and latency vary by inference provider, so benchmark figures are not guaranteed on any given endpoint
- We do not track detailed task benchmark scores (coding, math, reasoning) for this model, making direct capability comparisons harder
- Video input at long context lengths can consume tokens quickly, raising effective cost per request
Key Features
About Qwen 3.6 27B
Common Use Cases
Qwen 3.6 27B suits workloads that need multimodal input paired with a large context window: analyzing long PDFs alongside their figures and tables, summarizing or answering questions about video content, processing UI screenshots for automation and QA, and building assistants that reason over extended conversation history. Its mid-tier position in the Qwen 3.6 family makes it a reasonable default for general-purpose text tasks — summarization, structured extraction, translation, and coding assistance — where a compact model is not accurate enough but the largest family members are more capacity than the task requires. For very high-volume, latency-critical classification, a smaller Qwen 3.6 variant may be a better fit; for deep multi-step reasoning or long agentic workflows, the larger models in the family are worth evaluating first.
Frequently Asked Questions
How much does Qwen 3.6 27B cost to run?
Pricing depends on which inference provider you use and whether you are paying per token, per hour of dedicated capacity, or hosting the model yourself on rented GPUs. Because Qwen models are served by many hosts, rates can differ significantly between them. Check the pricing table on this page for current per-provider figures.
What is Qwen 3.6 27B best used for?
It works well for long-context multimodal tasks — document and screenshot analysis, video summarization, and question answering over large inputs — plus general text work like summarization, extraction, and coding help. The 262K context window and image/video input are its main differentiators at this size.
Can Qwen 3.6 27B process images and video?
Yes. Our metadata confirms text, image, and video input modalities, so it can handle image-based tasks such as chart and document reading as well as video understanding, subject to the limits of the specific provider serving the model.
How does it compare to other models in the Qwen 3.6 family?
At 27B parameters it is a mid-sized member of the family. Larger Qwen 3.6 models generally offer more headroom on complex reasoning and coding, while smaller ones are cheaper and faster for high-volume simple tasks. The 27B variant is a middle point between the two.
How fast is Qwen 3.6 27B?
Artificial Analysis measured roughly 57 output tokens per second with a time to first token of about 1,526 ms. These are provider-dependent measurements — actual speed varies with the host, hardware, batch size, and prompt length, so test on your chosen endpoint.