Skip to main content
Alibaba

Qwen 3.6 27B

Qwen 3.6 27B is a 27-billion-parameter multimodal model from Alibaba's Qwen 3.6 family, accepting text, image, and video input with a 262K token context window.

Context 262K
Modalities text, image, video
Input from
$0.320 / 1M tokens
across 5 providers

API Pricing

Cheapest on Deep Infra — 25% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.320$3.20-
$0.320$2.70$0.150
$0.343$3.01$0.172
$0.560$3.70-
$0.600$3.60-

Prices updated daily. Last check: Sep 25, 2026

Qwen 3.6 27B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
19.8 / 100
Coding
46.6 / 100

Reasoning & Knowledge

  • GPQA Diamond82.9%
  • Humanity's Last Exam15.1%

Agentic & Tool Use

  • Terminal-Bench Hard21.2%
  • Terminal-Bench v2.151.3%
  • τ²-bench93.6%
  • τ-bench Banking9.3%

Instruction & Long Context

  • IFBench45.7%
  • Long-Context Reasoning66.7%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Alibaba
Family
Qwen 3.6
Context Window
262K
Modalities
Text, Image, Video

Capabilities

Tool Calling
No
Open Source
No
Aliases
qwen3.6-27b, Qwen3.6-27B, Qwen/Qwen3.6-27B

Strengths & Limitations

Strengths

  • 262,144 token context window supports long documents, large codebases, and extended conversation histories in one request
  • Accepts image input in addition to text, enabling document, chart, and screenshot analysis
  • Accepts video input, which is less common among models in this parameter range
  • 27B parameter size sits between the small and large tiers of the Qwen 3.6 family, giving a middle capability/cost point
  • Measured at roughly 57 output tokens per second by Artificial Analysis, a workable rate for interactive applications
  • Available under multiple provider aliases (Qwen/Qwen3.6-27B), so buyers can compare hosts rather than depend on a single endpoint

Limitations

  • Smaller parameter count than the larger models in the Qwen 3.6 family, so complex reasoning and long agentic chains may be weaker
  • Time to first token measured around 1,526 ms, which is noticeable in latency-sensitive chat or voice use
  • Throughput and latency vary by inference provider, so benchmark figures are not guaranteed on any given endpoint
  • We do not track detailed task benchmark scores (coding, math, reasoning) for this model, making direct capability comparisons harder
  • Video input at long context lengths can consume tokens quickly, raising effective cost per request

Key Features

•262K (262,144) token context window
•Multimodal input: text, images, and video
•27B parameter mid-size configuration in the Qwen 3.6 family
•Measured throughput of approximately 57 output tokens per second (Artificial Analysis)
•Measured time to first token of approximately 1,526 ms (Artificial Analysis)
•Served under multiple provider aliases including qwen3.6-27b and Qwen/Qwen3.6-27B
•Long-context document and multi-frame video analysis in a single prompt

About Qwen 3.6 27B

Qwen 3.6 27B is a model in Alibaba's Qwen 3.6 family, positioned as a mid-sized option at roughly 27 billion parameters. It sits between the smaller, cheaper members of the family and the larger high-parameter variants, making it a middle option for teams that want more capability than a compact model provides without moving to the largest configurations in the lineup. The model accepts text, image, and video input and supports a context window of 262,144 tokens, which is large enough for long documents, extended chat histories, sizable code repositories, or multi-frame video analysis in a single request. Independent measurements from Artificial Analysis put throughput at approximately 57 output tokens per second with a time to first token of about 1,526 ms, though both figures vary meaningfully depending on which provider serves the model and under what load. In practice, Qwen 3.6 27B is used for tasks that combine long-context reading with visual understanding — document and screenshot analysis, video summarization, and multimodal assistants — as well as general text work such as summarization, extraction, and coding assistance. Compared with larger models in the Qwen 3.6 family, it trades some capability headroom for lower resource requirements; compared with the family's smaller variants, it offers more capacity at higher per-token cost and latency. Because Qwen models are frequently served by multiple inference providers, throughput and pricing for this model can differ substantially between hosts.

Common Use Cases

Qwen 3.6 27B suits workloads that need multimodal input paired with a large context window: analyzing long PDFs alongside their figures and tables, summarizing or answering questions about video content, processing UI screenshots for automation and QA, and building assistants that reason over extended conversation history. Its mid-tier position in the Qwen 3.6 family makes it a reasonable default for general-purpose text tasks — summarization, structured extraction, translation, and coding assistance — where a compact model is not accurate enough but the largest family members are more capacity than the task requires. For very high-volume, latency-critical classification, a smaller Qwen 3.6 variant may be a better fit; for deep multi-step reasoning or long agentic workflows, the larger models in the family are worth evaluating first.

Frequently Asked Questions

How much does Qwen 3.6 27B cost to run?

Pricing depends on which inference provider you use and whether you are paying per token, per hour of dedicated capacity, or hosting the model yourself on rented GPUs. Because Qwen models are served by many hosts, rates can differ significantly between them. Check the pricing table on this page for current per-provider figures.

What is Qwen 3.6 27B best used for?

It works well for long-context multimodal tasks — document and screenshot analysis, video summarization, and question answering over large inputs — plus general text work like summarization, extraction, and coding help. The 262K context window and image/video input are its main differentiators at this size.

Can Qwen 3.6 27B process images and video?

Yes. Our metadata confirms text, image, and video input modalities, so it can handle image-based tasks such as chart and document reading as well as video understanding, subject to the limits of the specific provider serving the model.

How does it compare to other models in the Qwen 3.6 family?

At 27B parameters it is a mid-sized member of the family. Larger Qwen 3.6 models generally offer more headroom on complex reasoning and coding, while smaller ones are cheaper and faster for high-volume simple tasks. The 27B variant is a middle point between the two.

How fast is Qwen 3.6 27B?

Artificial Analysis measured roughly 57 output tokens per second with a time to first token of about 1,526 ms. These are provider-dependent measurements — actual speed varies with the host, hardware, batch size, and prompt length, so test on your chosen endpoint.