Skip to main content
Alibaba

Qwen3.8 27B

Qwen3.8 27B is a 27-billion-parameter chat model from Alibaba's Qwen series, listed here so its hosted inference pricing can be compared across providers.

Input from
$0.200 / 1M tokens
across 12 providers

API Pricing

Cheapest on Deep Infra — 60% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.200$2.50$0.050
$0.383$3.01$0.191
$0.400$2.40$0.100
$0.400$3.00$0.050
$0.420$3.00$0.085
$0.420$3.00$0.085
$0.440$4.07-
$0.450$3.20$0.250
$0.470$3.19-
$0.678$3.73$0.136
Groq logo
GroqBatch
$0.800$4.00-
$0.990$1.49-

Prices updated daily. Last check: Oct 2, 2026

Qwen3.8 27B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
33.7 / 100
Coding
68.1 / 100
Output Speed
45.5 t/s
Latency (TTFT)
1.2s

Reasoning & Knowledge

  • GPQA Diamond90.5%
  • Humanity's Last Exam33.9%

Coding

  • SciCode46.6%

Agentic & Tool Use

  • Terminal-Bench v2.179.8%
  • τ-bench Banking48.0%

Instruction & Long Context

  • Long-Context Reasoning82.0%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Alibaba
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • 27B parameter scale sits between small high-throughput models and very large frontier-class systems, a common cost/quality middle ground
  • Part of Alibaba's Qwen series, which has broad ecosystem support across inference frameworks and hosting providers
  • Model size is small enough to be served on modest GPU configurations, which tends to keep hosted inference pricing competitive
  • Chat-tuned rather than base, so it is usable for assistant workloads without additional instruction tuning
  • Multiple hosting options for Qwen-family models make it possible to compare providers on the same checkpoint
  • Mid-size models in this class typically deliver lower time-to-first-token than much larger alternatives

Limitations

  • Our database has no confirmed context window figure for this model, so long-document suitability must be verified with the provider
  • Input and output modalities are not confirmed in our metadata — do not assume image or audio support without checking
  • Throughput and time-to-first-token benchmarks from Artificial Analysis are not yet populated for this entry
  • At 27B parameters it will generally trail much larger Qwen and competitor models on hard reasoning, math, and complex agentic tasks
  • Serving details such as quantization level and maximum sequence length can differ between providers hosting the same weights

Key Features

•Chat/instruction-following model from the Alibaba Qwen series
•Approximately 27 billion parameters
•Mid-size tier positioned above small 4B–8B variants
•Available through hosted inference APIs tracked on this page
•Suitable for single-node or small multi-GPU serving configurations
•Provider-by-provider price comparison in the table above
•Catalog entry tracked against Artificial Analysis benchmark data as it becomes available

About Qwen3.8 27B

Qwen3.8 27B is a chat-oriented large language model attributed to Alibaba, the group behind the Qwen model series. The "27B" in the name indicates a parameter count of roughly 27 billion, placing it in the mid-size class of open-weight-style models — larger than the small 4B-to-8B tiers that target cheap, high-volume work, but well below the very large mixture-of-experts models that sit at the top of most vendor lineups. Models in this size band are commonly chosen when a single accelerator or a small multi-GPU node needs to serve general assistant workloads. Our catalog entry for Qwen3.8 27B currently carries limited verified technical detail. We track it as a chat model, and we do not have confirmed figures for its context window, supported input modalities, or tool-calling interface, so this page does not assert values for those fields. Throughput and time-to-first-token measurements sourced from Artificial Analysis are not yet populated for this entry either; where a provider publishes its own latency numbers, those will generally be the more reliable reference for a specific deployment. Readers evaluating the model for production should confirm context length, quantization, and API feature support directly with whichever host they plan to use, since these vary by serving stack even for an identical checkpoint. In practice, a 27B-class Qwen chat model is typically deployed for general conversational assistants, summarization, drafting, and light code assistance, where the tradeoff between response quality and serving cost matters more than absolute peak capability. Compared with much smaller Qwen variants it should handle multi-step instructions and longer prompts more reliably; compared with the largest models in the ecosystem it will generally be cheaper to serve and faster per token, at some cost in reasoning depth. The pricing table on this page shows which providers currently host it and what they charge.

Common Use Cases

Qwen3.8 27B is aimed at general-purpose chat workloads where a mid-size model gives enough quality for customer-facing assistants, internal knowledge tools, summarization pipelines, content drafting, and everyday code explanation, without the per-token cost of a much larger model. Its 27B scale is a common choice for teams running a self-hosted or dedicated endpoint, since the weights fit comfortably on a small GPU footprint and can be scaled horizontally for concurrency. For workloads that hinge on very long inputs, multimodal input, or structured tool orchestration, confirm those capabilities with the specific provider first — our metadata does not yet record them for this model. Extremely simple classification or routing tasks are usually better served by a smaller Qwen variant, while multi-step research agents and hard math or scientific reasoning are typically better matched to a larger or explicitly reasoning-tuned model.

Frequently Asked Questions

How much does Qwen3.8 27B cost to use?

Pricing depends on the provider and on the pricing model — per-million-token serverless billing, dedicated endpoints, and hourly GPU rental all price differently, and rates change frequently. Check the pricing table on this page for the current figures from each host we track.

What is Qwen3.8 27B best used for?

General chat assistants, summarization, drafting, and light coding help, where a mid-size model balances response quality against serving cost. It is a reasonable default when a small 4B–8B model is not accurate enough but a very large model is more than the workload requires.

What context window does Qwen3.8 27B support?

We do not have a confirmed context window for this entry in our database, and effective maximum sequence length can also be capped by the serving provider. Verify the supported context length with the specific API host before designing long-document workflows.

Does Qwen3.8 27B support image input or tool calling?

Our metadata records it as a chat model but does not confirm modality or tool-calling support either way. Check the provider's API documentation for the endpoint you intend to use rather than assuming.

How does it compare to smaller and larger Qwen models?

At roughly 27 billion parameters it should handle multi-step instructions and longer prompts more reliably than the small Qwen variants, while costing less to serve and typically responding faster than the largest models in the series. The trade-off shows up on the hardest reasoning and agentic tasks, where larger models generally lead.

Are throughput benchmarks available for this model?

Not yet. Our Artificial Analysis figures for output tokens per second and time to first token are unpopulated for this entry, so provider-published latency numbers or your own load testing are the better reference for now.