Qwen3.8 27B
Qwen3.8 27B is a 27-billion-parameter chat model from Alibaba's Qwen series, listed here so its hosted inference pricing can be compared across providers.
API Pricing
Cheapest on Deep Infra — 60% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.200 | $2.50 | $0.050 | |
| $0.383 | $3.01 | $0.191 | |
| $0.400 | $2.40 | $0.100 | |
| $0.400 | $3.00 | $0.050 | |
| $0.420 | $3.00 | $0.085 | |
| $0.420 | $3.00 | $0.085 | |
| $0.440 | $4.07 | - | |
| $0.450 | $3.20 | $0.250 | |
| $0.470 | $3.19 | - | |
| $0.678 | $3.73 | $0.136 | |
| $0.800 | $4.00 | - | |
| $0.990 | $1.49 | - |
Prices updated daily. Last check: Oct 2, 2026
Qwen3.8 27B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond90.5%
- Humanity's Last Exam33.9%
Coding
- SciCode46.6%
Agentic & Tool Use
- Terminal-Bench v2.179.8%
- τ-bench Banking48.0%
Instruction & Long Context
- Long-Context Reasoning82.0%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Alibaba
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- 27B parameter scale sits between small high-throughput models and very large frontier-class systems, a common cost/quality middle ground
- Part of Alibaba's Qwen series, which has broad ecosystem support across inference frameworks and hosting providers
- Model size is small enough to be served on modest GPU configurations, which tends to keep hosted inference pricing competitive
- Chat-tuned rather than base, so it is usable for assistant workloads without additional instruction tuning
- Multiple hosting options for Qwen-family models make it possible to compare providers on the same checkpoint
- Mid-size models in this class typically deliver lower time-to-first-token than much larger alternatives
Limitations
- Our database has no confirmed context window figure for this model, so long-document suitability must be verified with the provider
- Input and output modalities are not confirmed in our metadata — do not assume image or audio support without checking
- Throughput and time-to-first-token benchmarks from Artificial Analysis are not yet populated for this entry
- At 27B parameters it will generally trail much larger Qwen and competitor models on hard reasoning, math, and complex agentic tasks
- Serving details such as quantization level and maximum sequence length can differ between providers hosting the same weights
Key Features
About Qwen3.8 27B
Common Use Cases
Qwen3.8 27B is aimed at general-purpose chat workloads where a mid-size model gives enough quality for customer-facing assistants, internal knowledge tools, summarization pipelines, content drafting, and everyday code explanation, without the per-token cost of a much larger model. Its 27B scale is a common choice for teams running a self-hosted or dedicated endpoint, since the weights fit comfortably on a small GPU footprint and can be scaled horizontally for concurrency. For workloads that hinge on very long inputs, multimodal input, or structured tool orchestration, confirm those capabilities with the specific provider first — our metadata does not yet record them for this model. Extremely simple classification or routing tasks are usually better served by a smaller Qwen variant, while multi-step research agents and hard math or scientific reasoning are typically better matched to a larger or explicitly reasoning-tuned model.
Frequently Asked Questions
How much does Qwen3.8 27B cost to use?
Pricing depends on the provider and on the pricing model — per-million-token serverless billing, dedicated endpoints, and hourly GPU rental all price differently, and rates change frequently. Check the pricing table on this page for the current figures from each host we track.
What is Qwen3.8 27B best used for?
General chat assistants, summarization, drafting, and light coding help, where a mid-size model balances response quality against serving cost. It is a reasonable default when a small 4B–8B model is not accurate enough but a very large model is more than the workload requires.
What context window does Qwen3.8 27B support?
We do not have a confirmed context window for this entry in our database, and effective maximum sequence length can also be capped by the serving provider. Verify the supported context length with the specific API host before designing long-document workflows.
Does Qwen3.8 27B support image input or tool calling?
Our metadata records it as a chat model but does not confirm modality or tool-calling support either way. Check the provider's API documentation for the endpoint you intend to use rather than assuming.
How does it compare to smaller and larger Qwen models?
At roughly 27 billion parameters it should handle multi-step instructions and longer prompts more reliably than the small Qwen variants, while costing less to serve and typically responding faster than the largest models in the series. The trade-off shows up on the hardest reasoning and agentic tasks, where larger models generally lead.
Are throughput benchmarks available for this model?
Not yet. Our Artificial Analysis figures for output tokens per second and time to first token are unpopulated for this entry, so provider-published latency numbers or your own load testing are the better reference for now.