Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is a large mixture-of-experts style model from Alibaba's Qwen family, measured at roughly 46.9 output tokens per second in third-party throughput testing.
API Pricing
Cheapest on Deep Infra — 5% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $2.00 | $6.00 | $0.200 | |
| $2.00 | $6.00 | - | |
| $2.00 | $6.00 | $0.250 | |
| $2.00 | $6.00 | $0.250 | |
| $2.50 | $6.00 | $0.630 |
Prices updated daily. Last check: Sep 28, 2026
Qwen3.8 2.4T A95B pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond93.5%
- Humanity's Last Exam42.4%
Coding
- SciCode54.1%
Agentic & Tool Use
- Terminal-Bench v2.182.0%
- τ-bench Banking49.1%
Instruction & Long Context
- Long-Context Reasoning80.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Alibaba
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Part of Alibaba's Qwen model line, which is widely served across multiple inference providers, giving buyers options for comparison shopping
- Name indicates a sparse mixture-of-experts design (2.4T total parameters, 95B active per token), a configuration aimed at large capacity without dense-model per-token compute
- Independently measured output speed of roughly 46.9 tokens per second by Artificial Analysis, providing a third-party performance reference
- Very large total parameter budget relative to its active-parameter count, the pattern typically used for broad knowledge coverage
- Multiple hosted endpoints in the pricing table on this page allow direct cost and speed comparison for the same weights
Limitations
- We do not have a confirmed context window length recorded for this model
- Time to first token measured at roughly 1,790 ms, which is slow for latency-sensitive interactive or voice applications
- Output speed of about 46.9 tokens per second is modest compared with smaller models optimised for throughput
- Limited verified benchmark coverage in our database beyond speed and latency measurements — quality comparisons should be sourced elsewhere
- Modality support, tool-calling behaviour, and reasoning modes are not tracked here and must be confirmed with the serving provider
Key Features
About Qwen3.8 2.4T A95B
Common Use Cases
Based on its scale and measured performance profile, Qwen3.8 2.4T A95B is oriented toward workloads where answer quality and breadth of knowledge matter more than response latency: long-form drafting and rewriting, document analysis and summarisation, code generation and review in batch pipelines, and multi-step agent or workflow steps that run asynchronously. The roughly 1.8-second time to first token makes it a weaker fit for real-time chat UIs, autocomplete, or voice interfaces where perceived responsiveness dominates. For high-volume, cost-sensitive tasks such as classification, routing, or short extraction jobs, a smaller Qwen model will usually be the more economical choice, with this model reserved for the harder requests in a tiered routing setup.
Frequently Asked Questions
How much does Qwen3.8 2.4T A95B cost to use?
Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, cached input, and batch rates are typically billed differently. Rates also change frequently. See the pricing table on this page for current per-provider figures.
What is Qwen3.8 2.4T A95B best used for?
It suits tasks where output quality outweighs latency: long-form generation, document analysis and summarisation, code generation and review, and asynchronous agent steps. Its roughly 1.8-second time to first token makes it less suitable for real-time chat, autocomplete, or voice applications.
What do the numbers in the name mean?
The naming follows the sparse mixture-of-experts convention used in recent Qwen releases: 2.4T refers to total parameter count and A95B to the parameters activated per token. The intent of that design is to combine a large total capacity with per-token compute closer to a mid-to-large dense model.
How fast is Qwen3.8 2.4T A95B?
Artificial Analysis measured a median output speed of about 46.9 tokens per second and a time to first token of roughly 1,790 milliseconds. Both figures vary by provider, hardware, and load, so treat them as one reference point and compare the providers listed on this page.
What context window does it support?
We do not have a confirmed context window recorded for Qwen3.8 2.4T A95B. Check the documentation of the specific provider you plan to use, since hosted endpoints for the same model sometimes cap the context length differently.