Skip to main content
Alibaba

Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is a large mixture-of-experts style model from Alibaba's Qwen family, measured at roughly 46.9 output tokens per second in third-party throughput testing.

Input from
$2.00 / 1M tokens
across 5 providers

API Pricing

Cheapest on Deep Infra — 5% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$2.00$6.00$0.200
$2.00$6.00-
$2.00$6.00$0.250
$2.00$6.00$0.250
$2.50$6.00$0.630

Prices updated daily. Last check: Sep 28, 2026

Qwen3.8 2.4T A95B pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
39.9 / 100
Coding
71.9 / 100
Output Speed
38.1 t/s
Latency (TTFT)
2.0s

Reasoning & Knowledge

  • GPQA Diamond93.5%
  • Humanity's Last Exam42.4%

Coding

  • SciCode54.1%

Agentic & Tool Use

  • Terminal-Bench v2.182.0%
  • τ-bench Banking49.1%

Instruction & Long Context

  • Long-Context Reasoning80.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Alibaba
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Part of Alibaba's Qwen model line, which is widely served across multiple inference providers, giving buyers options for comparison shopping
  • Name indicates a sparse mixture-of-experts design (2.4T total parameters, 95B active per token), a configuration aimed at large capacity without dense-model per-token compute
  • Independently measured output speed of roughly 46.9 tokens per second by Artificial Analysis, providing a third-party performance reference
  • Very large total parameter budget relative to its active-parameter count, the pattern typically used for broad knowledge coverage
  • Multiple hosted endpoints in the pricing table on this page allow direct cost and speed comparison for the same weights

Limitations

  • We do not have a confirmed context window length recorded for this model
  • Time to first token measured at roughly 1,790 ms, which is slow for latency-sensitive interactive or voice applications
  • Output speed of about 46.9 tokens per second is modest compared with smaller models optimised for throughput
  • Limited verified benchmark coverage in our database beyond speed and latency measurements — quality comparisons should be sourced elsewhere
  • Modality support, tool-calling behaviour, and reasoning modes are not tracked here and must be confirmed with the serving provider

Key Features

•Alibaba Qwen model family
•Sparse mixture-of-experts naming convention: 2.4T total parameters, 95B active per token
•Measured median output speed of ~46.9 tokens per second (Artificial Analysis)
•Measured time to first token of ~1,790 ms (Artificial Analysis)
•Available through hosted inference API endpoints listed in the pricing table on this page
•Per-provider pricing and throughput variation trackable across hosts

About Qwen3.8 2.4T A95B

Qwen3.8 2.4T A95B is a model released under Alibaba's Qwen line, the company's series of general-purpose language models. Its name follows the sparse mixture-of-experts naming convention used across recent Qwen releases, where the first figure refers to total parameter count (2.4T) and the "A" figure refers to the parameters active per token (95B). That structure is intended to give a model a very large total capacity while keeping the per-token compute cost closer to that of a mid-to-large dense model. Our database currently holds limited verified specifications for this model. The measurements we do track come from Artificial Analysis, which recorded a median output speed of about 46.9 tokens per second and a time to first token of roughly 1,790 milliseconds. Those figures place it in the range typical of large, high-capacity models rather than small latency-optimised ones: responses stream steadily, but the initial delay before the first token arrives is close to two seconds. Throughput and latency both vary by serving provider, hardware, and load, so the numbers above are best treated as one reference point rather than a guarantee. We do not currently track a confirmed context window, modality list, or reasoning-mode configuration for Qwen3.8 2.4T A95B, so buyers evaluating it for a specific workload should confirm those details against the serving provider's own documentation. Compare the providers listed in the pricing table on this page, since serving cost and measured speed for the same model can differ substantially between hosts.

Common Use Cases

Based on its scale and measured performance profile, Qwen3.8 2.4T A95B is oriented toward workloads where answer quality and breadth of knowledge matter more than response latency: long-form drafting and rewriting, document analysis and summarisation, code generation and review in batch pipelines, and multi-step agent or workflow steps that run asynchronously. The roughly 1.8-second time to first token makes it a weaker fit for real-time chat UIs, autocomplete, or voice interfaces where perceived responsiveness dominates. For high-volume, cost-sensitive tasks such as classification, routing, or short extraction jobs, a smaller Qwen model will usually be the more economical choice, with this model reserved for the harder requests in a tiered routing setup.

Frequently Asked Questions

How much does Qwen3.8 2.4T A95B cost to use?

Pricing depends on which provider serves the model and on the pricing type — input tokens, output tokens, cached input, and batch rates are typically billed differently. Rates also change frequently. See the pricing table on this page for current per-provider figures.

What is Qwen3.8 2.4T A95B best used for?

It suits tasks where output quality outweighs latency: long-form generation, document analysis and summarisation, code generation and review, and asynchronous agent steps. Its roughly 1.8-second time to first token makes it less suitable for real-time chat, autocomplete, or voice applications.

What do the numbers in the name mean?

The naming follows the sparse mixture-of-experts convention used in recent Qwen releases: 2.4T refers to total parameter count and A95B to the parameters activated per token. The intent of that design is to combine a large total capacity with per-token compute closer to a mid-to-large dense model.

How fast is Qwen3.8 2.4T A95B?

Artificial Analysis measured a median output speed of about 46.9 tokens per second and a time to first token of roughly 1,790 milliseconds. Both figures vary by provider, hardware, and load, so treat them as one reference point and compare the providers listed on this page.

What context window does it support?

We do not have a confirmed context window recorded for Qwen3.8 2.4T A95B. Check the documentation of the specific provider you plan to use, since hosted endpoints for the same model sometimes cap the context length differently.