Qwen3.5 397B A17B is a large mixture-of-experts language model from Alibaba's Qwen series, with roughly 397B total parameters and about 17B active per token.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.450 | $3.00 | $0.220 | |
| $0.450 | $3.00 | - | |
| $0.550 | $3.50 | $0.225 | |
| $0.600 | $3.60 | - | |
| $0.600 | $3.60 | - | |
| $0.600 | $3.60 | - | |
| $0.688 | $4.13 | - | |
| $0.830 | $5.55 | - |
Prices updated daily. Last check: Sep 22, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Qwen3.5 397B A17B suits workloads that benefit from a high-capacity model but where the compute cost of a dense model at similar scale would be prohibitive: general-purpose assistants, code generation and refactoring, document analysis, and multi-step reasoning tasks. Its sparse activation profile makes it a reasonable fit for batch and throughput-oriented pipelines, where the roughly 95 tokens per second measured output rate matters more than the ~1.4 second time to first token. For highly interactive, low-latency chat interfaces, a smaller Qwen variant may be a better trade. Teams already standardized on the Qwen ecosystem — or those needing strong Chinese-language handling alongside English — often evaluate this size class against comparable mixture-of-experts models from other labs on both cost per million tokens and measured throughput.
Pricing depends on which inference provider you use and how the endpoint is billed — per-input-token, per-output-token, and any serverless versus dedicated arrangement all differ between hosts. Because Qwen models are typically served by several independent providers, rates can vary substantially for the same model. See the pricing table on this page for the current figures from the providers we track.
It describes a mixture-of-experts model: roughly 397 billion total parameters across all experts, with approximately 17 billion parameters activated for any single token. The practical effect is that inference compute per token is closer to a 17B-class model, while total memory requirements scale with the full 397B parameter count.
It is generally aimed at general-purpose assistant work, coding, document analysis, and multi-step reasoning where a large-capacity model is useful. Its throughput profile — around 95 output tokens per second with a time to first token near 1,375 ms in Artificial Analysis testing — leans toward batch and generation-heavy workloads rather than very low-latency interactive use.
We do not have a confirmed context window for this specific variant in our database. Because context limits for Qwen models often differ between the model's native capability and what individual hosts expose, check the documentation of the provider you plan to use.
Within the Qwen3.5 generation, this variant sits at a much larger total parameter count than the compact models, which typically translates into more capacity for knowledge-heavy and reasoning tasks. Smaller siblings generally offer lower latency and lower per-token cost, so the trade-off is capacity versus responsiveness and price. Compare the specific endpoints in the pricing table to see how each option is served.