Skip to main content
Alibaba

Qwen3 Coder 480B A35B Instruct

Qwen3 Coder 480B A35B Instruct is a large mixture-of-experts chat model from Alibaba's Qwen team, oriented toward code generation and agentic software tasks.

Input from
$0.300 / 1M tokens
across 4 providers

API Pricing

Cheapest on Deep Infra — 66% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.300$1.00$0.100
$0.380$1.55-
$0.900$4.50$0.180
$2.00$2.00-

Prices updated daily. Last check: Sep 30, 2026

Qwen3 Coder 480B A35B Instruct pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
11.9 / 100
Math
39.3 / 100

Reasoning & Knowledge

  • MMLU-Pro78.8%
  • GPQA Diamond61.8%
  • Humanity's Last Exam4.5%

Coding

  • LiveCodeBench58.5%

Math

  • AIME 202539.3%
  • AIME47.7%
  • MATH-50094.2%

Agentic & Tool Use

  • Terminal-Bench Hard18.9%
  • τ²-bench43.6%

Instruction & Long Context

  • IFBench40.5%
  • Long-Context Reasoning45.7%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Alibaba
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Mixture-of-experts design activates roughly 35B parameters out of ~480B total, keeping serving cost lower than a dense model of comparable size
  • Purpose-tuned for code generation, refactoring, and debugging rather than repurposed from a general chat checkpoint
  • Instruct-tuned variant, so it follows task instructions directly without needing few-shot base-model prompting
  • Part of the broader Qwen3 family, making it straightforward to swap between coder and general-purpose siblings in the same stack
  • Served by multiple independent inference providers, allowing price and latency comparison across endpoints
  • Suited to agentic coding loops where the model iterates over edits and tool results across many turns

Limitations

  • Our database does not track a confirmed context window for this entry — verify the limit with the specific provider you use
  • Code specialization means general chat, creative writing, and open-domain reasoning may be better served by general-purpose Qwen3 variants
  • Large MoE models require substantial memory to self-host, so local deployment is impractical on single-GPU setups
  • Throughput and time-to-first-token measurements are not currently populated for this model in our data
  • Performance characteristics can vary between providers serving the same weights, depending on quantization and serving configuration

Key Features

•Mixture-of-experts architecture (~480B total parameters, ~35B active)
•Code-specialized instruct tuning within the Qwen3 family
•Chat-completions style API interface at hosting providers
•Multi-turn agentic coding workflows (edit, run, iterate)
•Available from multiple independent inference providers
•Instruct-following behavior without base-model prompt engineering

About Qwen3 Coder 480B A35B Instruct

Qwen3 Coder 480B A35B Instruct is a code-focused member of Alibaba's Qwen3 model family. As its name indicates, it is a mixture-of-experts (MoE) design with roughly 480 billion total parameters and about 35 billion active per token, and it is an instruct-tuned variant rather than a base checkpoint. Within the Qwen3 lineup, the Coder branch is the specialization aimed at programming work, while general Qwen3 Instruct and reasoning variants cover broader chat and analysis tasks. The MoE architecture is the defining technical characteristic here: only a fraction of the total parameter count is activated for any given token, which is what allows a model of this scale to be served at inference costs closer to a much smaller dense model. It is exposed as a chat-completions style model by the providers that host it, and is typically used for code writing, refactoring, debugging, and multi-step agentic coding loops where the model calls tools, reads files, and iterates. We do not track a confirmed context window, modality list, or tool-calling spec for this entry in our database — check the serving provider's documentation for the exact context length and feature support they expose, since these can differ between hosts of the same weights. In practice, this model is generally evaluated against other large code-specialized models and against general-purpose chat models used for programming. Its appeal is the combination of a code-tuned instruct model with MoE efficiency, and the fact that Qwen models are served by multiple independent inference providers, which gives buyers a choice of endpoints rather than a single vendor. Latency and throughput vary meaningfully between those hosts, so benchmark numbers from one provider should not be assumed to hold for another.

Common Use Cases

This model targets programming workloads: writing new functions and modules, translating code between languages, explaining or reviewing unfamiliar repositories, generating tests, and diagnosing failures from stack traces. Its instruct tuning and agentic orientation make it a candidate for coding agents and IDE assistants that run many sequential tool calls against a codebase, where the model must interpret file contents and command output rather than answer a single isolated question. The MoE architecture matters most for high-volume use — teams running continuous autocomplete, CI-integrated review, or batch refactoring jobs get a large-model capability profile at an activated-parameter cost closer to a mid-size dense model. For general assistant duties, summarization, or open-domain conversation, a general-purpose Qwen3 Instruct variant is usually the more appropriate pick.

Frequently Asked Questions

What does "480B A35B" mean in the model name?

It describes the mixture-of-experts configuration: approximately 480 billion total parameters across all experts, with roughly 35 billion parameters activated for any given token. The total count reflects the memory needed to hold the model; the active count is closer to what determines per-token compute.

What is Qwen3 Coder 480B A35B Instruct best used for?

Code-centric work — generating and refactoring code, debugging, writing tests, reviewing pull requests, and powering agentic coding tools that iterate over a repository with tool calls. It is a code-specialized branch of the Qwen3 family rather than a general-purpose chat model.

How much does it cost to use?

Pricing varies by provider and by pricing type, and input and output tokens are usually billed at different rates. Because multiple providers serve Qwen models, rates differ across endpoints. See the pricing table on this page for current figures.

How does it differ from a general Qwen3 Instruct model?

The Coder branch is tuned specifically for programming and agentic software tasks, while general Qwen3 Instruct variants cover broader conversational and analytical use. If most of your prompts involve source code, tests, or tool-driven repository work, the Coder variant is the closer fit; for mixed general assistant traffic, the general variant is usually more appropriate.

What context window does it support?

We do not have a confirmed context window recorded for this entry, and hosted context limits can differ between providers serving the same weights. Check the documentation of the specific inference provider you plan to use before designing long-context workflows around it.