Qwen3 Coder 480B A35B Instruct
Qwen3 Coder 480B A35B Instruct is a large mixture-of-experts chat model from Alibaba's Qwen team, oriented toward code generation and agentic software tasks.
API Pricing
Cheapest on Deep Infra — 66% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.300 | $1.00 | $0.100 | |
| $0.380 | $1.55 | - | |
| $0.900 | $4.50 | $0.180 | |
| $2.00 | $2.00 | - |
Prices updated daily. Last check: Sep 30, 2026
Qwen3 Coder 480B A35B Instruct pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro78.8%
- GPQA Diamond61.8%
- Humanity's Last Exam4.5%
Coding
- LiveCodeBench58.5%
Math
- AIME 202539.3%
- AIME47.7%
- MATH-50094.2%
Agentic & Tool Use
- Terminal-Bench Hard18.9%
- τ²-bench43.6%
Instruction & Long Context
- IFBench40.5%
- Long-Context Reasoning45.7%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Alibaba
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Mixture-of-experts design activates roughly 35B parameters out of ~480B total, keeping serving cost lower than a dense model of comparable size
- Purpose-tuned for code generation, refactoring, and debugging rather than repurposed from a general chat checkpoint
- Instruct-tuned variant, so it follows task instructions directly without needing few-shot base-model prompting
- Part of the broader Qwen3 family, making it straightforward to swap between coder and general-purpose siblings in the same stack
- Served by multiple independent inference providers, allowing price and latency comparison across endpoints
- Suited to agentic coding loops where the model iterates over edits and tool results across many turns
Limitations
- Our database does not track a confirmed context window for this entry — verify the limit with the specific provider you use
- Code specialization means general chat, creative writing, and open-domain reasoning may be better served by general-purpose Qwen3 variants
- Large MoE models require substantial memory to self-host, so local deployment is impractical on single-GPU setups
- Throughput and time-to-first-token measurements are not currently populated for this model in our data
- Performance characteristics can vary between providers serving the same weights, depending on quantization and serving configuration
Key Features
About Qwen3 Coder 480B A35B Instruct
Common Use Cases
This model targets programming workloads: writing new functions and modules, translating code between languages, explaining or reviewing unfamiliar repositories, generating tests, and diagnosing failures from stack traces. Its instruct tuning and agentic orientation make it a candidate for coding agents and IDE assistants that run many sequential tool calls against a codebase, where the model must interpret file contents and command output rather than answer a single isolated question. The MoE architecture matters most for high-volume use — teams running continuous autocomplete, CI-integrated review, or batch refactoring jobs get a large-model capability profile at an activated-parameter cost closer to a mid-size dense model. For general assistant duties, summarization, or open-domain conversation, a general-purpose Qwen3 Instruct variant is usually the more appropriate pick.
Frequently Asked Questions
What does "480B A35B" mean in the model name?
It describes the mixture-of-experts configuration: approximately 480 billion total parameters across all experts, with roughly 35 billion parameters activated for any given token. The total count reflects the memory needed to hold the model; the active count is closer to what determines per-token compute.
What is Qwen3 Coder 480B A35B Instruct best used for?
Code-centric work — generating and refactoring code, debugging, writing tests, reviewing pull requests, and powering agentic coding tools that iterate over a repository with tool calls. It is a code-specialized branch of the Qwen3 family rather than a general-purpose chat model.
How much does it cost to use?
Pricing varies by provider and by pricing type, and input and output tokens are usually billed at different rates. Because multiple providers serve Qwen models, rates differ across endpoints. See the pricing table on this page for current figures.
How does it differ from a general Qwen3 Instruct model?
The Coder branch is tuned specifically for programming and agentic software tasks, while general Qwen3 Instruct variants cover broader conversational and analytical use. If most of your prompts involve source code, tests, or tool-driven repository work, the Coder variant is the closer fit; for mixed general assistant traffic, the general variant is usually more appropriate.
What context window does it support?
We do not have a confirmed context window recorded for this entry, and hosted context limits can differ between providers serving the same weights. Check the documentation of the specific inference provider you plan to use before designing long-context workflows around it.