Qwen3 Coder Next
Qwen3 Coder Next is a coding-oriented chat model from Alibaba's Qwen3 Coder line, measured at roughly 130 output tokens per second with sub-second time to first token.
API Pricing
Cheapest on OpenRouter — 60% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.120 | $0.800 | $0.070 | |
| $0.180 | $1.23 | - | |
| $0.200 | $1.50 | - | |
| $0.300 | $1.50 | - | |
| $0.500 | $1.20 | - | |
| $0.500 | $1.20 | - |
Prices updated daily. Last check: Sep 30, 2026
Qwen3 Coder Next pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond73.7%
- Humanity's Last Exam10.1%
Coding
- SciCode36.2%
Agentic & Tool Use
- Terminal-Bench Hard18.2%
- Terminal-Bench v2.138.2%
- τ²-bench79.5%
- τ-bench Banking5.4%
Instruction & Long Context
- IFBench35.2%
- Long-Context Reasoning47.0%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Alibaba
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Purpose-built for software development tasks as part of Alibaba's Qwen3 Coder line, rather than a general chat model adapted to code
- Measured output throughput of roughly 129.6 tokens per second, which suits long code generation and file-rewrite tasks
- Time to first token around 926 ms, keeping interactive coding chat responsive
- Independent third-party benchmarking from Artificial Analysis rather than creator-reported figures only
- Part of the widely deployed Qwen3 family, so it is typically available from multiple inference providers, allowing price and latency comparison
- Distinct 'Next' release within the Coder line, giving teams an alternative checkpoint to test against earlier Qwen3 Coder models
Limitations
- Context window is not tracked in our database — confirm the maximum supported length with your chosen provider
- No coding benchmark scores (such as SWE-bench or HumanEval) are recorded in our data for this model
- Coder-tuned models are generally optimized for programming tasks and may be a weaker fit for open-ended creative or conversational workloads
- Serving characteristics vary significantly between providers, so the measured throughput and latency may not match what you observe in production
- Licensing and weight-availability terms are not tracked here and should be verified directly with Alibaba or the hosting provider
Key Features
About Qwen3 Coder Next
Common Use Cases
Qwen3 Coder Next is aimed at developer-facing workloads: generating new functions and modules, refactoring existing files, writing tests, translating between languages and frameworks, and explaining unfamiliar code. Its sub-second time to first token makes it workable for interactive assistants embedded in an IDE or terminal, where the perceived delay before the first character appears dominates user experience, while its roughly 130 tokens per second sustained rate helps on longer outputs such as full-file rewrites or multi-step agent trajectories that emit large diffs. Teams running high-volume automated code review, docstring generation, or batch migration jobs may also find the throughput profile relevant, since those pipelines are usually bound by decode speed rather than prompt processing. Before committing to it for agentic coding that requires very long repository context, verify the context window and any tool-calling support with the specific provider you plan to use, as our metadata does not cover those attributes.
Frequently Asked Questions
How much does Qwen3 Coder Next cost to run?
Pricing depends on which provider you use and whether you are billing per token through a serverless API, renting dedicated capacity, or self-hosting on GPUs. Rates also differ between input and output tokens. Check the pricing table on this page for the current per-provider figures rather than relying on any single quoted number.
What is Qwen3 Coder Next best used for?
It is a coding-oriented chat model, so it fits code generation, completion, refactoring, test writing, code explanation, and developer assistant workflows. The combination of about 926 ms to first token and roughly 130 output tokens per second suits both interactive IDE-style use and longer batch code generation.
How fast is Qwen3 Coder Next?
Artificial Analysis measured it at approximately 129.6 output tokens per second with a time to first token of about 926 milliseconds. Real-world figures vary with the provider, hardware, quantization, prompt length, and current load, so use these as a comparison baseline rather than a guaranteed service level.
How does it differ from general Qwen3 chat models?
Qwen3 Coder Next sits in the Coder branch of the Qwen3 family, which is tuned specifically for programming tasks. General-purpose Qwen3 chat models spread capability across conversation, writing, and reasoning. If your workload is predominantly code, the Coder variant is the more targeted choice; for mixed general assistant duties, a general Qwen3 model may be a better match.
What context window does Qwen3 Coder Next support?
We do not currently track a confirmed context window for this model. Hosted deployments of Qwen models often expose different maximum context lengths, so check the documentation of the specific provider you select in the pricing table above.