Skip to main content
Alibaba

Qwen3 Coder Next

Qwen3 Coder Next is a coding-oriented chat model from Alibaba's Qwen3 Coder line, measured at roughly 130 output tokens per second with sub-second time to first token.

Input from
$0.120 / 1M tokens
across 6 providers

API Pricing

Cheapest on OpenRouter — 60% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.120$0.800$0.070
$0.180$1.23-
$0.200$1.50-
$0.300$1.50-
$0.500$1.20-
$0.500$1.20-

Prices updated daily. Last check: Sep 30, 2026

Qwen3 Coder Next pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
9.2 / 100
Coding
36.2 / 100
Output Speed
130 t/s
Latency (TTFT)
926ms

Reasoning & Knowledge

  • GPQA Diamond73.7%
  • Humanity's Last Exam10.1%

Coding

  • SciCode36.2%

Agentic & Tool Use

  • Terminal-Bench Hard18.2%
  • Terminal-Bench v2.138.2%
  • τ²-bench79.5%
  • τ-bench Banking5.4%

Instruction & Long Context

  • IFBench35.2%
  • Long-Context Reasoning47.0%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Alibaba
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Purpose-built for software development tasks as part of Alibaba's Qwen3 Coder line, rather than a general chat model adapted to code
  • Measured output throughput of roughly 129.6 tokens per second, which suits long code generation and file-rewrite tasks
  • Time to first token around 926 ms, keeping interactive coding chat responsive
  • Independent third-party benchmarking from Artificial Analysis rather than creator-reported figures only
  • Part of the widely deployed Qwen3 family, so it is typically available from multiple inference providers, allowing price and latency comparison
  • Distinct 'Next' release within the Coder line, giving teams an alternative checkpoint to test against earlier Qwen3 Coder models

Limitations

  • Context window is not tracked in our database — confirm the maximum supported length with your chosen provider
  • No coding benchmark scores (such as SWE-bench or HumanEval) are recorded in our data for this model
  • Coder-tuned models are generally optimized for programming tasks and may be a weaker fit for open-ended creative or conversational workloads
  • Serving characteristics vary significantly between providers, so the measured throughput and latency may not match what you observe in production
  • Licensing and weight-availability terms are not tracked here and should be verified directly with Alibaba or the hosting provider

Key Features

•Coding-focused chat model in the Qwen3 Coder line
•Measured output speed of approximately 129.6 tokens per second
•Time to first token of approximately 926 ms
•Code generation, completion, refactoring, and explanation workloads
•Benchmarked independently by Artificial Analysis
•Available through multiple third-party inference providers listed in the pricing table
•Part of Alibaba's Qwen3 model family

About Qwen3 Coder Next

Qwen3 Coder Next is a chat model published by Alibaba as part of the Qwen3 Coder branch of the broader Qwen3 family. Where general-purpose Qwen3 chat models target a mix of conversation, reasoning, and writing, the Coder variants are positioned around software development work — code generation, completion, refactoring, and explanation. The "Next" designation marks it as a distinct release within that Coder line rather than a size variant, so it should be evaluated against sibling Qwen3 Coder checkpoints as well as against coding models from other creators. On throughput, independent measurement from Artificial Analysis puts Qwen3 Coder Next at about 129.6 output tokens per second with a time to first token of roughly 926 milliseconds. That combination — a fast steady-state generation rate paired with a just-under-one-second first-token latency — matters for coding workflows, where long file rewrites and multi-step agent turns are dominated by sustained decode speed, while inline completion and chat-style Q&A are more sensitive to the initial delay. Actual numbers will vary by serving provider, hardware, quantization, and prompt length, so treat these figures as a reference point rather than a guarantee. Other specifications for this model — including context window, licensing terms, parameter count, and multimodal input support — are not tracked in our database, and we do not assert them here. Readers evaluating Qwen3 Coder Next for a specific integration should confirm those details with the provider they intend to use, since hosted deployments of Qwen models frequently differ in the maximum context and feature set they expose.

Common Use Cases

Qwen3 Coder Next is aimed at developer-facing workloads: generating new functions and modules, refactoring existing files, writing tests, translating between languages and frameworks, and explaining unfamiliar code. Its sub-second time to first token makes it workable for interactive assistants embedded in an IDE or terminal, where the perceived delay before the first character appears dominates user experience, while its roughly 130 tokens per second sustained rate helps on longer outputs such as full-file rewrites or multi-step agent trajectories that emit large diffs. Teams running high-volume automated code review, docstring generation, or batch migration jobs may also find the throughput profile relevant, since those pipelines are usually bound by decode speed rather than prompt processing. Before committing to it for agentic coding that requires very long repository context, verify the context window and any tool-calling support with the specific provider you plan to use, as our metadata does not cover those attributes.

Frequently Asked Questions

How much does Qwen3 Coder Next cost to run?

Pricing depends on which provider you use and whether you are billing per token through a serverless API, renting dedicated capacity, or self-hosting on GPUs. Rates also differ between input and output tokens. Check the pricing table on this page for the current per-provider figures rather than relying on any single quoted number.

What is Qwen3 Coder Next best used for?

It is a coding-oriented chat model, so it fits code generation, completion, refactoring, test writing, code explanation, and developer assistant workflows. The combination of about 926 ms to first token and roughly 130 output tokens per second suits both interactive IDE-style use and longer batch code generation.

How fast is Qwen3 Coder Next?

Artificial Analysis measured it at approximately 129.6 output tokens per second with a time to first token of about 926 milliseconds. Real-world figures vary with the provider, hardware, quantization, prompt length, and current load, so use these as a comparison baseline rather than a guaranteed service level.

How does it differ from general Qwen3 chat models?

Qwen3 Coder Next sits in the Coder branch of the Qwen3 family, which is tuned specifically for programming tasks. General-purpose Qwen3 chat models spread capability across conversation, writing, and reasoning. If your workload is predominantly code, the Coder variant is the more targeted choice; for mixed general assistant duties, a general Qwen3 model may be a better match.

What context window does Qwen3 Coder Next support?

We do not currently track a confirmed context window for this model. Hosted deployments of Qwen models often expose different maximum context lengths, so check the documentation of the specific provider you select in the pricing table above.