Qwen2 Instruct 72B is Alibaba's 72-billion-parameter instruction-tuned language model from the Qwen2 generation, positioned as the largest dense model in that release.
| Provider | Input / 1M | Output / 1M |
|---|---|---|
| $0.900 | $0.900 |
Prices updated daily. Last check: Sep 21, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Qwen2 Instruct 72B suits general-purpose assistant workloads where a large dense model is preferred over a small one: multi-turn chat, document summarization, content drafting, code generation and code review assistance, and question answering over supplied text. Its position as the top size in the Qwen2 generation makes it a candidate for tasks where the smaller Qwen2 variants produce weaker instruction-following, and its strength in Chinese makes it a common choice for bilingual Chinese-English products, localization pipelines, and customer-facing assistants serving CJK markets. Teams already standardized on the Qwen family may also use it as a fine-tuning base or as a self-hosted option, though the 72B dense footprint means multi-GPU serving. For latency-sensitive, high-volume classification or extraction, a smaller Qwen model is usually the more economical fit.
Pricing depends on which provider you use and the pricing model they offer — per-token API billing versus renting GPU capacity to host the weights yourself. Rates change frequently and differ between hosts serving identical weights, so check the pricing table on this page for current per-provider figures.
It is best suited to general-purpose assistant work: multi-turn chat, summarization, drafting, code generation, and question answering, with particular relevance for Chinese-English bilingual applications given the Qwen series' language coverage. For very high-volume, simple classification tasks, a smaller model in the family is usually more cost-effective.
It is the 72-billion-parameter size, the largest dense model in Alibaba's Qwen2 release. Larger parameter counts generally improve instruction-following and reasoning quality but increase memory requirements and per-token cost, so the smaller Qwen2 sizes remain preferable for latency-sensitive or high-throughput pipelines.
Alibaba has released later Qwen generations since Qwen2, so this model represents an earlier point in the family. It is still hosted by several providers and remains a reasonable choice for existing deployments or where its price and availability fit, but newer Qwen releases are worth evaluating for new projects.
We do not have a confirmed context window recorded for this model in our database, and effective context can differ between hosts serving the same weights. Check the documentation of the specific provider you plan to use for the exact supported context length.