GLM-5.1 is a large language model from Zhipu (Z.ai) in the GLM-5 family, offering a 200,000-token context window.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.966 | $3.04 | $0.179 | |
| $1.00 | $3.20 | - | |
| $1.05 | $3.50 | $0.205 | |
| $1.22 | $3.96 | $0.610 | |
| $1.30 | $4.30 | $0.260 | |
| $1.38 | $4.40 | $0.260 | |
| $1.40 | $4.40 | - | |
| $1.40 | $4.40 | - |
Prices updated daily. Last check: Sep 6, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
GLM-5.1 fits workloads that benefit from a wide context window: analyzing long contracts, reports, or research papers in one pass; reviewing multi-file code changes; summarizing long transcripts; and powering assistants that need to retain lengthy conversation history. Its position in Zhipu's GLM-5 family and the series' bilingual heritage make it a candidate for applications serving Chinese and English users, or for teams evaluating non-US model providers for cost or sourcing reasons. Because it is hosted by several inference providers, it also suits teams that want the option to move the same model between endpoints. For latency-sensitive or extremely high-volume classification work, benchmark it against smaller models first, since we do not track confirmed throughput figures for GLM-5.1.
Pricing depends on which provider hosts the model and on the pricing type — input versus output tokens, cached input, and any batch or committed-use discounts. Rates for the same model can differ meaningfully between endpoints. See the pricing table on this page for the current per-provider figures we track.
It suits long-context tasks such as document analysis, multi-file code review, long transcript summarization, and assistants that carry extended conversation history, thanks to its 200,000-token context window. Its GLM lineage also makes it a common choice for bilingual Chinese and English applications.
GLM-5.1 supports a 200,000-token context window. That is enough for several hundred pages of text or a substantial portion of a codebase in a single request, though it is smaller than the 1M-token windows some competing long-context models offer.
GLM-5.1 is part of the same GLM-5 family from Zhipu and represents an incremental revision within that line rather than a separate model series. If you are choosing between them, test both on your own prompts, since our dataset does not include complete comparative benchmarks.
Providers list the model under several identifiers, including glm-5.1, zai-org/GLM-5.1, and THUDM/GLM-5.1. These refer to the same model; THUDM and zai-org are organization names associated with Zhipu's published releases.