GLM-5.1
GLM-5.1 is a large language model from Zhipu (Z.ai) in the GLM-5 family, offering a 200,000-token context window.
API Pricing
Cheapest on OpenRouter — 27% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.965 | $3.03 | $0.179 | |
| $1.05 | $3.50 | $0.205 | |
| $1.12 | $3.52 | $0.208 | |
| $1.36 | $4.27 | $0.679 | |
| $1.38 | $4.40 | $0.260 | |
| $1.40 | $4.40 | $0.350 | |
| $1.40 | $4.40 | $0.260 | |
| $1.40 | $4.40 | - | |
| $1.81 | $5.70 | - |
Prices updated daily. Last check: Sep 25, 2026
GLM-5.1 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond83.9%
- Humanity's Last Exam27.9%
Agentic & Tool Use
- Terminal-Bench Hard35.6%
- τ²-bench97.1%
Instruction & Long Context
- IFBench52.0%
- Long-Context Reasoning53.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Zhipu
- Family
- GLM-5
- Context Window
- 200K
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
- Aliases
- glm-5.1, GLM-5.1, zai-org/GLM-5.1, THUDM/GLM-5.1
Strengths & Limitations
Strengths
- 200,000-token context window supports long documents, large diffs, and extended conversation histories in a single request
- Part of Zhipu's GLM-5 family, an incremental revision over the initial GLM-5 release
- Published under multiple aliases (zai-org/GLM-5.1, THUDM/GLM-5.1), making it identifiable across different inference catalogs
- Available through more than one inference endpoint, allowing price and latency comparison between hosts
- GLM series models are commonly used for both Chinese and English workloads, which suits bilingual applications
- Large context reduces reliance on complex chunking or retrieval pipelines for long-input tasks
Limitations
- We do not have confirmed throughput or time-to-first-token measurements for GLM-5.1, so serving speed must be validated per provider
- Context window of 200K tokens is smaller than the 1M-token windows some competing long-context models advertise
- Modality support and tool-calling details are not tracked in our dataset and vary by host
- Provider coverage for GLM models is narrower than for the largest Western model families, which can limit regional availability and failover options
- Benchmark results in our dataset are incomplete, making direct quality comparison with peers harder
Key Features
About GLM-5.1
Common Use Cases
GLM-5.1 fits workloads that benefit from a wide context window: analyzing long contracts, reports, or research papers in one pass; reviewing multi-file code changes; summarizing long transcripts; and powering assistants that need to retain lengthy conversation history. Its position in Zhipu's GLM-5 family and the series' bilingual heritage make it a candidate for applications serving Chinese and English users, or for teams evaluating non-US model providers for cost or sourcing reasons. Because it is hosted by several inference providers, it also suits teams that want the option to move the same model between endpoints. For latency-sensitive or extremely high-volume classification work, benchmark it against smaller models first, since we do not track confirmed throughput figures for GLM-5.1.
Frequently Asked Questions
How much does GLM-5.1 cost to use?
Pricing depends on which provider hosts the model and on the pricing type — input versus output tokens, cached input, and any batch or committed-use discounts. Rates for the same model can differ meaningfully between endpoints. See the pricing table on this page for the current per-provider figures we track.
What is GLM-5.1 best used for?
It suits long-context tasks such as document analysis, multi-file code review, long transcript summarization, and assistants that carry extended conversation history, thanks to its 200,000-token context window. Its GLM lineage also makes it a common choice for bilingual Chinese and English applications.
How large is the context window?
GLM-5.1 supports a 200,000-token context window. That is enough for several hundred pages of text or a substantial portion of a codebase in a single request, though it is smaller than the 1M-token windows some competing long-context models offer.
How does GLM-5.1 relate to GLM-5?
GLM-5.1 is part of the same GLM-5 family from Zhipu and represents an incremental revision within that line rather than a separate model series. If you are choosing between them, test both on your own prompts, since our dataset does not include complete comparative benchmarks.
Why does GLM-5.1 appear under different names?
Providers list the model under several identifiers, including glm-5.1, zai-org/GLM-5.1, and THUDM/GLM-5.1. These refer to the same model; THUDM and zai-org are organization names associated with Zhipu's published releases.