GLM-4.5
GLM-4.5 is a chat-oriented large language model from Z AI, released as part of the company's GLM series of general-purpose assistant models.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.600 | $2.20 | $0.110 |
Prices updated daily. Last check: Sep 3, 2026
GLM-4.5 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro83.5%
- GPQA Diamond78.2%
- Humanity's Last Exam13.0%
Coding
- LiveCodeBench73.8%
- SciCode34.8%
Math
- AIME 202573.7%
- AIME87.3%
- MATH-50097.9%
Agentic & Tool Use
- Terminal-Bench Hard22.0%
- τ²-bench43.0%
Instruction & Long Context
- IFBench44.1%
- Long-Context Reasoning51.7%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Z AI
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- General-purpose chat and instruction-following model, applicable across drafting, summarization, and Q&A workloads
- Part of Z AI's GLM series, so teams standardizing on GLM models can keep a consistent prompting style across the family
- Listed on this page with per-provider pricing, making cost comparison across inference vendors straightforward
- Available through inference API endpoints rather than requiring self-hosted infrastructure to evaluate
- Adds a non-US model option for teams that want vendor diversity in their LLM stack
- Positioned as a broadly capable model rather than a task-specific one, so a single deployment can cover several use cases
Limitations
- We do not track a confirmed context window for this entry — check the serving provider's documentation before planning long-document workloads
- No throughput or time-to-first-token measurements are populated in our data, so latency must be verified independently
- Modality support (for example image input) is not recorded here and should be confirmed with the provider
- Benchmark scores are not tracked for this entry, so capability comparisons against peers require third-party evaluations
- Behavior and maximum settings can differ between providers serving the same model name
Key Features
About GLM-4.5
Common Use Cases
GLM-4.5 suits general assistant workloads: customer-facing or internal chatbots, drafting and rewriting text, summarizing supplied documents, answering questions over retrieved context, and everyday code explanation or generation tasks. Because it is a general-purpose chat model rather than a specialized reasoning, embedding, or vision entry, it works best where one model needs to cover a spread of text tasks at predictable per-token cost. It is also a reasonable candidate for teams evaluating alternatives to their incumbent chat model — the multi-provider listing on this page makes it easy to price a pilot before committing. For workloads that hinge on very long inputs, strict latency budgets, or image input, confirm those specifications with the serving provider first, since we do not track them for this entry.
Frequently Asked Questions
How much does GLM-4.5 cost to run?
Pricing depends on which inference provider you use and how they bill — input versus output tokens, batch versus real-time serving, and any commitment or throughput tiers. Because rates change frequently and differ by vendor, see the pricing table on this page for the current per-provider figures rather than relying on a fixed number.
What is GLM-4.5 best used for?
General-purpose text and chat work: assistant interfaces, drafting and rewriting, summarization, question answering over supplied context, and code assistance. It is a broadly applicable chat model rather than a narrow specialist.
Who created GLM-4.5?
GLM-4.5 was developed by Z AI, the company behind the GLM series of language models. It is served to end users through inference API providers, which are listed with their pricing on this page.
What context window does GLM-4.5 support?
Our database does not have a confirmed context window for this entry, and providers serving the same model can expose different maximums. Check the documentation of the specific provider you plan to use before designing long-context workloads.
How fast is GLM-4.5?
We do not have populated throughput or time-to-first-token measurements for GLM-4.5, so we cannot state a speed figure here. Output speed also varies by provider, hardware, and load — benchmark your own prompts against the specific endpoint you intend to use.