Skip to main content
Z AI

GLM-4.5

GLM-4.5 is a chat-oriented large language model from Z AI, released as part of the company's GLM series of general-purpose assistant models.

Input from
$0.600 / 1M tokens
across 1 provider

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.600$2.20$0.110

Prices updated daily. Last check: Sep 3, 2026

GLM-4.5 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
19.7 / 100
Math
73.7 / 100

Reasoning & Knowledge

  • MMLU-Pro83.5%
  • GPQA Diamond78.2%
  • Humanity's Last Exam13.0%

Coding

  • LiveCodeBench73.8%
  • SciCode34.8%

Math

  • AIME 202573.7%
  • AIME87.3%
  • MATH-50097.9%

Agentic & Tool Use

  • Terminal-Bench Hard22.0%
  • τ²-bench43.0%

Instruction & Long Context

  • IFBench44.1%
  • Long-Context Reasoning51.7%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Z AI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • General-purpose chat and instruction-following model, applicable across drafting, summarization, and Q&A workloads
  • Part of Z AI's GLM series, so teams standardizing on GLM models can keep a consistent prompting style across the family
  • Listed on this page with per-provider pricing, making cost comparison across inference vendors straightforward
  • Available through inference API endpoints rather than requiring self-hosted infrastructure to evaluate
  • Adds a non-US model option for teams that want vendor diversity in their LLM stack
  • Positioned as a broadly capable model rather than a task-specific one, so a single deployment can cover several use cases

Limitations

  • We do not track a confirmed context window for this entry — check the serving provider's documentation before planning long-document workloads
  • No throughput or time-to-first-token measurements are populated in our data, so latency must be verified independently
  • Modality support (for example image input) is not recorded here and should be confirmed with the provider
  • Benchmark scores are not tracked for this entry, so capability comparisons against peers require third-party evaluations
  • Behavior and maximum settings can differ between providers serving the same model name

Key Features

Chat-completion style API access for assistant and multi-turn conversation use
Instruction-following text generation for drafting, editing, and summarization
Part of the GLM model family from Z AI
Served through third-party inference providers with per-token pricing
Suitable for retrieval-augmented pipelines where context is supplied in the prompt
Code assistance and technical question answering as part of general-purpose text handling
Multi-provider availability enabling price and region comparison for the same model

About GLM-4.5

GLM-4.5 is a conversational large language model developed by Z AI (the developer behind the GLM series). It sits in the GLM-4 generation of the family as a general-purpose chat model, intended for assistant-style interaction, instruction following, and text generation tasks rather than for a single narrow specialty. On this page it is listed alongside the providers that serve it through inference APIs, so the same model name may be available from more than one endpoint. Our database records GLM-4.5 as a chat model from Z AI. We do not currently track a confirmed context window length, modality list, or per-provider throughput figures for this entry — the latency and tokens-per-second values we hold are unpopulated, so any speed comparison should be taken from the provider's own documentation or from independent measurement rather than from this listing. Where a specification matters for your deployment, verify it against the serving provider, since providers can differ in the maximum context they expose, the quantization they run, and the API surface they offer. In practice, models in this position are used for the broad middle of LLM workloads: chat assistants, drafting and rewriting, summarization, question answering over supplied text, and code assistance. GLM-4.5's practical appeal for many teams is provider choice — when a model is served by multiple inference vendors, buyers can compare per-token cost, throughput, and region availability for the same weights instead of being tied to a single vendor's terms. Use the pricing table on this page to see which providers currently serve GLM-4.5 and how their rates compare.

Common Use Cases

GLM-4.5 suits general assistant workloads: customer-facing or internal chatbots, drafting and rewriting text, summarizing supplied documents, answering questions over retrieved context, and everyday code explanation or generation tasks. Because it is a general-purpose chat model rather than a specialized reasoning, embedding, or vision entry, it works best where one model needs to cover a spread of text tasks at predictable per-token cost. It is also a reasonable candidate for teams evaluating alternatives to their incumbent chat model — the multi-provider listing on this page makes it easy to price a pilot before committing. For workloads that hinge on very long inputs, strict latency budgets, or image input, confirm those specifications with the serving provider first, since we do not track them for this entry.

Frequently Asked Questions

How much does GLM-4.5 cost to run?

Pricing depends on which inference provider you use and how they bill — input versus output tokens, batch versus real-time serving, and any commitment or throughput tiers. Because rates change frequently and differ by vendor, see the pricing table on this page for the current per-provider figures rather than relying on a fixed number.

What is GLM-4.5 best used for?

General-purpose text and chat work: assistant interfaces, drafting and rewriting, summarization, question answering over supplied context, and code assistance. It is a broadly applicable chat model rather than a narrow specialist.

Who created GLM-4.5?

GLM-4.5 was developed by Z AI, the company behind the GLM series of language models. It is served to end users through inference API providers, which are listed with their pricing on this page.

What context window does GLM-4.5 support?

Our database does not have a confirmed context window for this entry, and providers serving the same model can expose different maximums. Check the documentation of the specific provider you plan to use before designing long-context workloads.

How fast is GLM-4.5?

We do not have populated throughput or time-to-first-token measurements for GLM-4.5, so we cannot state a speed figure here. Output speed also varies by provider, hardware, and load — benchmark your own prompts against the specific endpoint you intend to use.