Skip to main content
Z AI

GLM-4.5V

GLM-4.5V is a multimodal chat model from Z AI, the vision-capable entry in the company's GLM-4.5 model family.

Input from
$0.600 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.600$1.80$0.110
$0.600$1.80$0.110

Prices updated daily. Last check: Oct 6, 2026

GLM-4.5V pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
6.7 / 100
Math
15.3 / 100

Reasoning & Knowledge

  • MMLU-Pro75.1%
  • GPQA Diamond57.3%
  • Humanity's Last Exam3.5%

Coding

  • LiveCodeBench35.2%

Math

  • AIME 202515.3%

Agentic & Tool Use

  • Terminal-Bench Hard6.8%
  • τ²-bench19.6%

Instruction & Long Context

  • IFBench28.6%
  • Long-Context Reasoning0.0%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Z AI
Modalities
Text

Capabilities

Tool Calling
No

Strengths & Limitations

Strengths

  • Multimodal chat interface — accepts visual input alongside text prompts in a single request
  • Part of the GLM-4.5 family from Z AI, so it shares prompting conventions with the family's text models
  • Served through standard chat-completions style APIs, reducing integration work for teams already calling OpenAI-compatible endpoints
  • Frequently available from more than one inference host, letting buyers compare providers rather than rely on a single endpoint
  • Handles document, chart, and screenshot interpretation workloads without a separate OCR or vision service
  • Positioned as the vision variant within its generation, giving a clear selection rule versus the text-only GLM-4.5 entries

Limitations

  • We do not have a confirmed context window figure for GLM-4.5V in our database
  • Throughput and time-to-first-token measurements in our benchmark record are unpopulated, so speed cannot be compared numerically against peers here
  • Independent benchmark scores for this model are not tracked in our catalog, making capability comparisons qualitative
  • Image-handling limits — resolution caps, images per request, file formats — vary by hosting provider and must be checked per endpoint
  • Less third-party tooling and community documentation than the most widely deployed Western multimodal models

Key Features

•Multimodal input: text prompts combined with visual material
•Chat-completions API surface compatible with common SDK patterns
•Member of the Z AI GLM-4.5 model generation
•Vision-oriented variant distinguished from the text-only GLM-4.5 entries
•Image and document understanding for screenshots, charts, and scanned pages
•Multi-provider availability across hosted inference endpoints
•Text output suitable for downstream parsing and application logic

About GLM-4.5V

GLM-4.5V is a chat model released by Z AI (the team behind the GLM series, also known as Zhipu AI). The "V" designation marks it as the vision-oriented member of the GLM-4.5 generation, positioned alongside the text-focused GLM-4.5 models rather than replacing them. It is served through chat-completions style APIs, so it slots into the same application code as other GLM chat endpoints. As a multimodal assistant, GLM-4.5V is aimed at prompts that mix natural language instructions with visual material — screenshots, documents, charts, photographs, and UI captures — and return text responses. Our catalog does not currently track a confirmed context window, parameter count, or verified throughput and latency figures for this model; the benchmark record we hold from Artificial Analysis has no measured tokens-per-second or time-to-first-token values. Readers who need exact context limits, supported image formats, or per-request size caps should confirm against the documentation of whichever provider they route through, since hosted deployments of the same model often differ in these details. In practice, GLM-4.5V is used where a team wants image understanding and text generation from one endpoint without maintaining a separate vision pipeline. It competes with other multimodal assistants in the open-weight-adjacent Chinese model ecosystem, and like other GLM releases it is frequently offered by multiple inference providers at once, which gives buyers a choice of host rather than a single vendor relationship. Compare the providers listed in the pricing table on this page before committing.

Common Use Cases

GLM-4.5V suits workloads that pair images with instructions: extracting structured information from invoices, forms, and scanned documents; describing or categorizing photographs; reading charts and dashboards and summarizing what they show; and answering questions about UI screenshots during QA or support triage. Because it exposes a conventional chat interface, it also works as a general assistant for text-only turns, which lets a single endpoint serve mixed traffic where only some requests carry an image. Teams evaluating it should prototype against their own image types first, since multimodal accuracy varies sharply by document quality and domain, and should confirm per-provider image size and request limits before scaling. For high-volume text-only classification or long-document processing, a text-focused model in the GLM-4.5 family may be the more economical pick.

Frequently Asked Questions

What is GLM-4.5V best used for?

It is best suited to tasks that combine images with text instructions — document and form extraction, chart and screenshot interpretation, image description and categorization, and visual question answering — while still handling ordinary text chat turns through the same endpoint.

How much does GLM-4.5V cost to run?

Pricing varies by provider and by pricing type, with separate rates typically applied to input and output tokens and additional handling for image input. Check the pricing table on this page for current per-provider rates rather than relying on any fixed figure.

How does GLM-4.5V differ from the other GLM-4.5 models?

The "V" suffix marks it as the vision-capable variant of the GLM-4.5 generation from Z AI. If your workload never includes images, a text-focused model in the same family is usually the simpler choice; if some requests carry screenshots, documents, or photos, GLM-4.5V lets you serve both from one endpoint.

What context window does GLM-4.5V support?

We do not have a confirmed context window figure for GLM-4.5V in our database. Hosted deployments can also impose their own request size limits, so verify the effective maximum with the specific provider you plan to use.

How fast is GLM-4.5V?

Our benchmark record for GLM-4.5V has no populated output-tokens-per-second or time-to-first-token values, so we cannot quote measured speed. Throughput also depends heavily on the serving provider, so run a short latency test against the endpoints listed in the pricing table before choosing one.