Skip to main content
Z AI

GLM-4.6V

GLM-4.6V is a vision-capable chat model from Z AI, the multimodal variant in the company's GLM-4.6 series of general-purpose language models.

Input from
$0.300 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$0.300$0.900$0.055
$0.300$0.900$0.055

Prices updated daily. Last check: Sep 3, 2026

GLM-4.6V pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
10.9 / 100
Math
26.3 / 100

Reasoning & Knowledge

  • MMLU-Pro75.2%
  • GPQA Diamond56.6%
  • Humanity's Last Exam3.7%

Coding

  • LiveCodeBench41.1%
  • SciCode27.2%

Math

  • AIME 202526.3%

Agentic & Tool Use

  • Terminal-Bench Hard3.0%
  • τ²-bench30.7%

Instruction & Long Context

  • IFBench27.9%
  • Long-Context Reasoning14.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Z AI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Part of the GLM-4.6 generation from Z AI, so it can be swapped into stacks already built around GLM models with minimal prompt rework
  • The "V" designation indicates a vision-oriented variant, giving GLM family users a multimodal option within the same generation
  • Exposed as a standard chat model, so it works with conventional chat-completions client libraries
  • Available through hosted API providers, avoiding the need to provision and manage GPUs directly
  • Multiple-provider availability means pricing and latency can be compared side by side rather than accepting a single vendor's terms
  • Z AI's GLM line has a track record of successive releases, which makes upgrade paths within the family predictable

Limitations

  • We do not have a confirmed context window length on file for this entry, so the usable prompt size must be verified with the serving provider
  • Published throughput and time-to-first-token measurements are not populated in our data, making speed comparisons against peers difficult without your own benchmarking
  • We do not track verified benchmark scores for GLM-4.6V, so capability claims cannot be compared numerically against other models here
  • Provider coverage for GLM models is generally narrower than for the most widely hosted Western model families, which can limit region and redundancy options
  • Documentation and community tooling for GLM models is less extensive in English than for more widely adopted families

Key Features

Chat completions interface for multi-turn conversational use
Vision-oriented variant within the GLM-4.6 generation
Developed by Z AI as part of the ongoing GLM model series
Available through hosted inference APIs listed in the pricing table on this page
Latency and throughput tracking sourced from Artificial Analysis where measurements are available
Drop-in compatibility with existing GLM-4.6 prompt and application patterns

About GLM-4.6V

GLM-4.6V is a chat model published by Z AI (the developer behind the GLM series). The "V" suffix marks it as the vision-oriented member of the GLM-4.6 generation, positioned alongside the text-focused GLM-4.6 models rather than as a separate product line. It is designed to be used through a standard chat completions interface, taking conversational turns and returning generated text. Our catalog entry for GLM-4.6V carries limited verified specifications: we confirm the creator (Z AI), the model type (chat), and that throughput and latency figures are sourced from Artificial Analysis. We do not currently track a confirmed context window length, parameter count, or license status for this entry, and the throughput and time-to-first-token measurements we have on file are unpopulated, so real-world serving speed will depend on which provider you route to. Readers who need an exact context limit or modality list should check the serving provider's own model card, since hosted deployments of the same GLM checkpoint sometimes differ in the maximum context they expose. In practice, GLM-4.6V is most relevant to teams already evaluating the GLM family — for example, those using GLM-4.6 for text workloads who need a sibling model in the same generation that handles visual input. Because multiple providers host GLM models, the practical comparison between GLM-4.6V and peers from other creators often comes down to per-provider pricing, context limits, and measured latency rather than the checkpoint alone. Use the pricing table on this page to compare the providers currently serving it.

Common Use Cases

GLM-4.6V suits teams that already run GLM-family models and want a same-generation option for workloads involving visual input alongside text — document and screenshot understanding, image-grounded question answering, and assistant flows where a user may attach an image mid-conversation. It also fits general chat assistant work, content drafting, and summarization where an organization is deliberately diversifying away from a single model vendor. Because we do not have confirmed context window or benchmark figures for this entry, teams planning long-document ingestion or latency-sensitive interactive products should run a short evaluation against their own prompts and confirm limits with the chosen provider before committing.

Frequently Asked Questions

How much does GLM-4.6V cost to run?

Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are billed separately, and the same model can differ noticeably in cost between hosts. Check the pricing table on this page for the current rates from each provider serving GLM-4.6V.

What is GLM-4.6V best used for?

It is a chat model, so it fits conversational assistants, drafting, and summarization. The "V" designation marks it as the vision-oriented variant of the GLM-4.6 generation, making it a reasonable choice for image-grounded question answering and document or screenshot understanding within a GLM-based stack.

Who created GLM-4.6V?

GLM-4.6V comes from Z AI, the developer behind the GLM series of language models.

How does GLM-4.6V differ from GLM-4.6?

They belong to the same generation of Z AI's GLM line. GLM-4.6V is the vision-oriented variant, while GLM-4.6 is the text-focused model. If your workload never involves images, the text variant is usually the simpler starting point; if you need visual input in the same generation, GLM-4.6V is the corresponding option.

What is GLM-4.6V's context window?

We do not have a confirmed context window length on file for this entry. Because hosted deployments of the same checkpoint sometimes expose different maximum context lengths, confirm the limit with the specific provider you plan to use.

How fast is GLM-4.6V?

Our throughput and time-to-first-token fields for this model, sourced from Artificial Analysis, are currently unpopulated. Serving speed also depends heavily on the provider and region, so we recommend benchmarking against your own prompts on the specific endpoint you intend to use.