GLM-4.6V
GLM-4.6V is a vision-capable chat model from Z AI, the multimodal variant in the company's GLM-4.6 series of general-purpose language models.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.300 | $0.900 | $0.055 | |
| $0.300 | $0.900 | $0.055 |
Prices updated daily. Last check: Sep 3, 2026
GLM-4.6V pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- MMLU-Pro75.2%
- GPQA Diamond56.6%
- Humanity's Last Exam3.7%
Coding
- LiveCodeBench41.1%
- SciCode27.2%
Math
- AIME 202526.3%
Agentic & Tool Use
- Terminal-Bench Hard3.0%
- τ²-bench30.7%
Instruction & Long Context
- IFBench27.9%
- Long-Context Reasoning14.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Z AI
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Part of the GLM-4.6 generation from Z AI, so it can be swapped into stacks already built around GLM models with minimal prompt rework
- The "V" designation indicates a vision-oriented variant, giving GLM family users a multimodal option within the same generation
- Exposed as a standard chat model, so it works with conventional chat-completions client libraries
- Available through hosted API providers, avoiding the need to provision and manage GPUs directly
- Multiple-provider availability means pricing and latency can be compared side by side rather than accepting a single vendor's terms
- Z AI's GLM line has a track record of successive releases, which makes upgrade paths within the family predictable
Limitations
- We do not have a confirmed context window length on file for this entry, so the usable prompt size must be verified with the serving provider
- Published throughput and time-to-first-token measurements are not populated in our data, making speed comparisons against peers difficult without your own benchmarking
- We do not track verified benchmark scores for GLM-4.6V, so capability claims cannot be compared numerically against other models here
- Provider coverage for GLM models is generally narrower than for the most widely hosted Western model families, which can limit region and redundancy options
- Documentation and community tooling for GLM models is less extensive in English than for more widely adopted families
Key Features
About GLM-4.6V
Common Use Cases
GLM-4.6V suits teams that already run GLM-family models and want a same-generation option for workloads involving visual input alongside text — document and screenshot understanding, image-grounded question answering, and assistant flows where a user may attach an image mid-conversation. It also fits general chat assistant work, content drafting, and summarization where an organization is deliberately diversifying away from a single model vendor. Because we do not have confirmed context window or benchmark figures for this entry, teams planning long-document ingestion or latency-sensitive interactive products should run a short evaluation against their own prompts and confirm limits with the chosen provider before committing.
Frequently Asked Questions
How much does GLM-4.6V cost to run?
Pricing varies by provider and by pricing type — input tokens, output tokens, and any cached or batch rates are billed separately, and the same model can differ noticeably in cost between hosts. Check the pricing table on this page for the current rates from each provider serving GLM-4.6V.
What is GLM-4.6V best used for?
It is a chat model, so it fits conversational assistants, drafting, and summarization. The "V" designation marks it as the vision-oriented variant of the GLM-4.6 generation, making it a reasonable choice for image-grounded question answering and document or screenshot understanding within a GLM-based stack.
Who created GLM-4.6V?
GLM-4.6V comes from Z AI, the developer behind the GLM series of language models.
How does GLM-4.6V differ from GLM-4.6?
They belong to the same generation of Z AI's GLM line. GLM-4.6V is the vision-oriented variant, while GLM-4.6 is the text-focused model. If your workload never involves images, the text variant is usually the simpler starting point; if you need visual input in the same generation, GLM-4.6V is the corresponding option.
What is GLM-4.6V's context window?
We do not have a confirmed context window length on file for this entry. Because hosted deployments of the same checkpoint sometimes expose different maximum context lengths, confirm the limit with the specific provider you plan to use.
How fast is GLM-4.6V?
Our throughput and time-to-first-token fields for this model, sourced from Artificial Analysis, are currently unpopulated. Serving speed also depends heavily on the provider and region, so we recommend benchmarking against your own prompts on the specific endpoint you intend to use.