GLM 5V Turbo
GLM 5V Turbo is a chat model from Z AI, part of the company's GLM series, positioned as a throughput-oriented "Turbo" variant within that lineup.
API Pricing
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $1.20 | $4.00 | $0.240 | |
| $1.20 | $4.00 | $0.240 |
Prices updated daily. Last check: Sep 30, 2026
GLM 5V Turbo pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond80.9%
- Humanity's Last Exam17.1%
Agentic & Tool Use
- Terminal-Bench Hard32.6%
- τ²-bench98.5%
Instruction & Long Context
- IFBench61.1%
- Long-Context Reasoning70.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Z AI
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Part of Z AI's GLM series, giving access to a model line with an established API surface and multiple hosting options
- "Turbo" positioning within the GLM 5V generation targets faster serving over maximum capability, which suits latency-sensitive deployments
- Available as a chat/completions-style endpoint, so it drops into existing OpenAI-compatible client code used for other GLM models
- GLM-series models are widely deployed for Chinese and English workloads, making the family a common choice for bilingual applications
- Provider availability and rates for this model can be compared side by side in the pricing table on this page
- Sits alongside larger GLM variants, allowing a tiered routing setup where simple turns go to the Turbo model and harder ones escalate
Limitations
- We do not track a confirmed context window for GLM 5V Turbo, so maximum input length must be verified with the serving provider
- No measured throughput or time-to-first-token values are available in our benchmark data yet — the Artificial Analysis figures are unpopulated
- No published benchmark scores in our database, making objective capability comparison against peers difficult
- Turbo-tier models in general trade some reasoning depth for speed, so complex multi-step tasks may be better served by a larger GLM variant
- Provider coverage for GLM models is narrower than for the most widely hosted open-weight families, which can limit region and redundancy choices
Key Features
About GLM 5V Turbo
Common Use Cases
GLM 5V Turbo fits workloads where response speed and per-request cost matter more than maximum reasoning depth: customer-facing chat assistants, in-product Q&A, summarization and rewriting pipelines, classification and tagging of user text, and the high-volume first pass in a tiered routing setup that escalates difficult requests to a larger GLM model. Teams building for Chinese-language or mixed Chinese/English audiences often evaluate the GLM series specifically for that reason. Before committing it to workloads that depend on long documents or image inputs, confirm the context window and supported modalities with your chosen provider, since our database does not yet carry verified values for either on this model.
Frequently Asked Questions
How much does GLM 5V Turbo cost to run?
Pricing depends on which provider you use and the pricing model they offer — per-token serverless rates, batch discounts, and dedicated capacity are all priced differently, and rates change frequently. See the pricing table on this page for current per-provider figures rather than relying on a fixed number.
What is GLM 5V Turbo best used for?
It is best suited to latency-sensitive and high-volume chat workloads: assistants, Q&A over short inputs, summarization, rewriting, and text classification. Its Turbo positioning within the GLM 5V generation signals a tilt toward fast serving, so very long multi-step reasoning or agentic chains are usually better matched to a larger model in the family.
What context window does GLM 5V Turbo support?
We do not have a verified context window for this model in our database. Because GLM models are served by multiple hosts that sometimes cap context differently, check the documentation of the specific provider listed in the pricing table before designing around a particular input length.
Who makes GLM 5V Turbo?
It is made by Z AI (Zhipu AI), the developer of the GLM model series. The "5V" indicates the model generation and the "Turbo" suffix marks it as the speed-oriented variant within that generation.
Does GLM 5V Turbo accept image input or support tool calling?
We do not track confirmed modality or tool-calling support for this specific model, so we cannot state either way. Both capabilities are frequently gated by the serving provider as well as the model itself — verify against your provider's API reference before building on them.