Skip to main content
Z AI

GLM 5V Turbo

GLM 5V Turbo is a chat model from Z AI, part of the company's GLM series, positioned as a throughput-oriented "Turbo" variant within that lineup.

Input from
$1.20 / 1M tokens
across 2 providers

API Pricing

ProviderInput / 1MOutput / 1MCached / 1M
$1.20$4.00$0.240
$1.20$4.00$0.240

Prices updated daily. Last check: Sep 30, 2026

GLM 5V Turbo pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
23.5 / 100

Reasoning & Knowledge

  • GPQA Diamond80.9%
  • Humanity's Last Exam17.1%

Agentic & Tool Use

  • Terminal-Bench Hard32.6%
  • τ²-bench98.5%

Instruction & Long Context

  • IFBench61.1%
  • Long-Context Reasoning70.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Z AI
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Part of Z AI's GLM series, giving access to a model line with an established API surface and multiple hosting options
  • "Turbo" positioning within the GLM 5V generation targets faster serving over maximum capability, which suits latency-sensitive deployments
  • Available as a chat/completions-style endpoint, so it drops into existing OpenAI-compatible client code used for other GLM models
  • GLM-series models are widely deployed for Chinese and English workloads, making the family a common choice for bilingual applications
  • Provider availability and rates for this model can be compared side by side in the pricing table on this page
  • Sits alongside larger GLM variants, allowing a tiered routing setup where simple turns go to the Turbo model and harder ones escalate

Limitations

  • We do not track a confirmed context window for GLM 5V Turbo, so maximum input length must be verified with the serving provider
  • No measured throughput or time-to-first-token values are available in our benchmark data yet — the Artificial Analysis figures are unpopulated
  • No published benchmark scores in our database, making objective capability comparison against peers difficult
  • Turbo-tier models in general trade some reasoning depth for speed, so complex multi-step tasks may be better served by a larger GLM variant
  • Provider coverage for GLM models is narrower than for the most widely hosted open-weight families, which can limit region and redundancy choices

Key Features

•Chat completion interface for conversational and instruction-following requests
•Member of the Z AI GLM 5V model generation
•Turbo variant tuned toward lower-latency serving within its generation
•Typically exposed through OpenAI-compatible API endpoints by GLM hosts
•Bilingual Chinese/English usage common across the GLM series
•Multi-turn conversation handling with system prompt support
•Provider-level price comparison available on this page

About GLM 5V Turbo

GLM 5V Turbo is a chat-oriented large language model released by Z AI (Zhipu AI), the developer behind the GLM model series. Within that series, Z AI uses the "V" designation for its vision-language line and the "Turbo" suffix for variants tuned toward faster, lower-overhead serving rather than maximum capability, so GLM 5V Turbo sits as a speed-oriented member of the GLM 5V generation rather than a separate architecture family. Our database currently holds limited verified specifications for this model. We do not track a confirmed context window, modality list, or public benchmark scores for GLM 5V Turbo — the throughput and time-to-first-token figures sourced from Artificial Analysis are not yet populated with measured values. Readers who need exact limits on context length, image input, structured output, or tool calling should check the documentation of whichever provider they plan to route requests through, since serving parameters for GLM models often differ between Z AI's own API and third-party hosts. In practice, models in the GLM 5V Turbo position are typically selected by teams that already want a GLM-series model and are optimizing for response latency and request volume rather than for the heaviest reasoning workloads. The pricing table on this page shows which providers currently serve GLM 5V Turbo and how their rates compare, which is usually the deciding factor between a Turbo-tier model and a larger sibling in the same family.

Common Use Cases

GLM 5V Turbo fits workloads where response speed and per-request cost matter more than maximum reasoning depth: customer-facing chat assistants, in-product Q&A, summarization and rewriting pipelines, classification and tagging of user text, and the high-volume first pass in a tiered routing setup that escalates difficult requests to a larger GLM model. Teams building for Chinese-language or mixed Chinese/English audiences often evaluate the GLM series specifically for that reason. Before committing it to workloads that depend on long documents or image inputs, confirm the context window and supported modalities with your chosen provider, since our database does not yet carry verified values for either on this model.

Frequently Asked Questions

How much does GLM 5V Turbo cost to run?

Pricing depends on which provider you use and the pricing model they offer — per-token serverless rates, batch discounts, and dedicated capacity are all priced differently, and rates change frequently. See the pricing table on this page for current per-provider figures rather than relying on a fixed number.

What is GLM 5V Turbo best used for?

It is best suited to latency-sensitive and high-volume chat workloads: assistants, Q&A over short inputs, summarization, rewriting, and text classification. Its Turbo positioning within the GLM 5V generation signals a tilt toward fast serving, so very long multi-step reasoning or agentic chains are usually better matched to a larger model in the family.

What context window does GLM 5V Turbo support?

We do not have a verified context window for this model in our database. Because GLM models are served by multiple hosts that sometimes cap context differently, check the documentation of the specific provider listed in the pricing table before designing around a particular input length.

Who makes GLM 5V Turbo?

It is made by Z AI (Zhipu AI), the developer of the GLM model series. The "5V" indicates the model generation and the "Turbo" suffix marks it as the speed-oriented variant within that generation.

Does GLM 5V Turbo accept image input or support tool calling?

We do not track confirmed modality or tool-calling support for this specific model, so we cannot state either way. Both capabilities are frequently gated by the serving provider as well as the model itself — verify against your provider's API reference before building on them.