Skip to main content
Zhipu

GLM 4.7 Flash

GLM 4.7 Flash is a chat model from Zhipu in the GLM family, positioned by its "Flash" naming as a lighter-weight, throughput-oriented variant of the GLM 4.7 line.

Input from
$0.060 / 1M tokens
across 5 providers

API Pricing

Cheapest on Deep Infra 7% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.060$0.400$0.010
$0.060$0.400$0.010
$0.063$0.400$0.031
$0.070$0.400-
$0.070$0.400$0.010

Prices updated daily. Last check: Sep 1, 2026

GLM 4.7 Flash pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
15.6 / 100

Reasoning & Knowledge

  • GPQA Diamond45.2%
  • Humanity's Last Exam5.0%

Coding

  • SciCode25.5%

Agentic & Tool Use

  • Terminal-Bench Hard3.8%
  • τ²-bench91.8%

Instruction & Long Context

  • IFBench46.3%
  • Long-Context Reasoning18.3%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Zhipu
Family
GLM
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Part of Zhipu's GLM family, a line with multiple published generations and continued iteration
  • "Flash" tier positioning targets latency-sensitive and high-request-volume chat workloads
  • Sibling to the standard GLM 4.7 model, allowing an easy upgrade path within the same family and prompt format
  • Chat-oriented interface fits standard assistant, summarization, and drafting pipelines without task-specific fine-tuning
  • Zhipu's GLM series has notable Chinese-language coverage in addition to English, relevant for bilingual deployments
  • Compact-tier models in this class are typically served by multiple providers, letting buyers compare hosting options

Limitations

  • We do not have a confirmed context window length recorded for this model
  • No published benchmark scores are tracked in our database for GLM 4.7 Flash
  • Throughput and time-to-first-token measurements in our record are unpopulated, so speed claims are not independently verified here
  • As a lighter tier, it is generally a weaker fit than the full GLM 4.7 model for long-horizon reasoning and complex agentic chains
  • Provider availability for Zhipu models is narrower in Western markets than for models from the largest US labs

Key Features

Chat/instruction-following interface for assistant-style interactions
Member of Zhipu's GLM 4.7 generation
"Flash" tier variant aimed at faster, lower-cost serving
Suited to English and Chinese text workloads, consistent with the GLM series' bilingual focus
Available through hosted inference APIs listed in the pricing table on this page
Drop-in sibling to the standard GLM 4.7 model for tier switching

About GLM 4.7 Flash

GLM 4.7 Flash is a text chat model developed by Zhipu (Zhipu AI, also known for publishing GLM models under the Z.ai brand). It belongs to the GLM family, a series of general-purpose assistant models that Zhipu has iterated across numbered generations. Within that family, entries carrying the "Flash" label have conventionally been the smaller, faster-serving siblings of the main GLM release of the same generation, intended for workloads where request volume and response latency matter more than maximum capability on the hardest reasoning tasks. Our database records GLM 4.7 Flash as a chat-oriented model; we do not currently track a confirmed context window length, modality list, or published benchmark scores for it, and the throughput and time-to-first-token figures in our record are unpopulated rather than measured. Readers who need exact context limits, tool-calling semantics, or multilingual coverage should confirm against Zhipu's own model documentation and against the specific provider they intend to route through, since hosted deployments of the same GLM checkpoint can differ in maximum context, streaming behavior, and available API parameters. In practice, models in this position are used for conversational assistants, summarization, content drafting, and batch text processing where a large number of requests must be served economically. GLM 4.7 Flash competes in the same slot as other vendors' compact chat tiers, and its main comparison point is the non-Flash GLM 4.7 model: the Flash variant trades some headroom on complex, multi-step problems for faster, cheaper serving. Compare live availability and rates across providers using the pricing table on this page.

Common Use Cases

GLM 4.7 Flash fits high-volume, latency-sensitive text work: customer-facing chat assistants, FAQ and support deflection, message classification and routing, short-form summarization, metadata and tag extraction, and first-draft content generation that a larger model or a human then refines. Its Flash tier positioning makes it a candidate for the cheap-and-fast leg of a routing setup, where straightforward requests are handled by the compact model and only ambiguous or multi-step ones escalate to the full GLM 4.7 or another larger model. Teams working with Chinese-language content alongside English may find the GLM series worth benchmarking against Western compact tiers. For long-context document analysis or complex agentic tool chains, verify the deployed context limit and tool-calling support with your provider before committing, since we do not track confirmed figures for those on this model.

Frequently Asked Questions

How much does GLM 4.7 Flash cost?

Pricing varies by provider and by pricing type — input tokens, output tokens, cached input, and any batch or committed-throughput discounts are all priced separately, and rates change frequently. Check the pricing table on this page for current per-provider rates rather than relying on a fixed figure.

What is GLM 4.7 Flash best used for?

It suits high-throughput chat and text-processing tasks — support assistants, classification and routing, summarization, extraction, and bulk drafting — where per-request cost and response latency matter more than peak reasoning performance.

How does GLM 4.7 Flash differ from GLM 4.7?

They are siblings in the same generation. The Flash variant is the lighter, faster-serving tier, while the standard GLM 4.7 is the larger model intended for harder reasoning and longer multi-step tasks. Many teams route easy traffic to the Flash tier and escalate the rest.

What context window does GLM 4.7 Flash support?

We do not have a confirmed context window recorded for this model. Hosted deployments of the same checkpoint can also expose different maximum context lengths, so check Zhipu's documentation and your chosen provider's model page before designing around a specific limit.

Who makes GLM 4.7 Flash?

It is developed by Zhipu, the Chinese AI lab behind the GLM series of language models, which it also publishes under the Z.ai brand.