GLM 4.7 Flash
GLM 4.7 Flash is a chat model from Zhipu in the GLM family, positioned by its "Flash" naming as a lighter-weight, throughput-oriented variant of the GLM 4.7 line.
API Pricing
Cheapest on Deep Infra — 7% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.060 | $0.400 | $0.010 | |
| $0.060 | $0.400 | $0.010 | |
| $0.063 | $0.400 | $0.031 | |
| $0.070 | $0.400 | - | |
| $0.070 | $0.400 | $0.010 |
Prices updated daily. Last check: Sep 1, 2026
GLM 4.7 Flash pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- GPQA Diamond45.2%
- Humanity's Last Exam5.0%
Coding
- SciCode25.5%
Agentic & Tool Use
- Terminal-Bench Hard3.8%
- τ²-bench91.8%
Instruction & Long Context
- IFBench46.3%
- Long-Context Reasoning18.3%
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Zhipu
- Family
- GLM
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No
Strengths & Limitations
Strengths
- Part of Zhipu's GLM family, a line with multiple published generations and continued iteration
- "Flash" tier positioning targets latency-sensitive and high-request-volume chat workloads
- Sibling to the standard GLM 4.7 model, allowing an easy upgrade path within the same family and prompt format
- Chat-oriented interface fits standard assistant, summarization, and drafting pipelines without task-specific fine-tuning
- Zhipu's GLM series has notable Chinese-language coverage in addition to English, relevant for bilingual deployments
- Compact-tier models in this class are typically served by multiple providers, letting buyers compare hosting options
Limitations
- We do not have a confirmed context window length recorded for this model
- No published benchmark scores are tracked in our database for GLM 4.7 Flash
- Throughput and time-to-first-token measurements in our record are unpopulated, so speed claims are not independently verified here
- As a lighter tier, it is generally a weaker fit than the full GLM 4.7 model for long-horizon reasoning and complex agentic chains
- Provider availability for Zhipu models is narrower in Western markets than for models from the largest US labs
Key Features
About GLM 4.7 Flash
Common Use Cases
GLM 4.7 Flash fits high-volume, latency-sensitive text work: customer-facing chat assistants, FAQ and support deflection, message classification and routing, short-form summarization, metadata and tag extraction, and first-draft content generation that a larger model or a human then refines. Its Flash tier positioning makes it a candidate for the cheap-and-fast leg of a routing setup, where straightforward requests are handled by the compact model and only ambiguous or multi-step ones escalate to the full GLM 4.7 or another larger model. Teams working with Chinese-language content alongside English may find the GLM series worth benchmarking against Western compact tiers. For long-context document analysis or complex agentic tool chains, verify the deployed context limit and tool-calling support with your provider before committing, since we do not track confirmed figures for those on this model.
Frequently Asked Questions
How much does GLM 4.7 Flash cost?
Pricing varies by provider and by pricing type — input tokens, output tokens, cached input, and any batch or committed-throughput discounts are all priced separately, and rates change frequently. Check the pricing table on this page for current per-provider rates rather than relying on a fixed figure.
What is GLM 4.7 Flash best used for?
It suits high-throughput chat and text-processing tasks — support assistants, classification and routing, summarization, extraction, and bulk drafting — where per-request cost and response latency matter more than peak reasoning performance.
How does GLM 4.7 Flash differ from GLM 4.7?
They are siblings in the same generation. The Flash variant is the lighter, faster-serving tier, while the standard GLM 4.7 is the larger model intended for harder reasoning and longer multi-step tasks. Many teams route easy traffic to the Flash tier and escalate the rest.
What context window does GLM 4.7 Flash support?
We do not have a confirmed context window recorded for this model. Hosted deployments of the same checkpoint can also expose different maximum context lengths, so check Zhipu's documentation and your chosen provider's model page before designing around a specific limit.
Who makes GLM 4.7 Flash?
It is developed by Zhipu, the Chinese AI lab behind the GLM series of language models, which it also publishes under the Z.ai brand.