Z AI
GLM-5.3-Flash
GLM-5.3-Flash by Z AI — compare inference API pricing across providers.
Input from
$0.075 / 1M tokens
across 4 providers
API Pricing
Cheapest on OpenRouter — 43% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.075 | $0.250 | $0.015 | |
| $0.150 | $0.500 | $0.030 | |
| $0.150 | $0.500 | $0.075 | |
| $0.150 | $0.500 | - |
Prices updated daily. Last check: Aug 28, 2026
GLM-5.3-Flash pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Intelligence
57.5 / 100
Coding
71.5 / 100
Output Speed
49.2 t/s
Latency (TTFT)
1.2s
Reasoning & Knowledge
- GPQA Diamond91.2%
- Humanity's Last Exam39.9%
Coding
- SciCode46.1%
Agentic & Tool Use
- Terminal-Bench v2.184.3%
- τ-bench Banking47.2%
Instruction & Long Context
- Long-Context Reasoning78.0%
Benchmarks measured Aug 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Z AI
- Modalities
- Text
Capabilities
- Tool Calling
- No
- Open Source
- No