GLM-4.5 Air is Zhipu's lightweight text model with a 128K token context window, optimized for speed with 82.55 tokens/second output.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.130 | $0.850 | $0.025 | |
| $0.130 | $0.850 | $0.025 | |
| $0.157 | $0.937 | $0.078 | |
| $0.200 | $1.10 | - |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
GLM-4.5 Air is designed for applications requiring fast, efficient text processing where response speed is crucial. Its lightweight architecture makes it suitable for high-volume chat applications, customer service automation, content generation workflows, and real-time text analysis tasks. The 128K context window enables document summarization and analysis, while tool calling support allows integration with external systems and APIs. Organizations prioritizing cost efficiency and speed over maximum model capability will find GLM-4.5 Air appropriate for production deployments requiring consistent, quick responses rather than complex reasoning or creative tasks.
GLM-4.5 Air pricing varies by provider and pricing type (standard vs batch). Check the pricing table above for current rates across all providers.
GLM-4.5 Air excels at high-volume text processing tasks requiring fast response times, such as customer service automation, content generation, and real-time chat applications. Its 128K context window and tool calling capabilities make it suitable for document analysis and API integration workflows where efficiency is prioritized over complex reasoning.
GLM-4.5 Air generates output at 82.55 tokens per second with a 644ms time to first token. This positions it as a speed-optimized option within the lightweight model category, though specific comparisons depend on the particular models and providers being evaluated.