Introducing GLM-4.5-AirZ.AIReleased July 28, 2025

GLM-4.5-Air

Cost-effective high-performance variant in the GLM-4.5 family, priced at only $0.20 per million input tokens.

text-to-textProprietary API$0.20 / 1M tok (Cost-Effective Air)Context: 128K (128,000 tokens)Arena ELO: 1265

Technical Specifications

Architecture Type
Cost-Effective High-Performance Transformer
Total Parameters
Efficient Scale
Context Window
128K (128,000 tokens)
Max Output Tokens
32.768K (32,768 tokens)
Knowledge Cutoff
June 2025
Supported Modalities
text, code
License & Access
Z.AI Terms of Service

Benchmark Evaluations

Mmlu
80.5
Humaneval
82
Chatbot Arena ELO
1265

Deep Architectural Overview

GLM-4.5-Air delivers balanced intelligence for high-throughput batch extraction, summarization, and daily coding assistance with cache read hits as low as $0.03 per million tokens.

Strengths & Considerations

Core Strengths
  • Very affordable at $0.20 input / $1.10 output
  • $0.03 cached input reads
  • 128k context retention
Known Limitations
  • Lightweight model, not designed for deep mathematical reasoning

Token & API Pricing

Input Tokens (1M)$0.20
Output Tokens (1M)$1.10
Cached Input (1M)$0.0300
Pricing is verified directly against Z.AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Z.AI and comparable reasoning engines.

Browse all models
Z.AI

Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.

200K ctx$0.37 / 1M tok (Ultra-Low Latency MoE)
Z.AI

The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.

200K ctx$0.15 / 1M tok (320B MoE Visual Coding)
Z.AI

Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.

1M ctx$1.40 / 1M tok (Flagship Coding & Cyber)
Z.AI

Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.

1M ctx$1.40 / 1M tok (1M Lossless Context)