Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.
GLM-4-32B-0414-128K
A 32B parameter model delivering high intelligence at unmatched cost efficiency with flat $0.10/1M input and output pricing.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
GLM-4-32B-0414-128K provides enterprise users with symmetric $0.10/1M pricing across both prompt input and generation output, offering a reliable, cost-effective workhorse for classification, summarization, and data pipelines over 128k tokens.
Strengths & Considerations
- Symmetric $0.10 / 1M pricing on both inputs and outputs
- 128k token context window
- Optimal throughput-to-cost ratio for 32B scale
- Non-reasoning architecture; succeeded by GLM-4.7-Flash for coding
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Z.AI and comparable reasoning engines.
The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.
Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.
Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.