Introducing GLM-4-32B-0414-128KZ.AIReleased April 14, 2025

GLM-4-32B-0414-128K

A 32B parameter model delivering high intelligence at unmatched cost efficiency with flat $0.10/1M input and output pricing.

text-to-textProprietary API$0.10 / 1M tok (Symmetric Price)Context: 128K (128,000 tokens)Arena ELO: 1240

Technical Specifications

Architecture Type
Dense Autoregressive Transformer
Total Parameters
32B
Context Window
128K (128,000 tokens)
Max Output Tokens
16.384K (16,384 tokens)
Knowledge Cutoff
March 2025
Supported Modalities
text, code
License & Access
Z.AI Terms of Service

Benchmark Evaluations

Mmlu
81.2
Gsm8k
84.5
Humaneval
78.4
Chatbot Arena ELO
1240

Deep Architectural Overview

GLM-4-32B-0414-128K provides enterprise users with symmetric $0.10/1M pricing across both prompt input and generation output, offering a reliable, cost-effective workhorse for classification, summarization, and data pipelines over 128k tokens.

Strengths & Considerations

Core Strengths
  • Symmetric $0.10 / 1M pricing on both inputs and outputs
  • 128k token context window
  • Optimal throughput-to-cost ratio for 32B scale
Known Limitations
  • Non-reasoning architecture; succeeded by GLM-4.7-Flash for coding

Token & API Pricing

Input Tokens (1M)$0.10
Output Tokens (1M)$0.10
Pricing is verified directly against Z.AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Z.AI and comparable reasoning engines.

Browse all models
Z.AI

Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.

200K ctx$0.37 / 1M tok (Ultra-Low Latency MoE)
Z.AI

The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.

200K ctx$0.15 / 1M tok (320B MoE Visual Coding)
Z.AI

Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.

1M ctx$1.40 / 1M tok (Flagship Coding & Cyber)
Z.AI

Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.

1M ctx$1.40 / 1M tok (1M Lossless Context)