Introducing GLM-5.3-FlashXZ.AIReleased August 26, 2026

GLM-5.3-FlashX

Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.

multimodal-realtimeProprietary API$0.37 / 1M tok (Ultra-Low Latency MoE)Context: 200K (200,000 tokens)Arena ELO: 1475

Technical Specifications

Architecture Type
High-Throughput Hybrid MoE
Total Parameters
320B
Active Parameters (MoE)
18B
Context Window
200K (200,000 tokens)
Max Output Tokens
64K (64,000 tokens)
Knowledge Cutoff
July 2026
Supported Modalities
text, code, vision, reasoning
License & Access
Z.AI Terms of Service

Benchmark Evaluations

Mmlu Pro
86.8
Humaneval
95.2
Time To First Token Ms
120
Chatbot Arena ELO
1475

Deep Architectural Overview

GLM-5.3-FlashX provides accelerated inference nodes delivering ultra-fast Time-to-First-Token (TTFT) for latency-sensitive applications like in-editor autocomplete, interactive terminal commands, and real-time visual inspection.

Strengths & Considerations

Core Strengths
  • Ultra-fast time-to-first-token response
  • Full multimodal visual understanding with 200k context
  • Priced at $0.37 / $1.25 per million tokens
Known Limitations
  • Slight price premium over standard Flash for dedicated throughput

Token & API Pricing

Input Tokens (1M)$0.37
Output Tokens (1M)$1.25
Cached Input (1M)$0.0750
Pricing is verified directly against Z.AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Z.AI and comparable reasoning engines.

Browse all models
Z.AI

The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.

200K ctx$0.15 / 1M tok (320B MoE Visual Coding)
Z.AI

Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.

1M ctx$1.40 / 1M tok (Flagship Coding & Cyber)
Z.AI

Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.

1M ctx$1.40 / 1M tok (1M Lossless Context)
Z.AI

Engineered for long-horizon tasks, able to work independently for up to 8 hours in a single run, aligned with Claude Opus 4.6.

200K ctx$1.40 / 1M tok (8h Autonomous Agent)