Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.
GLM-OCR
Compact optical character recognition model combining CogViT with GLM-0.5B for fast, highly accurate document and table extraction.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
GLM-OCR is powered by self-developed CogViT visual encoders and a GLM-0.5B decoder. Pre-trained on billions of image-text pairs via CLIP alignment, it parses dense financial statements, multi-column research papers, and complex tables at $0.03/1M tokens.
Strengths & Considerations
- 98.4% character accuracy on complex printed and scanned documents
- Precise table structure markdown extraction
- Flat $0.03 / 1M token pricing
- Specialized document extraction model, not designed for conversational dialogue
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Z.AI and comparable reasoning engines.
The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.
Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.
Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.