Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.
GLM-4.6V
Multimodal vision model with native function calling and controllable thinking mode switch over 128k tokens.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
GLM-4.6V excels at parsing dense technical charts, architectural diagrams, financial reports, and UI screenshots. It supports seamless function calling directly from image inputs at an economical $0.30/$0.90 price point.
Strengths & Considerations
- Native function calling from images and documents
- Controllable thinking mode switch
- Affordable $0.30 input / $0.05 cache hit pricing
- Replaced by GLM-5.3-Flash for native visual web coding
Token & API Pricing
Similar & Alternative Models
Explore other frontier models from Z.AI and comparable reasoning engines.
The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.
Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.
Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.