Introducing GLM-4-VoiceZ.AIReleased October 24, 2024

GLM-4-Voice

End-to-end speech conversational model capable of understanding and generating emotional, realtime human-like speech without intermediate ASR/TTS.

speech-to-speechOpen WeightsFree Open Weights / End-to-End VoiceContext: 32.768K (32,768 tokens)

Technical Specifications

Architecture Type
End-to-End Realtime Speech Transformer
Total Parameters
9B (estimated)
Context Window
32.768K (32,768 tokens)
Max Output Tokens
4.096K (4,096 tokens)
Knowledge Cutoff
June 2024
Supported Modalities
audio, text
License & Access
Apache 2.0
Weights Formats
safetensors, pth

Benchmark Evaluations

Latency Ms
180
Speech Naturalness Mos
4.35

Deep Architectural Overview

GLM-4-Voice is an open-source end-to-end voice model. By directly modeling audio tokens alongside text tokens, it achieves low-latency duplex conversation, emotion control, accent adaptation, and interruptible speech.

Strengths & Considerations

Core Strengths
  • True end-to-end speech-to-speech architecture
  • Sub-200ms conversational response latency
  • Open weights under Apache 2.0
Known Limitations
  • Audio token vocabulary increases token consumption compared to pure text

Token & API Pricing

Input Tokens (1M)$0.50
Output Tokens (1M)$1.00
Self-Hosted Min VRAMVaries by quantization
Pricing is verified directly against Z.AI's developer documentation and API rate sheets.

Similar & Alternative Models

Explore other frontier models from Z.AI and comparable reasoning engines.

Browse all models
Z.AI

Ultra-low latency version of GLM-5.3-Flash engineered for high-concurrency real-time IDE code completion and interactive agents.

200K ctx$0.37 / 1M tok (Ultra-Low Latency MoE)
Z.AI

The first native multimodal model in the GLM-5 series: 320B parameters (18B active) combining linear and sparse attention for low-cost visual coding.

200K ctx$0.15 / 1M tok (320B MoE Visual Coding)
Z.AI

Z.AI’s premier flagship model, delivering a 50% performance gain on Code Bench and matching Claude Mythos 5 in cybersecurity and vulnerability discovery.

1M ctx$1.40 / 1M tok (Flagship Coding & Cyber)
Z.AI

Flagship model built for project-scale context, supporting truly usable 1M-token context with 128k output.

1M ctx$1.40 / 1M tok (1M Lossless Context)