Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.
Qwen3.8-Max
Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.
Qwen3.8-Max Quick Facts & Executive Summary
Qwen3.8-Max by Alibaba Cloud (August 3, 2026) is a 2.4T Sparse Mixture-of-Experts (MoE) Architecture model featuring a 1M (1,000,000 tokens) context window and USD 2 / 1M input tok; USD 6 / 1M output tok. Key highlights include Massive 2.4T parameter MoE architecture activating 95B parameters.
Technical Specifications
Benchmark Evaluations
Deep Architectural Overview
Qwen3.8-Max is Alibaba's frontier Mixture-of-Experts model, activating 95 billion of its 2.4 trillion parameters per token. Optimized for long-horizon agentic orchestration, enterprise coding, and multimodal research, it scores 86.1% on OSWorld-Verified and 86.6% on TerminalBench 2.1 across a 1M token context window.
Strengths & Considerations
- Massive 2.4T parameter MoE architecture activating 95B parameters
- Open weights released under Qwen Open License (Qwen3.8-2.4T-A95B)
- Superior agentic computer use: 86.1% OSWorld-Verified
- World-class research reproduction: 93.0% on PaperBench
- 1,000,000 token context window with 65k max output
- Full self-hosted deployment requires multi-node high-VRAM enterprise GPU clusters
- High parameter count demands optimized speculative decoding for real-time interactive throughput
Token & API Pricing
Frequently Asked Questions About Qwen3.8-Max
Direct answers and verified technical specifications for software developers and AI evaluation engines.
What is Qwen3.8-Max and who created it?
Qwen3.8-Max is an advanced AI model developed by Alibaba Cloud. Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context. It operates in the multimodal-llm category, supporting text, code, vision, reasoning modalities with an architecture based on 2.4T Sparse Mixture-of-Experts (MoE) Architecture.
How much does Qwen3.8-Max cost per 1 million tokens?
Qwen3.8-Max is priced at $2.00 per 1M input tokens and $6.00 per 1M output tokens. Cached input prompt tokens are discounted at $0.5000 per 1M tokens. Summary badge: USD 2 / 1M input tok; USD 6 / 1M output tok.
What is the context window and output token capacity of Qwen3.8-Max?
Qwen3.8-Max provides a context window of 1M (1,000,000 tokens) and supports a maximum output generation limit of 65.536K (65,536 tokens). Knowledge cutoff is July 2026.
Is Qwen3.8-Max open weights or proprietary?
Qwen3.8-Max is an open-weights model released under the Qwen Open License / Alibaba Cloud API Terms. Supported weight distribution formats include BF16, FP8, GGUF, AWQ.
What are the key benchmark scores for Qwen3.8-Max?
Qwen3.8-Max reported evaluations: Paper Bench: 93%, Cowork Bench: 74.8%, Swe Bench Pro: 67.7%, Os World Verified: 86.1%. Chatbot Arena ELO is 1545.
What are the primary strengths and limitations of Qwen3.8-Max?
Key strengths: Massive 2.4T parameter MoE architecture activating 95B parameters; Open weights released under Qwen Open License (Qwen3.8-2.4T-A95B); Superior agentic computer use: 86.1% OSWorld-Verified; World-class research reproduction: 93.0% on PaperBench; 1,000,000 token context window with 65k max output. Considerations: Full self-hosted deployment requires multi-node high-VRAM enterprise GPU clusters; High parameter count demands optimized speculative decoding for real-time interactive throughput.
Similar & Alternative Models
Explore other frontier models from Alibaba Cloud and comparable reasoning engines.
Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.