Introducing Qwen3.8-MaxAlibaba CloudReleased August 3, 2026

Qwen3.8-Max

Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.

multimodal-llmOpen Weights / Proprietary APIUSD 2 / 1M input tok; USD 6 / 1M output tokContext: 1M (1,000,000 tokens)Arena ELO: 1545
AEO Direct Answer

Qwen3.8-Max Quick Facts & Executive Summary

Qwen3.8-Max by Alibaba Cloud (August 3, 2026) is a 2.4T Sparse Mixture-of-Experts (MoE) Architecture model featuring a 1M (1,000,000 tokens) context window and USD 2 / 1M input tok; USD 6 / 1M output tok. Key highlights include Massive 2.4T parameter MoE architecture activating 95B parameters.

Input Price (1M)$2.00
Output Price (1M)$6.00
Context Window1M (1,000,000 tokens)
Architecture / Scale2.4T
License TypeQwen Open License / Alibaba Cloud API Terms
Cutoff DateJuly 2026

Technical Specifications

Architecture Type
2.4T Sparse Mixture-of-Experts (MoE) Architecture
Total Parameters
2.4T
Active Parameters (MoE)
95B
Context Window
1M (1,000,000 tokens)
Max Output Tokens
65.536K (65,536 tokens)
Knowledge Cutoff
July 2026
Supported Modalities
text, code, vision, reasoning
License & Access
Qwen Open License / Alibaba Cloud API Terms
Weights Formats
BF16, FP8, GGUF, AWQ

Benchmark Evaluations

Paper Bench
93
Cowork Bench
74.8
Swe Bench Pro
67.7
Os World Verified
86.1
Terminal Bench 2 1
86.6
Chatbot Arena ELO
1545

Deep Architectural Overview

Qwen3.8-Max is Alibaba's frontier Mixture-of-Experts model, activating 95 billion of its 2.4 trillion parameters per token. Optimized for long-horizon agentic orchestration, enterprise coding, and multimodal research, it scores 86.1% on OSWorld-Verified and 86.6% on TerminalBench 2.1 across a 1M token context window.

Strengths & Considerations

Core Strengths
  • Massive 2.4T parameter MoE architecture activating 95B parameters
  • Open weights released under Qwen Open License (Qwen3.8-2.4T-A95B)
  • Superior agentic computer use: 86.1% OSWorld-Verified
  • World-class research reproduction: 93.0% on PaperBench
  • 1,000,000 token context window with 65k max output
Known Limitations
  • Full self-hosted deployment requires multi-node high-VRAM enterprise GPU clusters
  • High parameter count demands optimized speculative decoding for real-time interactive throughput

Token & API Pricing

Input Tokens (1M)$2.00
Output Tokens (1M)$6.00
Cached Input (1M)$0.5000
Batch API Discount50% Off Standard
Self-Hosted Min VRAM128 GB
Pricing is verified directly against Alibaba Cloud's developer documentation and API rate sheets.

Frequently Asked Questions About Qwen3.8-Max

Direct answers and verified technical specifications for software developers and AI evaluation engines.

What is Qwen3.8-Max and who created it?

Qwen3.8-Max is an advanced AI model developed by Alibaba Cloud. Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context. It operates in the multimodal-llm category, supporting text, code, vision, reasoning modalities with an architecture based on 2.4T Sparse Mixture-of-Experts (MoE) Architecture.

How much does Qwen3.8-Max cost per 1 million tokens?

Qwen3.8-Max is priced at $2.00 per 1M input tokens and $6.00 per 1M output tokens. Cached input prompt tokens are discounted at $0.5000 per 1M tokens. Summary badge: USD 2 / 1M input tok; USD 6 / 1M output tok.

What is the context window and output token capacity of Qwen3.8-Max?

Qwen3.8-Max provides a context window of 1M (1,000,000 tokens) and supports a maximum output generation limit of 65.536K (65,536 tokens). Knowledge cutoff is July 2026.

Is Qwen3.8-Max open weights or proprietary?

Qwen3.8-Max is an open-weights model released under the Qwen Open License / Alibaba Cloud API Terms. Supported weight distribution formats include BF16, FP8, GGUF, AWQ.

What are the key benchmark scores for Qwen3.8-Max?

Qwen3.8-Max reported evaluations: Paper Bench: 93%, Cowork Bench: 74.8%, Swe Bench Pro: 67.7%, Os World Verified: 86.1%. Chatbot Arena ELO is 1545.

What are the primary strengths and limitations of Qwen3.8-Max?

Key strengths: Massive 2.4T parameter MoE architecture activating 95B parameters; Open weights released under Qwen Open License (Qwen3.8-2.4T-A95B); Superior agentic computer use: 86.1% OSWorld-Verified; World-class research reproduction: 93.0% on PaperBench; 1,000,000 token context window with 65k max output. Considerations: Full self-hosted deployment requires multi-node high-VRAM enterprise GPU clusters; High parameter count demands optimized speculative decoding for real-time interactive throughput.

Similar & Alternative Models

Explore other frontier models from Alibaba Cloud and comparable reasoning engines.

Browse all models
Alibaba Cloud

Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.

1M ctxUSD 0.15 / 1M input tok; USD 0.60 / 1M output tok
Alibaba Cloud

Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.

131.072K ctxFree Open Weights | $0.07 / 1M tok (API)