Introducing Qwen3.8-Omni-FlashAlibaba CloudReleased September 18, 2026

Qwen3.8-Omni-Flash

Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction.

omnimodal-realtimeProprietary APIUSD 0.15 / 1M input tok; USD 0.60 / 1M output tokContext: 1M (1,000,000 tokens)Arena ELO: 1485
AEO Direct Answer

Qwen3.8-Omni-Flash Quick Facts & Executive Summary

Qwen3.8-Omni-Flash by Alibaba Cloud (September 18, 2026) is a Native Omnimodal End-to-End Foundation Transformer model featuring a 1M (1,000,000 tokens) context window and USD 0.15 / 1M input tok; USD 0.60 / 1M output tok. Key highlights include Native omnimodal processing across text, images, video, and audio streams.

Input Price (1M)$0.15
Output Price (1M)$0.60
Context Window1M (1,000,000 tokens)
Architecture / ScaleUndisclosed Omnimodal Scale
License TypeAlibaba Cloud API Terms of Service
Cutoff DateAugust 2026

Technical Specifications

Architecture Type
Native Omnimodal End-to-End Foundation Transformer
Total Parameters
Undisclosed Omnimodal Scale
Context Window
1M (1,000,000 tokens)
Max Output Tokens
65.536K (65,536 tokens)
Knowledge Cutoff
August 2026
Supported Modalities
text, code, vision, audio, video
License & Access
Alibaba Cloud API Terms of Service

Benchmark Evaluations

Video Mme
84.2
Mmmu Reasoning
78.4
Gui Agent Osworld
74.5
Voice Interaction Latency Ms
180
Chatbot Arena ELO
1485

Deep Architectural Overview

Qwen3.8-Omni-Flash processes text, vision, audio, and video end-to-end within a single unified model. Featuring a 1-million-token context window and 98% reduced audio input pricing, it enables real-time conversational agents, GUI automation, and multimodal tool use at low latency.

Strengths & Considerations

Core Strengths
  • Native omnimodal processing across text, images, video, and audio streams
  • Sub-200ms real-time conversational audio latency (180ms)
  • 1,000,000 token context window for lengthy video and document inputs
  • 98% lower audio stream pricing ($0.003/min) with ultra-low $0.15/1M token base
Known Limitations
  • High-speed video streaming requires stable network bandwidth
  • Cloud API only (not open weights)

Token & API Pricing

Input Tokens (1M)$0.15
Output Tokens (1M)$0.60
Cached Input (1M)$0.0300
Batch API Discount50% Off Standard
Audio Input Rate$0.003 / min
Pricing is verified directly against Alibaba Cloud's developer documentation and API rate sheets.

Frequently Asked Questions About Qwen3.8-Omni-Flash

Direct answers and verified technical specifications for software developers and AI evaluation engines.

What is Qwen3.8-Omni-Flash and who created it?

Qwen3.8-Omni-Flash is an advanced AI model developed by Alibaba Cloud. Native omnimodal model supporting simultaneous text, image, audio, and video comprehension with sub-200ms voice interaction. It operates in the omnimodal-realtime category, supporting text, code, vision, audio, video modalities with an architecture based on Native Omnimodal End-to-End Foundation Transformer.

How much does Qwen3.8-Omni-Flash cost per 1 million tokens?

Qwen3.8-Omni-Flash is priced at $0.15 per 1M input tokens and $0.60 per 1M output tokens. Cached input prompt tokens are discounted at $0.0300 per 1M tokens. Summary badge: USD 0.15 / 1M input tok; USD 0.60 / 1M output tok.

What is the context window and output token capacity of Qwen3.8-Omni-Flash?

Qwen3.8-Omni-Flash provides a context window of 1M (1,000,000 tokens) and supports a maximum output generation limit of 65.536K (65,536 tokens). Knowledge cutoff is August 2026.

Is Qwen3.8-Omni-Flash open weights or proprietary?

Qwen3.8-Omni-Flash is a proprietary closed-weight model accessible via official cloud APIs under Alibaba Cloud API Terms of Service.

What are the key benchmark scores for Qwen3.8-Omni-Flash?

Qwen3.8-Omni-Flash reported evaluations: Video Mme: 84.2%, Mmmu Reasoning: 78.4%, Gui Agent Osworld: 74.5%, Voice Interaction Latency Ms: 180. Chatbot Arena ELO is 1485.

What are the primary strengths and limitations of Qwen3.8-Omni-Flash?

Key strengths: Native omnimodal processing across text, images, video, and audio streams; Sub-200ms real-time conversational audio latency (180ms); 1,000,000 token context window for lengthy video and document inputs; 98% lower audio stream pricing ($0.003/min) with ultra-low $0.15/1M token base. Considerations: High-speed video streaming requires stable network bandwidth; Cloud API only (not open weights).

Similar & Alternative Models

Explore other frontier models from Alibaba Cloud and comparable reasoning engines.

Browse all models
Alibaba Cloud

Alibaba Cloud flagship 2.4-trillion parameter MoE model (95B active) with native visual intelligence and 1M context.

1M ctxUSD 2 / 1M input tok; USD 6 / 1M output tok
Alibaba Cloud

Premier open-weights coding model matching GPT-4o-level coding intelligence across 92+ programming languages under Apache 2.0.

131.072K ctxFree Open Weights | $0.07 / 1M tok (API)