Qwen3.8-Omni-Flash 发布,价格低于 Gemini Flash 且多模态基准相当

The Decoder:AI News(RSS)·2026-09-19 22:30·3小时前·Matthias Bastian
AI 导读

Qwen 发布首个面向 AI 智能体的多模态模型 Qwen3.8-Omni-Flash,可联合处理音频和视频并自主使用工具,上下文窗口达 100 万 token。

The Decoder:AI News(RSS)
68AI 编辑部评分,满分 100

Qwen3.8-Omni-Flash 发布,价格低于 Gemini Flash 且多模态基准相当

2026-09-19 22:30· 3小时前· Matthias Bastian
AI 导读

Qwen 发布首个面向 AI 智能体的多模态模型 Qwen3.8-Omni-Flash,可联合处理音频和视频并自主使用工具,上下文窗口达 100 万 token。

Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents. It processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies. The context window spans one million tokens. On audio-video tasks, Qwen says it comes close to matching Gemini 3.8 Flash.

Qwen 3.8 Omni Flash performs on par with Gemini Flash 3.8 in multimodal benchmarks but is much more affordable. | Image: Qwen

API pricing sits at $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour, while 720p video with audio at one frame per second runs about $0.20, not counting response costs. For comparison, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at its introductory rate, with prices set to double on January 1, 2027.

The model is available through Qwen Studio, Qwen Cloud, and the API. The open-source Qwen-MM-Plugins add video editing, speaker recognition, PDF video notes, and reusable workflows to agents like Claude Code, Gemini CLI, and Qwen Code. Qwen-Live Harness enables real-time interaction using a camera and microphone.