Qwen 团队发布 Qwen3.8-Omni-Flash 报告:面向原生全模态智能体

DAIR.AI · @dair_ai · X·2026-09-26 03:53·2小时前
AI 导读

Qwen 团队发布报告,介绍 Qwen3.8-Omni-Flash,一个原生多模态模型,面向跨文本、音频和视频的长程智能体任务,如视频编辑和长音频视频翻译。模型采用 Qwen3.8-Next 的稀疏 MoE 架构,上下文窗口达一百万 token,并通过协同训练策略在保持文本性能的同时将智能体能力迁移到音频和视频任务。

DAIR.AI@dair_ai
70AI 编辑部评分,满分 100

Qwen 团队发布 Qwen3.8-Omni-Flash 报告:面向原生全模态智能体

2026-09-26 03:53· 2小时前
AI 导读

Qwen 团队发布报告,介绍 Qwen3.8-Omni-Flash,一个原生多模态模型,面向跨文本、音频和视频的长程智能体任务,如视频编辑和长音频视频翻译。模型采用 Qwen3.8-Next 的稀疏 MoE 架构,上下文窗口达一百万 token,并通过协同训练策略在保持文本性能的同时将智能体能力迁移到音频和视频任务。

The era of omni agents is upon us.

This is a great report by the Qwen Team on their omni-modal agents.

They present Qwen3.8-Omni-Flash, a natively multimodal model trained for long-horizon agent tasks across text, audio, and video, such as video editing and long-form audio and video translation.

It uses the sparse mixture-of-experts design of Qwen3.8-Next with a context window of one million tokens. A co-training strategy keeps text performance while carrying agent skills over to audio and video tasks.

Two open-source frameworks come with it. Qwen-MM-Plugins adds audio and video support to existing agent harnesses, and Qwen-Live-Harness handles real-time multimodal interaction with context and memory management, tool use, and sub-agent delegation.

Paper: https://academy.dair.ai/papers/qwen3-8-omni-towards-native-omni-modal-agents-2609.25611