elvis · @omarsar0 · X·2026-08-26 23:23·27天前
AI 导读

阿里发布Qwen3.8-Flash,一款多模态MoE模型,也是Qwen4架构的早期预览,现已开源权重。该模型总参数量125B,每token仅激活6B,原生上下文262K,可扩展至1M。

elvis@omarsar0
49AI 编辑部评分,满分 100
2026-08-26 23:23· 27天前
AI 导读

阿里发布Qwen3.8-Flash,一款多模态MoE模型,也是Qwen4架构的早期预览,现已开源权重。该模型总参数量125B,每token仅激活6B,原生上下文262K,可扩展至1M。

What's better than an open-weight multimodal model release?

Well, the technical report. I just love how these labs like Qwen and DeepSeek continue to drop gem after gem.

Qwen3.8-Flash is the latest in efficient multimodal MoE models. Worth reading the report.

Qwen⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available so...