本周最佳, Week 37
没有 AI slop, 唯一的 DeepSeek-V4.1-Flash: Causal Encoder-Decoder (CED)
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
DeepSeek 发布 DeepSeek-V4.1-Flash,一款 552B 参数(每 token 激活 16B/预填充 8B)的多模态 MoE 模型,支持最长一百万 token 上下文。其 Causal Encoder-Decoder(CED)架构结合 CSA2 与 FP4 KV caching,将全局 KV cache 降至每 token 890 字节,约为 DeepSeek-V4-Flash 的 1/4、DeepSeek-V1 的 1/437,配合 SWA Bounded Replay 将持久 KV cache 降至约 1/8,性能仍优于基线。模型基于 45T token 多模态语料预训练,checkpoint 已在 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 开放。
DeepSeek 发布 DeepSeek-V4.1-Flash,一款 552B 参数(每 token 激活 16B/预填充 8B)的多模态 MoE 模型,支持最长一百万 token 上下文。其 Causal Encoder-Decoder(CED)架构结合 CSA2 与 FP4 KV caching,将全局 KV cache 降至每 token 890 字节,约为 DeepSeek-V4-Flash 的 1/4、DeepSeek-V1 的 1/437,配合 SWA Bounded Replay 将持久 KV cache 降至约 1/8,性能仍优于基线。模型基于 45T token 多模态语料预训练,checkpoint 已在 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 开放。
本周最佳, Week 37
没有 AI slop, 唯一的 DeepSeek-V4.1-Flash: Causal Encoder-Decoder (CED)
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
来源:Dongxi 东锡 NLP· x.com