模型 信源:X:马东锡 NLP (@dongxi_nlp) · 👁️ 13358 次研读

DeepSeek 发布 DeepSeek-V4.1-Flash:以 CED 架构压缩 KV cache

💡 灵机 AI 深度洞见与核心提炼
DeepSeek 发布 DeepSeek-V4.1-Flash,一款 552B 参数(每 token 激活 16B/预填充 8B)的多模态 MoE 模型,支持最长一百万 token 上下文。其 Causal Encoder-Decoder(CED)架构结合 CSA2 与 FP4 KV caching,将全局 KV cache 降至每 token 890 字节,约为 DeepSeek-V4-Flash 的 1/4、DeepSeek-V1 的 1/437,配合 SWA Bounded Replay 将持久 KV cache 降至约 1/8,性能仍优于基线。模型基于 45T token 多模态语料预训练,checkpoint 已在 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 开放。
信源媒体:X:马东锡 NLP (@dongxi_nlp)
访问出处网页 ↗
阅读原文出处 ↗