跳到正文
原文
AGI Hunt· TheZachMueller·· 5 小时前AI 评分34

Prime Intellect 用 NVFP4 压缩 MLA KV 缓存,可多缓存约 50% token

AI 导读

Prime Intellect 发布 Decode 项目的 NVFP4 KV 压缩方案:把 MLA latent 以 NVFP4 存储,每行从 576 字节降到 352 字节,相比 FP8 每个解码器可多缓存约 50% 的 token。

来源:AGI Hunt · agihunt.info