AGI Hunt· NationalUniversityofSingapore·· 12 小时前AI 评分39
QuantWM:免训练 2-bit KV cache 量化,视频世界模型压缩 6.2 倍
QuantWM:免训练 2-bit KV cache 量化,视频世界模型压缩 6.2 倍
AI 导读
新加坡国立大学提出免训练量化框架 QuantWM,视频世界模型 KV cache 最高压缩 6.20 倍。现有 2-bit KV cache 量化在 VBench 上近乎无损,但用于视频世界模型会引发严重时序闪烁,因为 Key 扰动会改变 attention logits 并挪动 Query 选中的时空 token。
来源:AGI Hunt · agihunt.info