HuggingFace Blog·· 2026-08-25AI 评分49
Quantization-Aware Healing:GPT-OSS 120B 压缩到 60B 的 4-bit 模型在 7/9 基准上超越全精度版本
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
AI 导读
Quantization-Aware Healing(QAH)让 GPT-OSS 120B 压缩到 60B、量化为 MXFP4 后,在 9 项基准中 7 项超过其 bfloat16 全精度版本。
来源:HuggingFace Blog · huggingface.co