跳到正文
原文
AGI Hunt· PyTorch·· 3 小时前AI 评分36

PyTorch 用 Triton 替换 CUDA,推荐系统 kernel 前向提速 1.28 倍

AI 导读

Meta 团队将推荐系统核心的 Table Batched Embedding(TBE)kernel 从 CUDA 迁移到 Triton(FBTriton),前向传播最高提速 1.28x,反向传播最高 2x。

来源:AGI Hunt · agihunt.info