跳到正文
原文
AGI Hunt· ngxson·· 9 小时前AI 评分53

ngxson 实测 RTX 5060 Ti,llama.cpp 跑 Clef Flash 9B 的 prompt 评估比 Ollama 快约 30%

RTX 5060 Ti 实测:llama.cpp 跑 9B 模型 prompt 评估快约 30%

AI 导读

开发者 ngxson 在 RTX 5060 Ti 上以 Q8_0 量化全本地运行 Clef Flash 9B,实测 llama.cpp 的 prompt 评估约 3000 tokens/s,Ollama 约 2100 tokens/s,前者快约 30%。

来源:AGI Hunt · agihunt.info