AGI Hunt· PawelHuryn·· 4 小时前AI 评分62
Paweł Huryn 自建 Bug Hunt 基准实测新模型,Sonnet 5.5 满分档夺魁
自建Bug Hunt基准实测新模型:Sonnet 5.5夺魁
AI 导读
Paweł Huryn 自建 Bug Hunt Benchmark,用 2 个真实仓库中 105 个 2026 年初前沿模型均未发现的 bug 实测新模型,Sonnet 5.5 拿下满分档夺魁。结果显示 Opus 5.5 性能追平 Fable 5.1 而成本仅为其三分之二,GPT-6.1 Sol 比 GPT-5.6 便宜 10 倍。他还据此核算了各家订阅计划的真实 API 价格。
来源:AGI Hunt · agihunt.info