AGI Hunt· teortaxesTex·· 4 小时前AI 评分39
ExploitBench 成绩高度依赖 harness:GLM 5.3 Flash 表现远超 GLM 4.1 Flash
ExploitBench 成绩高度依赖 harness:GLM 5.3 Flash 表现远超 4.1
AI 导读
GLM 5.3 Flash 在 ExploitBench 上明显强于 GLM 4.1 Flash,成绩高度依赖 harness:Claude Code 中表现最好,Codex 居次。该 harness 并非 ZCode 或 DSH,而是 ExploitBench 默认的约等于 Python while 循环的环境。
来源:AGI Hunt · agihunt.info