跳到正文
原文
arXiv cs.CL· Haowei Liu, Hsin-Tai Wu, Yi Fang·· 2 天前AI 评分48

生产级 text-to-SQL 流水线中 LLM-as-judge 失败的审计与修复

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

AI 导读

生产 text-to-SQL 流水线中部署的 gpt-4o-mini LLM-as-judge 与人工标注一致性极低,分歧富集集上 Cohen's kappa 仅 0.04、均匀随机抽查 0.42,并误标 77.1% 的人工判定 FAITHFUL 案例。

来源:arXiv cs.CL · arxiv.org