跳到正文
原文
arXiv cs.CL· Ian Arawjo·· 4 小时前AI 评分45

如何对 LLM judge 做统计并信任结果:evalstats 的小样本 AI 评估校准推断

How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats

AI 导读

开源 Python 包 evalstats 通过 prediction-powered inference(PPI)为小样本 AI 评估提供九种假设检验,并自动选择校准方法。

来源:arXiv cs.CL · arxiv.org