arXiv cs.CL· Ian Arawjo·· 4 小时前AI 评分45
如何对 LLM judge 做统计并信任结果:evalstats 的小样本 AI 评估校准推断
How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats
AI 导读
开源 Python 包 evalstats 通过 prediction-powered inference(PPI)为小样本 AI 评估提供九种假设检验,并自动选择校准方法。
来源:arXiv cs.CL · arxiv.org