Mistral AI·· 2025-04-09AI 评分41
Mistral 如何用 LLM as a Judge 评估 RAG 系统
Evaluating RAG with LLM as a Judge
AI 导读
Mistral 展示用 LLM as a Judge 评估 RAG:由 judge LLM 按数值或定性量表为 generator LLM 的答案打分。评判采用 TruLens 的 RAG Triad 三项指标,并通过 Mistral API 近期上线的结构化输出,按自定义 schema 返回可机读的评分与解释。
来源:Mistral AI · mistral.ai