arXiv cs.CL· Bhavik Mangla·· 5 小时前AI 评分60
VoxParity 基准测试语音智能体是否依据音频线索决策,23 个系统中仅 11 个通过
Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
AI 导读
论文提出 VoxParity 基准,用同一份文本配不同音频的方式检验语音智能体是否据此改变工具调用,在 23 个可同时跑文本管线的系统中仅 11 个通过。
来源:arXiv cs.CL · arxiv.org