跳到正文
原文
arXiv cs.CV· Zhixi Zhu, Kristina Gligoric·· 4 小时前AI 评分46

为过多 Token 付费?用简单启发式实现有效且低成本的多模态 LLM 标注

Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

AI 导读

多模态 LLM 视频标注的系统评估发现,用简单镜头切换检测构建单个 2×8 图像网格,即可接近全视频理解(κ 差异在 .05 内),token 成本仅约 15%。

来源:arXiv cs.CV · arxiv.org