跳到正文
原文
Google DeepMind·· 2026-06-09精选AI 评分77

Google DeepMind 发布 Gemma 4 12B 无编码器多模态模型

Introducing Gemma 4 12B: a unified, encoder-free multimodal model

AI 导读

Google DeepMind 发布 Gemma 4 12B,一款无编码器的多模态模型,可在笔记本本地运行。该模型把视觉输入改为单次矩阵乘法的轻量嵌入模块,并去掉音频编码器,将原始音频信号投影到与文本 token 相同的维度,是 Gemma 4 系列首款支持原生音频输入的中等规模模型。

推荐理由

Gemma 4 12B 用无编码器架构让视觉和音频直接进入 LLM,性能接近 26B MoE 而内存占用不到一半。

来源:Google DeepMind · deepmind.google