Google DeepMind·· 2026-06-11精选AI 评分82
Google DeepMind 发布 DiffusionGemma 文本扩散模型,推理速度最高提升 4 倍
DiffusionGemma: 4x faster text generation
AI 导读
Google DeepMind 发布实验性开放模型 DiffusionGemma,采用文本扩散和并行生成,在专用 GPU 上实现最高 4 倍的文本生成速度。该 26B MoE 模型推理时激活 3.8B 参数,量化后可适配 18GB VRAM,并以 Apache 2.0 许可发布。DiffusionGemma 面向本地、低并发的交互式工作流,但整体输出质量低于标准 Gemma 4。
推荐理由
文章具体说明了 DiffusionGemma 在本地低并发推理中的速度优势,也交代了输出质量与标准 Gemma 4 的取舍。
来源:Google DeepMind · deepmind.google