DiffusionGemma Technical Report
谷歌的DiffusionGemma是一款实验性开放权重模型,通过并行打磨256个token组成的区块,让文本生成速度远超逐词生成的传统AR模型
与传统自回归(AR)模型一次生成一个token不同,DiffusionGemma使用离散扩散技术,一次性并行迭代打磨256个token组成的区块来生成文本。它并非从零训练,而是在混合专家(MoE)架构的Gemma 4模型基础上微调而成,该模型激活参数38亿、总参数252亿,微调仅使用了原AR模型训练token预算的不到10%。在完整评测集上平均,它每次前向传播生成约20个token,在单张H100 GPU上可达到约每秒1500个输出token,比使用最先进推测解码的AR模型还要快得多。
METAL LAB 解读图
DiffusionGemma两阶段流水线如何实现快速文本生成
证据状态已报告实测结果
- 起点:Gemma 4 AR模型从预训练的Gemma 4 26B A4B MoE权重初始化(激活参数38亿,总参数252亿),而非从零训练
- 第一阶段:SFT监督微调教模型在256-token画布上学习双向去噪,仅使用原AR模型训练token预算的不到10%
- 第二阶段:SD·RL采样器蒸馏结合强化学习,同时提升基于奖励的生成质量并压缩所需的去噪步数
- 推理:区块式扩散解码从随机噪声画布出发,熵约束采样器配合自适应停止平均约12步完成一个256-token区块,再拼接进KV缓存
- 结果:全新的速度-质量帕累托前沿单张H100上每秒约1500个token、每次前向传播约20个token,超越了配备最先进推测解码的AR模型
他们做了什么
- 传统AR模型一次只能生成一个token,在请求量较低时GPU算力被浪费,因为时间主要花在把模型权重和KV缓存从显存搬运到计算单元上(受限于显存带宽)。
- DiffusionGemma用离散扩散绕开了这个瓶颈:从一个由随机噪声组成的256-token画布出发,用双向注意力同时并行地逐步打磨整个区块。
- 训练分两个阶段:第一阶段用监督微调(SFT)教模型学会双向去噪;第二阶段称为SD·RL,结合采样器蒸馏与强化学习,同时提升生成质量并压缩所需的去噪步数。
- 熵约束采样器搭配自适应停止机制,能在模型预测足够自信且稳定时提前结束去噪,最多可将延迟降低约4倍而不牺牲质量。
- 由于与Gemma 4共享完全相同的transformer架构,微调后的权重仍可切换回标准自回归模式运行,为扩散与AR的混合解码留下了空间。

| Total | 25.2B |
|---|---|
| Activated | 3.85B |
| Vision Encoder | 550M |
| Embedder | 740M |
| Self-Conditioning | 7.8M |
| Active / Total Experts | 8 / 128 |
| + 1 shared |

| Open-weight Models | Closed-weight Model | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| DiffusionGemma | Gemma 4 | LLaDA 2.1 Flash | Nemotron Diffusion | Mercury 2 | ||||||
| 26B A4B | 26B A4B | 100B | 14B | Unknown | ||||||
| Mode | TD | TD (No-think) | AR | AR (No-think) | AR (MTP) | AR (MTP, No-think) | TD (S Mode) | TD (Diffusion Mode) | High | Medium |
| AIME 2026 | 69.1 | 50.8 | 84.2 | 57.5 | 88.3 | 80.0 | 80.0 | 40.0 | 91.7 | 82.5 |
| GPQA Diamond | 73.2 | 64.6 | 79.8 | 67.2 | 82.3 | 73.7 | 68.7 | 47.0 | 75.2 | 66.7 |
| LiveCodeBench-V6 | 69.1 | 60.6 | 71.4 | 58.3 | 77.1 | 72.6 | 39.4 | 28.6 | 79.4 | 74.9 |
| Codeforces ELO | 1429 | 959 | 1569 | 1059 | 1718 | 1529 | 718 | - | 1986 | 1629 |
| BigBench EH | 47.6 | 40.0 | 59.1 | 42.2 | 64.8 | 56.2 | - | - | 48.9 | 43.8 |
| GSM8K | 96.3 | 95.8 | 96.6 | 96.1 | 96.7 | 96.4 | 45.0 | - | 96.5 | 95.8 |
| MGSM | 84.8 | 80.7 | 87.9 | 84.3 | 92.9 | 91.5 | 6.8 | 69.3 | 91.9 | 91.2 |
| MMMLU | 81.5 | 76.3 | 82.2 | 78.0 | 86.3 | 78.0 | - | - | 81.9 | 80.6 |
| MMMU Pro | 54.3 | 66.0 | 63.3 | 66.7 | 73.8 | 72.5 | - | - | - | - |
| Putnam | 67.4 | 57.1 | 74.7 | 59.7 | 81.0 | 72.9 | - | 45.8 | 73.6 | 73.6 |
| HumanEval | 94.5 | 92.7 | 98.2 | 97.6 | 98.8 | 97.6 | 90.2 | 86.0 | 98.2 | 98.2 |
| BigCodeBench | 46.0 | 41.9 | 47.7 | 45.9 | 50.2 | 48.1 | - | 33.5 | 47.6 | 45.3 |
| LBPP | 81.0 | 68.9 | 86.3 | 74.1 | 89.5 | 77.3 | 45.7 | 40.7 | 89.2 | 85.0 |
| IFEval | 97.4 | 94.5 | 97.2 | 95.7 | 98.7 | 97.8 | - | 72.1 | 97.0 | 94.5 |
| Tau2 Retail | 71.5 | 57.5 | 75.4 | 61.0 | 85.5 | 79.0 | - | - | - | - |
| Tau2 Airline | 69.0 | 49.0 | 72.0 | 50.0 | 76.0 | 51.0 | - | - | - | - |
| Tau2 Telecom | 28.1 | 32.0 | 33.8 | 32.0 | 43.0 | 34.2 | - | - | - | - |
| MMLU-Pro | 77.6 | 77.9 | 78.8 | 79.1 | 82.6 | 82.6 | - | - | 77.6 | 75.5 |
| Natural2Code | 94.0 | 90.1 | 96.2 | 92.3 | 96.3 | 94.7 | 86.9 | 73.3 | 79.1 | 71.3 |
| HiddenMath | 80.6 | 74.3 | 85.4 | 77.5 | 87.2 | 81.6 | - | 44.3 | 82.7 | 82.3 |
| Output Speed (TPS) | 1479 | 1512 | 204 | 204 | 303 | 303 | 375 | 49 | 600 | 547 |
| Tokens Per Forward (TPF) | 19.74 | 18.76 | 1.00 | 1.00 | 1.40 | 1.40 | 4.63 | 1.79 | - | - |
| Average Total Tokens | 4,001 | 829 | 5,184 | 1,025 | 7,207 | 1,816 | 4,371 | 941 | 3,882 | 1,222 |


| Score (↑) | TPF (↑) | TPS (↑) | Effective DNS (↓) | Total Forwards (↓) | Total Tokens | E2E Time (s) (↓) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Think | No-Think | Think | No-Think | Think | No-Think | Think | No-Think | Think | No-Think | Think | No-Think | Think | No-Think |
| AIME 2026 | 69.1 | 50.8 | 19.3 | 16.7 | 1365.4 | 1333.0 | 12.6 | 14.1 | 390.6 | 91.1 | 6,445 | 1,309 | 4.72 | 0.98 |
| GPQA Diamond | 73.2 | 64.6 | 16.7 | 16.5 | 1207.8 | 1330.2 | 15.1 | 13.4 | 443.4 | 48.1 | 5,647 | 726 | 4.68 | 0.55 |
| LiveCodeBench-V6 | 69.1 | 60.6 | 18.5 | 16.9 | 1278.3 | 1333.4 | 13.8 | 14.0 | 581.8 | 195.7 | 7,534 | 1,847 | 5.89 | 1.39 |
| Codeforces ELO | 1429 | 959 | 15.1 | 14.0 | 950.5 | 1040.3 | 17.1 | 18.5 | 959.6 | 521.1 | 11,622 | 4,279 | 12.23 | 4.11 |
| BigBench EH | 47.6 | 40.0 | 20.8 | 17.7 | 1390.2 | 1415.1 | 11.9 | 12.8 | 434.9 | 68.1 | 9,062 | 1,233 | 6.52 | 0.87 |
| GSM8K | 96.3 | 95.8 | 23.2 | 24.1 | 1866.2 | 1966.4 | 9.1 | 7.5 | 43.4 | 13.2 | 883 | 298 | 0.47 | 0.15 |
| MGSM | 84.8 | 80.7 | 19.0 | 16.8 | 1526.7 | 1367.5 | 11.5 | 11.9 | 63.6 | 19.5 | 1,085 | 291 | 0.71 | 0.21 |
| MMMU Pro | 54.3 | 66.0 | 17.7 | 15.4 | 1351.3 | 1255.1 | 13.8 | 15.3 | 191.6 | 31.6 | 3,178 | 472 | 2.35 | 0.38 |
| Putnam | 67.4 | 57.1 | 18.1 | 16.0 | 1330.8 | 1282.4 | 13.4 | 14.4 | 303.8 | 77.4 | 4,725 | 1,103 | 3.55 | 0.86 |
| HumanEval | 94.5 | 92.7 | 23.0 | 24.3 | 1838.2 | 1981.2 | 9.4 | 8.0 | 55.6 | 14.3 | 1,174 | 305 | 0.64 | 0.15 |
| BigCodeBench | 46.0 | 41.9 | 19.6 | 19.4 | 1560.1 | 1579.0 | 11.3 | 10.1 | 77.3 | 21.8 | 1,410 | 394 | 0.90 | 0.25 |
| LBPP | 81.0 | 68.9 | 20.5 | 19.2 | 1509.8 | 1545.4 | 11.6 | 10.9 | 264.3 | 79.4 | 4,730 | 859 | 3.13 | 0.56 |
| IFEval | 97.4 | 94.5 | 17.2 | 9.0 | 1368.3 | 732.1 | 13.0 | 14.4 | 100.8 | 24.0 | 1,464 | 239 | 1.07 | 0.33 |
| Natural2Code | 94.0 | 90.1 | 21.1 | 21.0 | 1682.7 | 1706.5 | 10.5 | 9.6 | 70.5 | 24.7 | 1,391 | 465 | 0.83 | 0.27 |
| HiddenMath | 80.6 | 74.3 | 21.0 | 19.4 | 1591.2 | 1564.7 | 11.3 | 11.5 | 206.4 | 49.7 | 3,501 | 844 | 2.20 | 0.54 |


| Model | Effective Denoising Steps | Accuracy (%) | BLEU |
|---|---|---|---|
| DiffusionGemma | 18.09 | 75.6 | 10.76 |
| + LoRA finetuning | 31.57 | 76.62 | 20.67 |


| Hyperparameter | Sudoku (LoRA) | Sudoku (Full) | PubMedQA |
|---|---|---|---|
| LoRA rank | 8 | — | 4 |
| Canvas size | 256 | 256 | 128 |
| Number of canvases | 1 | 1 | 2 |
| Prompt length | 256 | 256 | 1024 |
| Batch size | 2 | 8 | 2 |
| Peak learning rate | 3×10−4 | 1.125×10−4 | 1.0×10−4 |
| End learning rate | 3×10−5 | 1.125×10−5 | 1.0×10−5 |
| Training steps | 8,000 | 2,000 | 2,000 |
| Optimizer | Adam | Adafactor | Adam |
| LR schedule | Cosine with warmup | Cosine with warmup | Cosine with warmup |
| Warmup iterations | 400 | 100 | 100 |
| Weight decay | 10−4 | 10−4 | 10−4 |
| Min. hardware | 2× A100 80GB | 8× A100 80GB | 2× A100 80GB |


研究结果
- 在完整评测集上平均,DiffusionGemma每次前向传播生成约20个token,在单张NVIDIA H100 GPU上达到约每秒1500个输出token(TPS)。
- 这比使用最先进推测解码的AR模型(通常每次前向传播仅生成约3到6个token)快得多。
- 启用自适应停止后,模型在最大48步预算中平均只需约12个有效去噪步,延迟降低约4倍且不牺牲生成质量。
- 尽管每步处理256个token,但每步延迟相比单token的AR生成仅增加3.2倍(在H100、FP8精度、批大小为1的条件下测得)。
- 微调后的权重仍可用于标准AR生成,性能仅有轻微下降。
可应用场景
- 为对响应速度敏感的聊天机器人或代码助手提供低延迟服务
- 在并发请求量较低(小批大小)、GPU算力容易被浪费的推理场景中使用
- 借助Apache 2.0的宽松许可,针对特定领域(如医疗问答、语音识别)进行轻量级微调
- 根据延迟约束和任务复杂度,在扩散解码与AR解码之间动态路由请求的混合服务方案
局限与待验证事项
- 当并发请求量达到中等规模(约32个并发请求)时,AR模型在吞吐量上会重新占据优势。
- 扩散微调相比原始AR基线会带来一定的性能下降。
- 对Mercury 2等闭源竞品的速度评估是通过OpenRouter API间接估算得出,而非直接测量,存在一定误差。
- 论文正文中间大段内容及部分附录(如附录F、G.3)在本次材料中被省略,未能查看全部细节结果。
- 由于离散扩散在更新每个位置时无法看到其他位置的同步决定,偶尔会出现局部语法不一致的问题。
为什么重要
对于聊天机器人、代码助手等对响应速度敏感的服务来说,这意味着单张GPU就有可能实现比当前AR服务方案快得多的文本生成。该模型以Apache 2.0协议开放权重发布,研究者和开发者可以直接检视离散扩散的运作机制,并以较低算力成本针对自己的场景进行微调。
本文术语
- 离散扩散(discrete diffusion) · 直接在词表中的实际token状态之间添加和去除噪声的生成方法,不同于图像常用的连续空间扩散
- 自回归(AR)模型 · 从左到右逐个生成token的传统语言模型方式
- 混合专家(MoE) · 每次输入只激活部分专门子网络以降低计算量的模型架构
- 自适应停止(adaptive stopping) · 当模型预测足够自信且稳定时提前结束去噪过程,而不是每次都跑满最大步数
- 采样器蒸馏(sampler distillation) · 训练模型用更少的去噪步数就能复现高质量但缓慢的生成结果的技术
论文原文摘要(英文)
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI助手在该向你提问的时候,却更愿意自己去核实事实
- SynFlow: A Multidimensional Diachronic Semantic Analysis Toolkit一款把单词意义变化拆解到语法细节的开源分析工具
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)用AI总结股市新闻发现:简单的摘要方法反而比时髦的检索增强技术更靠谱
METAL LAB 最新报道
图片来源: DiffusionGemma Team et al., arXiv:2608.00146, CC BY 4.0