Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model
Google DeepMind has adapted the Gemma 4 language model into a diffusion model, allowing it to generate text in parallel rather than token by token. By using less than 10 percent of the original training budget, this technique achieves speeds of approximately 1,500 tokens per second. This development demonstrates that existing large language models can be repurposed for faster text generation without requiring entirely new training processes.
Covered by 1 source
- TThe Decoder↗Jonathan KemperAug 9