Skip to main content
Back to News Hub
💡Google DeepMind
June 10, 2026
General AI

DiffusionGemma: 4x faster text generation

Overview

Google DeepMind released DiffusionGemma, an experimental open-weights model that writes text through a diffusion process instead of one word at a time, running up to 4x faster than standard models on GPUs. The model is a 26B Mixture of Experts design that activates only about 3.8B parameters per step and generates 256 tokens in parallel. It reaches more than 1,000 tokens per second on a single NVIDIA H100 and ships free to use under the Apache 2.0 license.

Stats & Key Facts

  • #Google DeepMind released DiffusionGemma, an experimental open-weights model that writes text through a diffusion process instead of one word at a time, running up to 4x faster than standard models on GPUs.
  • #The model is a 26B Mixture of Experts design that activates only about 3.8B parameters per step and generates 256 tokens in parallel.
  • #It reaches more than 1,000 tokens per second on a single NVIDIA H100 and ships free to use under the Apache 2.0 license.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by Google DeepMind
Read the original