An overview of DiffusionGemma, an exceptionally fast text generation model with up to 4x faster speeds.
Our newest open experimental model delivers up to 4x faster inference on dedicated GPUs and opens the door to exploring speed-critical, interactive local workflows.
Brendan O'Donoghue
Research Scientist
Sebastian Flennerhag
Research Scientist
Today, we’re introducing DiffusionGemma, an experimental open model that explores text diffusion, an exceptionally fast approach to text generation. Released under an Apache 2.0 license, this 26B Mixture of Experts (MoE) model moves beyond the sequential token-by-token processing of typical autoregressive Large Language Models (LLMs). Instead, it generates entire blocks of text simultaneously, delivering up to 4x faster text generation on GPUs.
Built upon the industry-leading intelligence-per-parameter of our Gemma 4 family and cutting-edge Gemini Diffusion research, DiffusionGemma integrates a novel diffusion head designed to maximize generation speed. While autoregressive Gemma 4 models remain the standard for high-quality production outputs, DiffusionGemma is designed for researchers and developers exploring speed-critical, interactive local workflows such as in-line editing, rapid iteration, and generating non-linear text structures.
Unlocking new value for developersDevelopers building real-time interactive AI applications often struggle with the latency bottlenecks of local inference. DiffusionGemma addresses these challenges directly, with some key trade-offs:
You can improve DiffusionGemma's performance on specific tasks through fine-tuning. In the example below, Unsloth fine-tuned DiffusionGemma to play Sudoku — a task autoregressive models struggle with because each token depends on future tokens. DiffusionGemma's bi-directional attention makes this much easier.
Fine-tuned DiffusionGemma solving Sudoku.
While the AI research community has explored diffusion-based text generation for years, applying it to large models has remained a challenge. DiffusionGemma changes this by shifting how models use hardware.
The trade-off with traditional modelsMost language models act like a typewriter, generating one token at a time from left to right. In the cloud, this is efficient because servers can batch thousands of user requests together to share the hardware load. But when run locally for a single user, this word-by-word process leaves your dedicated GPU or TPU underutilized — it spends most of its time simply waiting for the next "keystroke."
DiffusionGemma reverses this inefficiency. Instead of predicting words sequentially, it drafts an entire 256-token paragraph simultaneously. By giving the computer's processor a larger chunk of work at once, DiffusionGemma utilizes your hardware to its full potential. It upgrades your model inference from a single, sequential typewriter to a massive printing press that stamps the entire block of text simultaneously.
DiffusionGemma text-to-3D SVG demo by Hugging Face. Step-by-step generation.
This means DiffusionGemma's speedup is designed for local and low-concurrency inference. In high-QPS cloud serving, autoregressive models can be deployed to saturate compute efficiently, so DiffusionGemma's parallel decoding offers diminishing returns and can result in higher serving costs. The throughput advantage is strongest at low-to-medium batch sizes on a single accelerator.
How text diffusion worksSimilar to AI image generators that start with visual static and iteratively refine it into a clear picture, DiffusionGemma applies this to text:
Because the model can process the whole paragraph while generating, it unlocks new patterns of model behavior, like perfectly closing complex markdown formatting or generating and rendering code in near real-time.
Get started today| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Introducing Gemma 4 12B: a unified, encoder-free multimodal model | 0 | 19.02 | 03-06-2026 |
| 2 | See what 3 builders are making with Gemma 4 | 0 | 19.07 | 09-06-2026 |
| 3 | Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency | 0 | 24.36 | 05-06-2026 |
| 4 | Эксперимент по подстройке Gemma 3 для вызова процедур | 0 | 21.07 | 13-01-2026 |
| 5 | Neue KI-Modelle: Gemini 3.7 Flash und DeepSeek V4-Pro im Vergleich | 0 | 17.75 | 13-08-2026 |
| 6 | Comparatif Google Gemini vs ChatGPT : quelle est la meilleure IA en marketing ? | 0 | 10.76 | 24-04-2026 |
| 7 | Introducing Gemini 3.7 Flash | 0 | 22.13 | 13-08-2026 |
| 8 | Start building with Nano Banana 2 Lite and Gemini Omni Flash | 5 | 6 | 30-06-2026 |
| 9 | New ways to create faster with Gemini in Docs, Sheets, Slides and Drive | 5 | 7 | 10-03-2026 |
| 10 | NVIDIA выкладывает на Hugging Face по несколько моделей в неделю, ... | 0 | 9.68 | 28-07-2026 |