
ModelsLead story
Google DeepMind's DiffusionGemma generates text 4x faster on GPUs
The 26B MoE model hits 1,000+ tokens per second on an H100 by drafting 256 tokens in parallel instead of one at a time.
Jaeden SchaferEditor in Chief5 min read
Tag · 1 story
Every story tagged DiffusionGemma on AI Chat Daily.

The 26B MoE model hits 1,000+ tokens per second on an H100 by drafting 256 tokens in parallel instead of one at a time.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at