
ModelsLead story
Google DeepMind's DiffusionGemma generates text 4x faster on GPUs
The 26B MoE model hits 1,000+ tokens per second on an H100 by drafting 256 tokens in parallel instead of one at a time.
Jaeden SchaferEditor in Chief5 min read
Tag · 3 stories
Every story tagged Gemma 4 on AI Chat Daily.

The 26B MoE model hits 1,000+ tokens per second on an H100 by drafting 256 tokens in parallel instead of one at a time.

The mid-sized open model runs on 16GB of VRAM and nears 26B benchmark performance, with native audio and vision baked into the LLM backbone.

The new mid-weight Gemma slots between mobile and workstation variants, with model weights just under 18GB available on Hugging Face and Kaggle.
The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.
The briefing read inside teams at