
BusinessLead story
Kog claims 30x faster LLM inference on existing Nvidia and AMD GPUs
The 11-person French startup hit 3,000 tokens per second on a 2B model and now needs to prove the same trick works on real LLMs by September.
Jaeden SchaferEditor in Chief5 min read