Skip to main content
Live
Main content

Kog claims 30x faster LLM inference on existing Nvidia and AMD GPUs

The 11-person French startup hit 3,000 tokens per second on a 2B model and now needs to prove the same trick works on real LLMs by September.

Jaeden Schafer
Editor in Chief · · 5 min read
Kog claims 30x faster LLM inference on existing Nvidia and AMD GPUs

French startup Kog claims it can extract 30x faster LLM inference from the same Nvidia H200 and AMD MI300X GPUs that enterprises already run, a bet that inference speed is a software problem more than a silicon problem. Its May tech preview hit 3,000 tokens per second per request on a purpose-built 2-billion-parameter model called Laneformer 2B, which the company has since open-sourced. The demo landed on the front page of Hacker News and generated 200 business leads, according to CEO Gaël Delalleau.

The pitch arrives as inference cost and latency become the binding constraint across the industry. Cerebras rode that thesis to a warm IPO debut in May with purpose-built chips. Kog is arguing the opposite: the H200 and MI300X have enough memory bandwidth to deliver Cerebras-class single-request decoding, provided the software goes deep enough into the hardware.

extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own
Gaël Delalleau, Kog CEO

The Kog Inference Engine, or KIE, is the product. Delalleau frames the case bluntly, saying that on the hardware enterprises already own, extremely fast single-request decoding is possible if the stack is engineered for it.

Key facts

  • 01Kog's tech preview clocked 3,000 tokens per second per request on a 2B-parameter model running on AMD MI300X and Nvidia H200 GPUs.
  • 02The company is targeting 30x faster LLM inference on standard datacenter GPUs, with a September deadline to hit 10x on its first major model.
  • 03CEO Gaël Delalleau says Kog collected 200 business leads after its May Hacker News debut, with software engineering the leading use case.
  • 04The 11-person team spends several weeks or months per GPU on low-level engineering, an approach Delalleau traces to his DEFCON CTF background.
  • 05Kog is backed by Varsity VC, Scaleway, Bpifrance, and the French Tech 2030 program, and plans to raise a Series A after proving out LLM performance.

The initial customer set skews toward workflows where latency directly costs money. Veteran users of Claude Code routinely wait hours for long-running agentic tasks, and Anthropic already charges a premium for Claude's Fast Mode. Kog is targeting the same pain point from the infrastructure side, along with design partners running prompt-to-app and prompt-to-game generation products where faster output translates into more revenue per user.

The commercial catch is that Kog's prospective customers do not want to fine-tune small models to get the speedup. That forced a strategic pivot within weeks of the demo.

The 3,000 TPS demo ran on a 2B-parameter model, and skeptics argued the approach would not scale to frontier-sized LLMs, where model weights dominate memory bandwidth. Delalleau disagrees, arguing that newer GPUs keep adding memory bandwidth that most inference stacks fail to exploit. Proving that on a real LLM is now the company's central technical bet.

Kog is not the only French team trying to squeeze more out of commodity accelerators. ZML has released hardware-agnostic software that bypasses Nvidia's CUDA to run inference across competing chips. Delalleau positions Kog closer to Stanford University's Hazy Research lab, with a deeper focus on per-GPU acceleration rather than portability.

The methodology is hands-on and slow. Delalleau, an École Polytechnique-trained solid-state physicist and a four-time finalist at DEFCON's CTF tournament, says Kog will dedicate several weeks or even months per new GPU to low-level engineering, including reverse-engineering down to assembly. With 11 people on the team, that caps how many chips Kog can support at once. The longer-term plan is to fold the methodology into agent-based pipelines that widen coverage across hardware and models.

Related · from this week
Ramp launches Router, an AI model routing service to rival OpenRouter
Jaeden Schafer · 4 min read →

The near-term milestone is a September delivery of Kog's first major model running at 10x speed, which Delalleau frames as the trigger for a Series A. The company is already backed by Varsity VC, whose partner Kamel Zeroual co-founded Delalleau's 2009 startup Stribe, and receives support from Scaleway, Bpifrance, and the French Tech 2030 program.

The counterweight is that a 2B-parameter demo does not prove a 70B or 400B claim. Memory bandwidth helps decoding, but frontier LLMs are constrained by weight-loading and KV-cache pressure in ways small models are not, and inference-optimization startups have a long history of demos that shrink when models grow. Kog also has to fend off well-funded incumbents — vLLM, TensorRT-LLM, and SGLang all ship improvements monthly — while shipping GPU-specific work at the pace an 11-person team allows.

For the broader market, Kog is a useful stress test of the Cerebras thesis. If a small European team can pull 10x out of an H200 with pure software, the case for specialty inference silicon narrows and the value migrates back to the CUDA and ROCm software layers. If Kog misses September, the specialty-chip camp gets a data point that GPU inference has a ceiling closer to today's numbers than its critics claim. Either outcome reshapes where the next wave of inference dollars gets spent, and the answer arrives in weeks, not quarters.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

Ramp launches Router, an AI model routing service to rival OpenRouter
Business

Ramp launches Router, an AI model routing service to rival OpenRouter

The corporate expense platform opens a toll house for AI inference, connecting eight model providers with a $26 launch credit.

Jaeden Schafer4 min read
Cerebras stock drops 20% after margin guidance spooks first earnings call
Business

Cerebras stock drops 20% after margin guidance spooks first earnings call

The AI chipmaker beat Q1 estimates with $193M in revenue, but a 38–41% full-year margin forecast triggered the selloff.

Jaeden Schafer5 min read
Baseten nears $1.5B round at $13B valuation, up 160% in five months
Business

Baseten nears $1.5B round at $13B valuation, up 160% in five months

The AI inference startup is closing a split-priced round five months after a $300M Series E, riding the inference gold rush.

Jaeden Schafer4 min read