Skip to main content
Live
Main content

ZML launches free LLMD inference server across Nvidia, AMD, Google TPU and Apple chips

The Paris startup, backed by $20M and Yann LeCun, wants to break silicon lock-in as inference costs climb.

Jaeden Schafer
Editor in Chief · · 5 min read
ZML launches free LLMD inference server across Nvidia, AMD, Google TPU and Apple chips

ZML, a Paris-based AI infrastructure startup backed by $20 million in venture funding and endorsed by Turing Award winner Yann LeCun, launched LLMD today — a free inference server that runs open-source large language models across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc silicon. The 20-person team is targeting one of the fastest-growing bottlenecks in AI: how to move prompts through hardware quickly and cheaply without being locked to a single chip vendor.

Inference has overtaken training as the operational cost center for most AI deployments, and the software stack for serving models remains fragmented. ZML founder Steeve Morin argues that enterprises and cloud operators want the option to mix chips — some cheaper, some more energy-efficient — but current tooling forces them into vendor silos. LLMD is pitched as the layer that lets a single deployment span across those silos while extracting peak performance from each.

Morin's framing is that unlocking multi-chip inference is not just an engineering exercise but a market lever. If a workload can move fluidly between Nvidia H-series parts, AMD Instinct accelerators, Google TPUs and specialty European silicon, the pricing power of any single vendor weakens.

The idea is to give people back the power to create their own system and achieve real efficiency gains that allow [AI] to be disseminated
Steeve Morin, ZML founder

Key facts

  • 01ZML released LLMD, a free LLM inference server that runs open-source models across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc chips.
  • 02The Paris startup has 20 employees and has raised $20 million from investors including 20VC, Kima Ventures, Kindred Capital and LocalGlobe.
  • 03ZML competes with Baseten, recently valued at $13 billion, plus Inferact from the vLLM creators and RadixArk behind SGLang.
  • 04Founder Steeve Morin previously served as VP of engineering at Zenly, which Snapchat acquired for nine figures in 2017.
  • 05LLMD launches free, unlike ZML's open-source inference framework released in 2024 and updated in March 2026.

The competitive field is already crowded and well-capitalized. Baseten was recently valued at $13 billion. Inferact, spun out of the team behind the open-source vLLM project, and RadixArk, the commercial company behind SGLang, both compete on portions of what LLMD does. Morin says ZML's ambitions are broader — spanning novel silicon that vLLM and SGLang do not target.

That novel silicon largely comes from Europe. Morin cited Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud and VSORA as chipmakers ZML is working with on capabilities he says have not been built anywhere else. He is not bearish on Nvidia — ZML has a working relationship with the incumbent, which has been repositioning itself around the inference gold rush — but the strategic bet is that a heterogeneous chip landscape is coming, and someone needs to write the software for it.

The co-design angle matters commercially. When an inference framework is tuned to a specific chip at the compiler and kernel level, both sides gain: the chipmaker gets a viable software story to sell against Nvidia's CUDA moat, and the framework vendor gets performance numbers no generic runtime can match.

We have reached the point where we are co-designing silicon
Steeve Morin, ZML founder

ZML is releasing LLMD as free software, but unlike the ML framework the company released in 2024 and updated in March 2026, LLMD is not open source. Morin said the plan is to measure usage first and monetize where the product is most useful, rather than pricing aggressively from day one and hindering adoption. There is no announced timeline for a paid tier.

Morin's credibility comes partly from Zenly, the location-sharing app where he served as VP of engineering before Snapchat acquired it in 2017 for nine figures. That exit helped him raise $20 million for ZML from a roster that includes Harry Stebbings' 20VC, Xavier Niel's Kima Ventures, Kindred Capital, LocalGlobe, Puzzle Ventures, AALVC, Drysdale Ventures and >commit. The cap table also includes Dagger and Docker founder Solomon Hykes, Hugging Face's Clément Delangue and Julien Chaumond, and LeCun himself, now at AMI Labs.

I couldn't do ZML anywhere but in Paris
Steeve Morin, ZML founder
Related · from this week
AMD unveils Helios rack system to challenge Nvidia in AI data centers
Jaeden Schafer · 5 min read →

The open questions are adoption and durability. LLMD's value proposition depends on enterprises actually running mixed-chip fleets in production, and most large deployments today still standardize on a single vendor for operational simplicity. Portability software has historically struggled to displace vertically integrated stacks — CUDA remains dominant precisely because it is deep, not just wide. Whether LLMD can convert its breadth into measurable cost savings that justify the switching effort is unproven, and the company has not yet published benchmark comparisons against vLLM or SGLang on shared hardware.

For the AI infrastructure market, ZML is a bet that the inference layer eventually looks less like a monoculture and more like the CPU market of the late 1990s — multiple vendors, standardized software abstractions, and margin pressure on whoever tries to hold a proprietary moat. That thesis only pays off if the European chip startups Morin is co-designing with actually ship at scale, and if the hyperscalers decide multi-vendor sourcing is worth the operational tax. If both happen, LLMD is well positioned. If Nvidia's supply catches demand and CUDA keeps its grip, ZML becomes a specialist tool rather than a category-defining platform.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

AMD unveils Helios rack system to challenge Nvidia in AI data centers
Models

AMD unveils Helios rack system to challenge Nvidia in AI data centers

Lisa Su calls Helios the highest-performance AI rack, with OpenAI, Meta, Oracle, Anthropic, and Microsoft lined up to deploy it at gigawatt scale.

Jaeden Schafer5 min read
Intel's Crescent Island AI chip ships this year, undercutting Nvidia on cost and cooling
Models

Intel's Crescent Island AI chip ships this year, undercutting Nvidia on cost and cooling

The air-cooled inference GPU uses LPDDR5 memory instead of HBM, betting that cheaper silicon wins the inference market Nvidia hasn't locked down.

Jaeden Schafer5 min read
AMD commits $10B to Taiwan's AI chip ecosystem as shares double
Business

AMD commits $10B to Taiwan's AI chip ecosystem as shares double

The investment targets advanced packaging and manufacturing for Helios, AMD's AI server system shipping in H2 2026.

Jaeden Schafer5 min read