ZML, a Paris-based AI infrastructure startup backed by $20 million in venture funding and endorsed by Turing Award winner Yann LeCun, launched LLMD today — a free inference server that runs open-source large language models across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc silicon. The 20-person team is targeting one of the fastest-growing bottlenecks in AI: how to move prompts through hardware quickly and cheaply without being locked to a single chip vendor.
Inference has overtaken training as the operational cost center for most AI deployments, and the software stack for serving models remains fragmented. ZML founder Steeve Morin argues that enterprises and cloud operators want the option to mix chips — some cheaper, some more energy-efficient — but current tooling forces them into vendor silos. LLMD is pitched as the layer that lets a single deployment span across those silos while extracting peak performance from each.
Morin's framing is that unlocking multi-chip inference is not just an engineering exercise but a market lever. If a workload can move fluidly between Nvidia H-series parts, AMD Instinct accelerators, Google TPUs and specialty European silicon, the pricing power of any single vendor weakens.
“The idea is to give people back the power to create their own system and achieve real efficiency gains that allow [AI] to be disseminated”— Steeve Morin, ZML founder
Key facts
- 01ZML released LLMD, a free LLM inference server that runs open-source models across Nvidia, AMD, Google TPU, Apple Metal and Intel Arc chips.
- 02The Paris startup has 20 employees and has raised $20 million from investors including 20VC, Kima Ventures, Kindred Capital and LocalGlobe.
- 03ZML competes with Baseten, recently valued at $13 billion, plus Inferact from the vLLM creators and RadixArk behind SGLang.
- 04Founder Steeve Morin previously served as VP of engineering at Zenly, which Snapchat acquired for nine figures in 2017.
- 05LLMD launches free, unlike ZML's open-source inference framework released in 2024 and updated in March 2026.
The competitive field is already crowded and well-capitalized. Baseten was recently valued at $13 billion. Inferact, spun out of the team behind the open-source vLLM project, and RadixArk, the commercial company behind SGLang, both compete on portions of what LLMD does. Morin says ZML's ambitions are broader — spanning novel silicon that vLLM and SGLang do not target.
That novel silicon largely comes from Europe. Morin cited Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud and VSORA as chipmakers ZML is working with on capabilities he says have not been built anywhere else. He is not bearish on Nvidia — ZML has a working relationship with the incumbent, which has been repositioning itself around the inference gold rush — but the strategic bet is that a heterogeneous chip landscape is coming, and someone needs to write the software for it.
The co-design angle matters commercially. When an inference framework is tuned to a specific chip at the compiler and kernel level, both sides gain: the chipmaker gets a viable software story to sell against Nvidia's CUDA moat, and the framework vendor gets performance numbers no generic runtime can match.
“We have reached the point where we are co-designing silicon”— Steeve Morin, ZML founder
ZML is releasing LLMD as free software, but unlike the ML framework the company released in 2024 and updated in March 2026, LLMD is not open source. Morin said the plan is to measure usage first and monetize where the product is most useful, rather than pricing aggressively from day one and hindering adoption. There is no announced timeline for a paid tier.
Morin's credibility comes partly from Zenly, the location-sharing app where he served as VP of engineering before Snapchat acquired it in 2017 for nine figures. That exit helped him raise $20 million for ZML from a roster that includes Harry Stebbings' 20VC, Xavier Niel's Kima Ventures, Kindred Capital, LocalGlobe, Puzzle Ventures, AALVC, Drysdale Ventures and >commit. The cap table also includes Dagger and Docker founder Solomon Hykes, Hugging Face's Clément Delangue and Julien Chaumond, and LeCun himself, now at AMI Labs.
“I couldn't do ZML anywhere but in Paris”— Steeve Morin, ZML founder
The open questions are adoption and durability. LLMD's value proposition depends on enterprises actually running mixed-chip fleets in production, and most large deployments today still standardize on a single vendor for operational simplicity. Portability software has historically struggled to displace vertically integrated stacks — CUDA remains dominant precisely because it is deep, not just wide. Whether LLMD can convert its breadth into measurable cost savings that justify the switching effort is unproven, and the company has not yet published benchmark comparisons against vLLM or SGLang on shared hardware.
For the AI infrastructure market, ZML is a bet that the inference layer eventually looks less like a monoculture and more like the CPU market of the late 1990s — multiple vendors, standardized software abstractions, and margin pressure on whoever tries to hold a proprietary moat. That thesis only pays off if the European chip startups Morin is co-designing with actually ship at scale, and if the hyperscalers decide multi-vendor sourcing is worth the operational tax. If both happen, LLMD is well positioned. If Nvidia's supply catches demand and CUDA keeps its grip, ZML becomes a specialist tool rather than a category-defining platform.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



