<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz" -->

---
title: ZML releases free LLM inference server that runs across...
description: ZML, a Paris-based AI startup with $20M in funding, has launched LLMD — a free LLM inference server designed to run open-source models across Nvidia, AMD,...
canonical: https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: ZML releases free LLM inference server that runs across Nvidia, AMD, TPU, and Apple chips | daily.dev
og:description: ZML, a Paris-based AI startup with $20M in funding, has launched LLMD — a free LLM inference server designed to run open-source models across Nvidia, AMD,...
og:url: https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz
og:image: https://api.daily.dev/og/posts/CzKK5rHrZ.png
og:image:alt: ZML releases free LLM inference server that runs across Nvidia, AMD, TPU, and Apple chips
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# ZML releases free LLM inference server that runs across Nvidia, AMD, TPU, and Apple chips

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

ZML, a Paris-based AI startup with $20M in funding, has launched LLMD — a free LLM inference server designed to run open-source models across Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc hardware. The key differentiator is cross-chip compatibility, allowing enterprises with mixed hardware setups to avoid vendor lock-in and optimize for cost and energy. LLMD is free but not open source, unlike ZML's earlier ML framework. The inference server market is competitive, with ZML going up against vLLM-backed Inferact, SGLang-backed RadixArk, and Baseten. ZML also co-designs silicon with chip partners, an unusual move for a 20-person team. The launch was amplified by a retweet from Yann LeCun.

## Content

Paris-based AI startup ZML has launched ZML/LLMD, a free LLM inference server designed to run open-source models across virtually every major chip architecture — Nvidia, AMD, Google TPUs, Apple Metal, and Intel Arc.

The core pitch is cross-chip compatibility. Most inference tools are effectively built around Nvidia's CUDA ecosystem, which means companies either pay Nvidia's prices or deal with painful software rewrites when they want to try something different. ZML wants to make chip choice a business decision rather than a technical constraint, letting enterprises mix hardware types to optimize for cost and energy use.

The company was founded by Steeve Morin and has raised $20M in VC funding. Yann LeCun is among its backers, along with angel investors including Docker and Dagger founder Solomon Hykes and several Hugging Face co-founders. The team is 20 people.

ZML/LLMD is free but not open source — a different approach from ZML's earlier ML framework, which was open source. The company says it's holding off on monetization until it has a clearer picture of how people actually use the product.

The server competes with vLLM (backed by Inferact), SGLang (backed by RadixArk), and Baseten. ZML's differentiation is the cross-chip focus and a co-design relationship with chip partners. That positioning also makes it potentially useful for European chipmakers like Axelera, Fractile, and Kalray, which have capable hardware but often lack the software ecosystem to compete with established players.

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-inference](https://daily.dev/tags/ai-inference), [#vllm](https://daily.dev/tags/vllm)

[View this post on daily.dev](https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"ZML releases free LLM inference server that runs across Nvidia, AMD, TPU, and Apple chips","url":"https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz"},"datePublished":"2026-07-08T08:57:01.223Z","dateModified":"2026-07-08T17:28:22.929Z","description":"ZML, a Paris-based AI startup with $20M in funding, has launched LLMD — a free LLM inference server designed to run open-source models across Nvidia, AMD,...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/zml-releases-free-llm-inference-server-that-runs-across-nvidia-amd-tpu-and-apple-chips-czkk5rhrz","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-inference,vllm","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"ZML releases free LLM inference server that runs across Nvidia, AMD, TPU, and Apple chips"}]}
```

