<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3" -->

---
title: AI Development on Windows: from PyTorch and llama.cpp to...
description: Microsoft's Windows ML now includes experimental support for running GGUF models locally via llama.cpp, adding new task-specific APIs (Text Generation, Speech...
canonical: https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: AI Development on Windows: from PyTorch and llama.cpp to Windows ML | daily.dev
og:description: Microsoft's Windows ML now includes experimental support for running GGUF models locally via llama.cpp, adding new task-specific APIs (Text Generation, Speech...
og:url: https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3
og:image: https://api.daily.dev/og/posts/f2T66Cns3.png
og:image:alt: AI Development on Windows: from PyTorch and llama.cpp to Windows ML
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Development on Windows: from PyTorch and llama.cpp to Windows ML

**[DevBlogs](https://daily.dev/sources/devblogs)** · 10 min read · 0 upvotes · 0 comments

## Summary

Microsoft's Windows ML now includes experimental support for running GGUF models locally via llama.cpp, adding new task-specific APIs (Text Generation, Speech Recognition) with an OpenAI-compatible endpoint. A new Windows-native Runtime API offers zero-copy data handling, deterministic multi-model pipelines, and ahead-of-time compilation, running alongside the existing ONNX Runtime APIs. Microsoft also contributed performance work to llama.cpp (CUDA kernel optimization, speculative decoding, multi-GPU execution) with NVIDIA, and PyTorch and Triton now have native Windows Arm64 builds, enabling a full PyTorch-train-to-Windows-ML-deploy workflow demonstrated with a ResNet18 export and the Windows ML CLI.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://devblogs.microsoft.com/foundry-on-windows/build-on-winml-oct-7-26>

## Questions this post answers

### How do I run a GGUF model locally on Windows using Windows ML?

Windows ML now ships an experimental Text Generation API that accepts both GGUF and ONNX models through one surface, automatically selecting llama.cpp as the execution engine for GGUF. You start a local server with a command like WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080, then point the standard OpenAI Python SDK at its OpenAI-compatible local endpoint to run inference.

_daily.dev surfaces practical guidance like this for developers wiring local GGUF inference into Windows apps._

### Does PyTorch support native Windows Arm64 builds?

Yes, PyTorch now offers official native Windows Arm64 CPU builds, and NVIDIA separately publishes CUDA-enabled Windows Arm64 packages for supported hardware. This gives model developers a native foundation for training, fine-tuning, and inference directly on Arm-based Windows AI systems, rather than relying on emulation or cloud-only workflows.

_developers evaluating Arm-based AI workstations can track PyTorch and Triton ecosystem shifts on daily.dev._

### What is the difference between the Windows ML Runtime API and the existing ONNX Runtime APIs in Windows ML?

The ONNX Runtime APIs remain fully supported for broad compatibility, while the new experimental Windows ML Runtime API is a Windows-native inferencing path built for deeper OS integration and finer control. It supports zero-copy Windows-native data types, deterministic multi-model pipelines with per-stage device placement across CPU, GPU, and NPU, and ahead-of-time model compilation; both APIs ship side by side.

_teams choosing an inferencing path for Windows apps can follow Windows ML API changes on daily.dev._

## Similar posts on daily.dev

- [Official home for llama.cpp](https://daily.dev/posts/official-home-for-llama-cpp-bgx69sgjl) · Hacker News · 2 upvotes · 0 comments
- [LiteRT: The Universal Framework for On-Device AI](https://daily.dev/posts/litert-the-universal-framework-for-on-device-ai-dkvzmc8hv) · Google Developers · 0 upvotes · 0 comments

---

Tags: [#pytorch](https://daily.dev/tags/pytorch), [#llama-cpp](https://daily.dev/tags/llama-cpp)

[View this post on daily.dev](https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"AI Development on Windows: from PyTorch and llama.cpp to Windows ML","url":"https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3"},"datePublished":"2026-10-07T18:03:35.638Z","dateModified":"2026-10-07T20:15:48.730Z","description":"Microsoft's Windows ML now includes experimental support for running GGUF models locally via llama.cpp, adding new task-specific APIs (Text Generation, Speech...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8a854210c6fb96d2c386f8fd37d5a869?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8a854210c6fb96d2c386f8fd37d5a869?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"DevBlogs","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"DevBlogs","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/8f68b453325f482ebeb73fb780092713","url":"https://daily.dev/sources/devblogs"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"pytorch,llama-cpp","timeRequired":"PT10M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"DevBlogs","item":"https://daily.dev/sources/devblogs"},{"@type":"ListItem","position":3,"name":"AI Development on Windows: from PyTorch and llama.cpp to Windows ML"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/ai-development-on-windows-from-pytorch-and-llama-cpp-to-windows-ml-f2t66cns3#faq","mainEntity":[{"@type":"Question","name":"How do I run a GGUF model locally on Windows using Windows ML?","acceptedAnswer":{"@type":"Answer","text":"Windows ML now ships an experimental Text Generation API that accepts both GGUF and ONNX models through one surface, automatically selecting llama.cpp as the execution engine for GGUF. You start a local server with a command like WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080, then point the standard OpenAI Python SDK at its OpenAI-compatible local endpoint to run inference. daily.dev surfaces practical guidance like this for developers wiring local GGUF inference into Windows apps."}},{"@type":"Question","name":"Does PyTorch support native Windows Arm64 builds?","acceptedAnswer":{"@type":"Answer","text":"Yes, PyTorch now offers official native Windows Arm64 CPU builds, and NVIDIA separately publishes CUDA-enabled Windows Arm64 packages for supported hardware. This gives model developers a native foundation for training, fine-tuning, and inference directly on Arm-based Windows AI systems, rather than relying on emulation or cloud-only workflows. developers evaluating Arm-based AI workstations can track PyTorch and Triton ecosystem shifts on daily.dev."}},{"@type":"Question","name":"What is the difference between the Windows ML Runtime API and the existing ONNX Runtime APIs in Windows ML?","acceptedAnswer":{"@type":"Answer","text":"The ONNX Runtime APIs remain fully supported for broad compatibility, while the new experimental Windows ML Runtime API is a Windows-native inferencing path built for deeper OS integration and finer control. It supports zero-copy Windows-native data types, deterministic multi-model pipelines with per-stage device placement across CPU, GPU, and NPU, and ahead-of-time model compilation; both APIs ship side by side. teams choosing an inferencing path for Windows apps can follow Windows ML API changes on daily.dev."}}]}
```

