<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j" -->

---
title: The Open Source AI Stack | daily.dev
description: An overview of the open-source AI development stack for building agentic coding workflows, broken into the &#x27;MIGHT&#x27; layers: Model, Inference, Gateways/routers,...
canonical: https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: The Open Source AI Stack | daily.dev
og:description: An overview of the open-source AI development stack for building agentic coding workflows, broken into the &#x27;MIGHT&#x27; layers: Model, Inference, Gateways/routers,...
og:url: https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j
og:image: https://api.daily.dev/og/posts/cAaNFPy9j.png
og:image:alt: The Open Source AI Stack
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# The Open Source AI Stack

**[Together AI](https://daily.dev/sources/togetherai)** · 16 min read · 2 upvotes · 0 comments

## Summary

An overview of the open-source AI development stack for building agentic coding workflows, broken into the 'MIGHT' layers: Model, Inference, Gateways/routers, Harness, and Tools (skills/MCP). Covers when to use large vs. small open models (e.g., Kimi K3 vs. GLM 5.3 Flash), inference providers like Together AI, gateways like OpenRouter and Vercel AI Gateway, harnesses like OpenCode, PI, and Amp, and practices like managing context, starting fresh sessions, and a plan-implement-review workflow across multiple models. Argues the key advantage of open models is a composable stack where each layer can be swapped independently.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.together.ai/blog/the-open-source-ai-stack>

## Questions this post answers

### When should I use a large model like Kimi K3 versus a small model like GLM 5.3 Flash for coding tasks?

Use large models such as Kimi K3 for ambiguous, multi-step work like refactoring authentication systems, upgrading frameworks, or reviewing pull requests, since their extra capacity handles unclear requirements well. Use small models such as GLM 5.3 Flash for narrowly scoped tasks like adding a function option or writing tests, since they are roughly 6 times smaller and 20 times cheaper than Kimi K3 while matching its performance on well-specified work.

_daily.dev helps developers compare model tradeoffs like this when picking tools for a coding stack._

### What is the difference between an AI gateway and a harness in an AI coding agent stack?

A gateway sits in front of multiple inference providers, aggregating them behind one API so requests can be routed, compared on price and latency, and switched without code changes; examples include OpenRouter, Vercel AI Gateway, and self-hosted LiteLLM. A harness is the application layer that manages the conversation, gives the model tools to search code, read files, and apply patches, and connects it to a codebase; examples include OpenCode, PI, and Amp.

_developers assembling their own agent stack track these layer distinctions on daily.dev._

### What is a plan-implement-review workflow using multiple AI models for coding?

It splits a coding task across three model calls: a large model plans by breaking an open-ended prompt into isolated tasks, a small model implements each task in its own fresh session to keep context focused, and a large model reviews the completed work, looping back to planning if problems appear. This lets teams use cheaper small models for routine implementation while reserving larger models for reasoning-heavy planning and review.

_daily.dev keeps developers experimenting with multi-model agent workflows like this one moving._

## Similar posts on daily.dev

- [The State of Open Source AI — V1.0 · July 2026](https://daily.dev/posts/the-state-of-open-source-ai-v1-0-july-2026-gjzwgazbw) · Hacker News · 1 upvotes · 0 comments
- [Medium](https://daily.dev/posts/medium-83fdoijq7) · Medium · 0 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#mcp](https://daily.dev/tags/mcp), [#ai-inference](https://daily.dev/tags/ai-inference)

[View this post on daily.dev](https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"The Open Source AI Stack","url":"https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j"},"datePublished":"2026-09-09T17:26:31.806Z","dateModified":"2026-09-14T08:52:41.866Z","description":"An overview of the open-source AI development stack for building agentic coding workflows, broken into the 'MIGHT' layers: Model, Inference, Gateways/routers,...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f47321cb095f2b3656f9547f91cc8781?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/f47321cb095f2b3656f9547f91cc8781?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Together AI","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Together AI","logo":"https://media.daily.dev/image/upload/s--tCjWcJfJ--/f_auto,q_auto/v1780213200/logos/togetherai?_a=BAMAMiWQ0","url":"https://daily.dev/sources/togetherai"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,mcp,ai-inference","timeRequired":"PT16M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Together AI","item":"https://daily.dev/sources/togetherai"},{"@type":"ListItem","position":3,"name":"The Open Source AI Stack"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/the-open-source-ai-stack-caanfpy9j#faq","mainEntity":[{"@type":"Question","name":"When should I use a large model like Kimi K3 versus a small model like GLM 5.3 Flash for coding tasks?","acceptedAnswer":{"@type":"Answer","text":"Use large models such as Kimi K3 for ambiguous, multi-step work like refactoring authentication systems, upgrading frameworks, or reviewing pull requests, since their extra capacity handles unclear requirements well. Use small models such as GLM 5.3 Flash for narrowly scoped tasks like adding a function option or writing tests, since they are roughly 6 times smaller and 20 times cheaper than Kimi K3 while matching its performance on well-specified work. daily.dev helps developers compare model tradeoffs like this when picking tools for a coding stack."}},{"@type":"Question","name":"What is the difference between an AI gateway and a harness in an AI coding agent stack?","acceptedAnswer":{"@type":"Answer","text":"A gateway sits in front of multiple inference providers, aggregating them behind one API so requests can be routed, compared on price and latency, and switched without code changes; examples include OpenRouter, Vercel AI Gateway, and self-hosted LiteLLM. A harness is the application layer that manages the conversation, gives the model tools to search code, read files, and apply patches, and connects it to a codebase; examples include OpenCode, PI, and Amp. developers assembling their own agent stack track these layer distinctions on daily.dev."}},{"@type":"Question","name":"What is a plan-implement-review workflow using multiple AI models for coding?","acceptedAnswer":{"@type":"Answer","text":"It splits a coding task across three model calls: a large model plans by breaking an open-ended prompt into isolated tasks, a small model implements each task in its own fresh session to keep context focused, and a large model reviews the completed work, looping back to planning if problems appear. This lets teams use cheaper small models for routine implementation while reserving larger models for reasoning-heavy planning and review. daily.dev keeps developers experimenting with multi-model agent workflows like this one moving."}}]}
```

