<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz" -->

---
title: Your Agent Needs a Harness, Not a Framework | daily.dev
description: An open-source reference project called Utah demonstrates building AI agents on durable, event-driven infrastructure instead of a traditional agent framework....
canonical: https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Your Agent Needs a Harness, Not a Framework | daily.dev
og:description: An open-source reference project called Utah demonstrates building AI agents on durable, event-driven infrastructure instead of a traditional agent framework....
og:url: https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz
og:image: https://api.daily.dev/og/posts/sfYhmDNWz.png
og:image:alt: Your Agent Needs a Harness, Not a Framework
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Your Agent Needs a Harness, Not a Framework

**[Inngest Blog](https://daily.dev/sources/inngest)** · 13 min read · 0 upvotes · 0 comments

## Summary

An open-source reference project called Utah demonstrates building AI agents on durable, event-driven infrastructure instead of a traditional agent framework. Built with Inngest, Utah connects Telegram and Slack via webhooks to a think-act-observe agent loop where every LLM call and tool call is an independently retryable Inngest step. It uses pi-ai for provider-agnostic LLM calls, pi-coding-agent for battle-tested tools (read, write, edit, bash, grep), sub-agent delegation via step.invoke(), singleton concurrency for one conversation at a time, and a multi-tier context pruning and compaction system to manage LLM context windows. The team also shares lessons learned around context management, multi-provider LLM support, unsolved steering problems, and plans to explore multi-player, sandboxed, and self-modifying versions of the system.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.inngest.com/blog/your-agent-needs-a-harness-not-a-framework>

## Questions this post answers

### How do you prevent an AI agent conversation loop from losing context or crashing on tool call failures in Inngest?

Each LLM call and tool call is wrapped in an Inngest step, making it independently retryable and durable. If a step fails, for example an LLM API returning a 500 on iteration 3, only that step retries while results from prior iterations stay persisted and do not re-execute. Inngest auto-indexes duplicate step IDs like think:0, think:1 across loop iterations automatically.

_daily.dev surfaces practical patterns like this for engineers wiring durable execution into agent loops._

### How do you stop an LLM agent's context window from ballooning when tools return large outputs?

A two-tier pruning and compaction system handles this: old tool results are soft-trimmed to a max of 4000 characters (keeping 1500 head and tail characters) or hard-cleared with a placeholder once total context exceeds 50,000 characters, while the last three assistant turns always stay intact. Separately, session-level compaction summarizes conversation history once estimated tokens exceed a threshold, and overflow recovery force-compacts messages and retries if the LLM returns a context-too-large error.

_developers wrestling with LLM context limits can find agent architecture deep dives like this via daily.dev._

### How do you handle a user sending a new message while an AI agent is still mid-run processing a previous one?

One approach is Inngest's singleton concurrency configuration keyed on a session identifier with cancel mode: only one agent run executes per conversation at a time, and if a new message arrives while a run is active, that run is cancelled and a fresh one starts with the latest message and persisted session state. In-flight work from the cancelled run is lost, and seamless mid-run steering remains unsolved.

_teams designing concurrency handling for chat agents can track approaches like this through daily.dev._

## Similar posts on daily.dev

- [AI Agents at work: real-time platform insights in Slack](https://daily.dev/posts/ai-agents-at-work-real-time-platform-insights-in-slack-as53pzdrz) · monday Engineering · 17 upvotes · 0 comments
- [A no-nonsense explainer to Agentic AI](https://daily.dev/posts/a-no-nonsense-explainer-to-agentic-ai-zu62mz1ox) · Tailscale · 3 upvotes · 0 comments
- [What I learned building an opinionated and minimal coding agent](https://daily.dev/posts/what-i-learned-building-an-opinionated-and-minimal-coding-agent-myqj9tthj) · Hacker News · 1 upvotes · 0 comments
- [Three Years of Building Agents in Production \(Part 2\)](https://daily.dev/posts/three-years-of-building-agents-in-production-part-2--9uvhtnda1) · Mabl Engineering Blog · 2 upvotes · 1 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#architecture](https://daily.dev/tags/architecture), [#typescript](https://daily.dev/tags/typescript)

[View this post on daily.dev](https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Your Agent Needs a Harness, Not a Framework","url":"https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz"},"datePublished":"2026-09-13T19:57:20.391Z","dateModified":"2026-09-13T20:02:10.620Z","description":"An open-source reference project called Utah demonstrates building AI agents on durable, event-driven infrastructure instead of a traditional agent framework....","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b10e36591896577d4a1b8ca930559dec?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/b10e36591896577d4a1b8ca930559dec?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Inngest Blog","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Inngest Blog","logo":"https://media.daily.dev/image/upload/s--ieZKUa5y--/c_limit,w_256/f_auto,q_auto/v1789329423/logos/inngest?_a=BAMAMicg0","url":"https://daily.dev/sources/inngest"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,architecture,typescript","timeRequired":"PT13M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Inngest Blog","item":"https://daily.dev/sources/inngest"},{"@type":"ListItem","position":3,"name":"Your Agent Needs a Harness, Not a Framework"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/your-agent-needs-a-harness-not-a-framework-sfyhmdnwz#faq","mainEntity":[{"@type":"Question","name":"How do you prevent an AI agent conversation loop from losing context or crashing on tool call failures in Inngest?","acceptedAnswer":{"@type":"Answer","text":"Each LLM call and tool call is wrapped in an Inngest step, making it independently retryable and durable. If a step fails, for example an LLM API returning a 500 on iteration 3, only that step retries while results from prior iterations stay persisted and do not re-execute. Inngest auto-indexes duplicate step IDs like think:0, think:1 across loop iterations automatically. daily.dev surfaces practical patterns like this for engineers wiring durable execution into agent loops."}},{"@type":"Question","name":"How do you stop an LLM agent's context window from ballooning when tools return large outputs?","acceptedAnswer":{"@type":"Answer","text":"A two-tier pruning and compaction system handles this: old tool results are soft-trimmed to a max of 4000 characters (keeping 1500 head and tail characters) or hard-cleared with a placeholder once total context exceeds 50,000 characters, while the last three assistant turns always stay intact. Separately, session-level compaction summarizes conversation history once estimated tokens exceed a threshold, and overflow recovery force-compacts messages and retries if the LLM returns a context-too-large error. developers wrestling with LLM context limits can find agent architecture deep dives like this via daily.dev."}},{"@type":"Question","name":"How do you handle a user sending a new message while an AI agent is still mid-run processing a previous one?","acceptedAnswer":{"@type":"Answer","text":"One approach is Inngest's singleton concurrency configuration keyed on a session identifier with cancel mode: only one agent run executes per conversation at a time, and if a new message arrives while a run is active, that run is cancelled and a fresh one starts with the latest message and persisted session state. In-flight work from the cancelled run is lost, and seamless mid-run steering remains unsolved. teams designing concurrency handling for chat agents can track approaches like this through daily.dev."}}]}
```

