<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp" -->

---
title: How to Build AI Agents That Don’t Start Over When They Fail
description: Long-running AI agent workflows fail differently than simple background jobs: LLM calls are slow and expensive, outputs are non-deterministic, side effects...
canonical: https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How to Build AI Agents That Don’t Start Over When They Fail | daily.dev
og:description: Long-running AI agent workflows fail differently than simple background jobs: LLM calls are slow and expensive, outputs are non-deterministic, side effects...
og:url: https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp
og:image: https://api.daily.dev/og/posts/hDtNQPnLp.png
og:image:alt: How to Build AI Agents That Don’t Start Over When They Fail
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How to Build AI Agents That Don’t Start Over When They Fail

**[System Design Newsletter](https://daily.dev/sources/systemdesignnews)** · 31 min read · 3 upvotes · 1 comments

## Summary

Long-running AI agent workflows fail differently than simple background jobs: LLM calls are slow and expensive, outputs are non-deterministic, side effects like sending emails or charging cards can't be safely repeated, and some steps wait on humans for days. Durable execution solves this by checkpointing completed steps so failures resume from the last successful point rather than restarting entirely. Inngest is used as a case study, contrasted with BullMQ (queue-first, more operational control) and Temporal (workflow-first, more complex). Key building blocks covered include checkpointing via step.run(), waiting via step.sleep() and step.waitForEvent(), retry policies, idempotency for side effects, flow control (concurrency, throttling, rate limiting, debouncing, priority), and observability through tracing and replay. Developers still must handle idempotent side effects, step boundary design, retry policies, and cross-system data consistency themselves.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://newsletter.systemdesign.one/p/durable-ai-agents>

## Questions this post answers

### How is durable execution different from just retrying a failed job?

A retry reruns an entire failed task or workflow from the start, while durable execution (checkpointing) remembers which steps already completed successfully and only reruns the failed step. For a multi-step AI agent that retrieved documents, extracted evidence, and then failed during drafting, durable execution skips the retrieval and extraction and resumes at drafting, saving time, tokens, and avoiding duplicate side effects like re-sent emails.

_daily.dev surfaces this kind of workflow-recovery reasoning for engineers designing resilient agent pipelines._

### When should I choose BullMQ, Temporal, or Inngest for building a durable AI agent workflow?

BullMQ fits teams with a dedicated infrastructure team who want fine-grained control over queues, workers, and retries but are willing to build checkpointing and resumability themselves. Temporal suits highly complex workflow graphs with many interdependent workflows, at the cost of added operational complexity. Inngest keeps workflow logic in application code via step.run(), handling checkpointing, retries, and waiting automatically, trading some low-level control for simplicity.

_developers weighing queue-first versus workflow-first tools for agent infrastructure can track these tradeoffs on daily.dev._

### How does step.waitForEvent() let an AI agent workflow pause for human approval without wasting compute?

step.waitForEvent() in Inngest pauses a workflow's execution and releases its compute while waiting for a matching external event, such as an editor approving a draft, instead of holding a worker process open. A timeout can be configured for cases where the event never arrives, and once the event lands, the workflow resumes with its prior progress intact, even after waiting for days.

_teams building human-in-the-loop agent approvals can compare these waiting patterns on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@agustinbarrientos** · 0 upvotes

> Checkpointing doesn't make side effects safe by itself. The dangerous boundary is a write succeeding before its result is recorded, which still needs idempotency keys or compensation

## Similar posts on daily.dev

- [All your workflows are about to be long running. Build for it.](https://daily.dev/posts/all-your-workflows-are-about-to-be-long-running-build-for-it--ff8epi154) · Inngest Blog · 0 upvotes · 0 comments
- [5 Resilience Patterns for AI Agents](https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah) · The T-Shaped Dev · 1 upvotes · 0 comments
- [Durable AI Workflows 2026: Inngest, Trigger.dev, Vercel Workflow](https://daily.dev/posts/durable-ai-workflows-2026-inngest-trigger-dev-vercel-workflow-un2tgyyqs) · Alex CloudStar · 0 upvotes · 0 comments
- [Building Reliable Production AI with Durable Workflows](https://daily.dev/posts/building-reliable-production-ai-with-durable-workflows-8k7gdooen) · Salesforce Engineering · 2 upvotes · 1 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#workflow-orchestration](https://daily.dev/tags/workflow-orchestration)

[View this post on daily.dev](https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How to Build AI Agents That Don’t Start Over When They Fail","url":"https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp"},"datePublished":"2026-08-20T18:38:21.099Z","dateModified":"2026-09-14T08:45:52.573Z","description":"Long-running AI agent workflows fail differently than simple background jobs: LLM calls are slow and expensive, outputs are non-deterministic, side effects...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e4bf4de362fac74159839398705cbfa5?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/e4bf4de362fac74159839398705cbfa5?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"System Design Newsletter","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"System Design Newsletter","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/eec821422ebe4672a738b8fbfa2c0dbd","url":"https://daily.dev/sources/systemdesignnews"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"ai-agents,workflow-orchestration","timeRequired":"PT31M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"System Design Newsletter","item":"https://daily.dev/sources/systemdesignnews"},{"@type":"ListItem","position":3,"name":"How to Build AI Agents That Don’t Start Over When They Fail"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp","comment":[{"@type":"Comment","text":"Checkpointing doesn’t make side effects safe by itself. The dangerous boundary is a write succeeding before its result is recorded, which still needs idempotency keys or compensation","datePublished":"2026-08-21T20:25:04.844Z","url":"https://daily.dev/posts/hDtNQPnLp#c-erj3aiYzX","author":{"@type":"Person","name":"Agustin Barrientos","url":"https://daily.dev/agustinbarrientos","image":"https://media.daily.dev/image/upload/s--5ayxQnqn--/f_auto/v1788281802/avatars/avatar_wQYYVe5Tbj0NJ7C7qPoa8?_a=BAMAMicg0"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp#faq","mainEntity":[{"@type":"Question","name":"How is durable execution different from just retrying a failed job?","acceptedAnswer":{"@type":"Answer","text":"A retry reruns an entire failed task or workflow from the start, while durable execution (checkpointing) remembers which steps already completed successfully and only reruns the failed step. For a multi-step AI agent that retrieved documents, extracted evidence, and then failed during drafting, durable execution skips the retrieval and extraction and resumes at drafting, saving time, tokens, and avoiding duplicate side effects like re-sent emails. daily.dev surfaces this kind of workflow-recovery reasoning for engineers designing resilient agent pipelines."}},{"@type":"Question","name":"When should I choose BullMQ, Temporal, or Inngest for building a durable AI agent workflow?","acceptedAnswer":{"@type":"Answer","text":"BullMQ fits teams with a dedicated infrastructure team who want fine-grained control over queues, workers, and retries but are willing to build checkpointing and resumability themselves. Temporal suits highly complex workflow graphs with many interdependent workflows, at the cost of added operational complexity. Inngest keeps workflow logic in application code via step.run(), handling checkpointing, retries, and waiting automatically, trading some low-level control for simplicity. developers weighing queue-first versus workflow-first tools for agent infrastructure can track these tradeoffs on daily.dev."}},{"@type":"Question","name":"How does step.waitForEvent() let an AI agent workflow pause for human approval without wasting compute?","acceptedAnswer":{"@type":"Answer","text":"step.waitForEvent() in Inngest pauses a workflow's execution and releases its compute while waiting for a matching external event, such as an editor approving a draft, instead of holding a worker process open. A timeout can be configured for cases where the event never arrives, and once the event lands, the workflow resumes with its prior progress intact, even after waiting for days. teams building human-in-the-loop agent approvals can compare these waiting patterns on daily.dev."}}]}
```

