<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah" -->

---
title: 5 Resilience Patterns for AI Agents | daily.dev
description: Production AI agents fail differently than single requests because runs are long, stateful, and take real-world actions. The piece lays out five resilience...
canonical: https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: 5 Resilience Patterns for AI Agents | daily.dev
og:description: Production AI agents fail differently than single requests because runs are long, stateful, and take real-world actions. The piece lays out five resilience...
og:url: https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah
og:image: https://api.daily.dev/og/posts/OiijbLpaH.png
og:image:alt: 5 Resilience Patterns for AI Agents
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# 5 Resilience Patterns for AI Agents

**[The T-Shaped Dev](https://daily.dev/sources/thetshaped-dev)** · 8 min read · 1 upvotes · 0 comments

## Summary

Production AI agents fail differently than single requests because runs are long, stateful, and take real-world actions. The piece lays out five resilience patterns: retry the failed step rather than the whole run and sort failures into time-fixable, model-fixable, or unfixable categories; make every write tool idempotent with a stable key; checkpoint state after each step so crashes mean resume rather than restart; fall back to the same model on a different cloud provider rather than a different model to avoid shared outages and eval drift; and bound every run by steps, tokens, and wall-clock time, escalating stuck runs to a human. Code examples use Python with the Inngest durable-execution SDK to implement step-level retries, checkpointing, and human-in-the-loop approval waits.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://thetshaped.dev/p/5-resilience-patterns-for-ai-agents-retries-checkpoints-idempotency>

## Questions this post answers

### Why do AI agent runs fail so often even when each model or tool call succeeds 98% of the time?

Failure compounds across steps: at 98% success per call, only about 67% of a 20-step agent run completes without a single failure, meaning roughly one in three runs hits a failure. Even at 99% success per call, one in five 20-step runs fails, which is why failure has to be treated as the main path rather than an edge case in agent design.

_Anyone architecting multi-step agent workflows can track failure-handling patterns like these through daily.dev._

### How do I make a tool call safe to retry in an AI agent workflow?

Give every tool that writes data an idempotency key built from stable IDs, such as a model's tool-call ID, rather than a key generated fresh inside the step. Stripe popularized this pattern for payment APIs and keeps idempotency keys for 24 hours; applying the same approach lets a retried step re-run safely without duplicating side effects like refunds or ticket creation.

_Developers wiring up idempotent tool calls can keep up with patterns like this via daily.dev._

### What is the difference between retrying an entire AI agent run versus retrying just a failed step?

Retrying the whole run re-executes and re-pays for every step that already succeeded, replaying their side effects like duplicate emails or refunds, while retrying only the failed step preserves completed work. Failures should be sorted into three categories: time-fixable errors like 429s or timeouts get retried later, model-fixable errors get sent back as a tool result for the model to correct, and unfixable errors like a revoked API key should fail the step immediately.

_Teams debugging flaky agent retries can follow step-level failure-handling approaches on daily.dev._

## Similar posts on daily.dev

- [The journey to shippable AI systems: Patterns that work](https://daily.dev/posts/the-journey-to-shippable-ai-systems-patterns-that-work-8nthoadqh) · Temporal · 3 upvotes · 0 comments
- [How to Build Reliable AI Agents: 5 Engineering Patterns](https://daily.dev/posts/how-to-build-reliable-ai-agents-5-engineering-patterns-gpdk9dujm) · Salesforce Engineering · 10 upvotes · 0 comments
- [How to Build AI Agents That Don’t Start Over When They Fail](https://daily.dev/posts/how-to-build-ai-agents-that-don-t-start-over-when-they-fail-hdtnqpnlp) · System Design Newsletter · 3 upvotes · 1 comments
- [Why Most AI Agents Fail in Production](https://daily.dev/posts/why-most-ai-agents-fail-in-production-gwr8e39ne) · Medium · 0 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#distributed-systems](https://daily.dev/tags/distributed-systems), [#inngest](https://daily.dev/tags/inngest)

[View this post on daily.dev](https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"5 Resilience Patterns for AI Agents","url":"https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah"},"datePublished":"2026-09-24T10:22:11.888Z","dateModified":"2026-09-24T10:22:39.136Z","description":"Production AI agents fail differently than single requests because runs are long, stateful, and take real-world actions. The piece lays out five resilience...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/54a12be2a5c7de8e5f8a3ea011f8d99b?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/54a12be2a5c7de8e5f8a3ea011f8d99b?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"The T-Shaped Dev","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"The T-Shaped Dev","logo":"https://media.daily.dev/image/upload/s--TaOz2Ogr--/f_auto,q_auto/v1774964066/logos/thetshaped-dev?_a=BAMAMiWQ0","url":"https://daily.dev/sources/thetshaped-dev"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,distributed-systems,inngest","timeRequired":"PT8M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"The T-Shaped Dev","item":"https://daily.dev/sources/thetshaped-dev"},{"@type":"ListItem","position":3,"name":"5 Resilience Patterns for AI Agents"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/5-resilience-patterns-for-ai-agents-oiijblpah#faq","mainEntity":[{"@type":"Question","name":"Why do AI agent runs fail so often even when each model or tool call succeeds 98% of the time?","acceptedAnswer":{"@type":"Answer","text":"Failure compounds across steps: at 98% success per call, only about 67% of a 20-step agent run completes without a single failure, meaning roughly one in three runs hits a failure. Even at 99% success per call, one in five 20-step runs fails, which is why failure has to be treated as the main path rather than an edge case in agent design. Anyone architecting multi-step agent workflows can track failure-handling patterns like these through daily.dev."}},{"@type":"Question","name":"How do I make a tool call safe to retry in an AI agent workflow?","acceptedAnswer":{"@type":"Answer","text":"Give every tool that writes data an idempotency key built from stable IDs, such as a model's tool-call ID, rather than a key generated fresh inside the step. Stripe popularized this pattern for payment APIs and keeps idempotency keys for 24 hours; applying the same approach lets a retried step re-run safely without duplicating side effects like refunds or ticket creation. Developers wiring up idempotent tool calls can keep up with patterns like this via daily.dev."}},{"@type":"Question","name":"What is the difference between retrying an entire AI agent run versus retrying just a failed step?","acceptedAnswer":{"@type":"Answer","text":"Retrying the whole run re-executes and re-pays for every step that already succeeded, replaying their side effects like duplicate emails or refunds, while retrying only the failed step preserves completed work. Failures should be sorted into three categories: time-fixable errors like 429s or timeouts get retried later, model-fixable errors get sent back as a tool result for the model to correct, and unfixable errors like a revoked API key should fail the step immediately. Teams debugging flaky agent retries can follow step-level failure-handling approaches on daily.dev."}}]}
```

