<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze" -->

---
title: AI Performance Testing: How to Scale Agentic AI | daily.dev
description: A principal architect and performance engineering veteran discusses how to performance test agentic AI applications at enterprise scale, drawing on his free...
canonical: https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: AI Performance Testing: How to Scale Agentic AI | daily.dev
og:description: A principal architect and performance engineering veteran discusses how to performance test agentic AI applications at enterprise scale, drawing on his free...
og:url: https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze
og:image: https://api.daily.dev/og/posts/fESlD6Xze.png
og:image:alt: AI Performance Testing: How to Scale Agentic AI
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Performance Testing: How to Scale Agentic AI

**[Automation Testing with Joe Colantonio](https://daily.dev/sources/joecolantonio)** · 43 min read · 2 upvotes · 0 comments

## Summary

A principal architect and performance engineering veteran discusses how to performance test agentic AI applications at enterprise scale, drawing on his free book 'Rethinking Performance Engineering for Agentic AI.' Topics include handling non-deterministic agent behavior, setting outcome-based SLOs at the trace level instead of span level, implementing 'harness' guardrails to bound agent tool calls and reasoning loops, shifting performance testing left into CI pipelines and right into production via synthetic monitoring, using latency stubbing sampled from production traces for load testing, and cutting token costs through prompt caching, model routing based on query complexity, and summarizing data instead of feeding raw JSON between steps. He also covers tooling choices (JMeter, Langfuse, Arize) and the importance of streaming responses to improve perceived performance.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=v9ULBQYcW80>

## Questions this post answers

### How can I reduce LLM token costs for an AI agent in production?

Token costs can be cut through several layered optimizations: prompt caching (cached tokens cost about 10% versus 100% for regular tokens), summarizing tool call outputs instead of feeding raw JSON data back into the context, removing duplicated system instructions, and routing simple queries to cheaper, smaller models while reserving complex reasoning tasks for premium models, which can cut costs by 30-50%.

_Anyone tuning agent spend against quality tracks these token optimization patterns on daily.dev._

### What is an outcome-based SLO for an AI agent and how is it different from span-level monitoring?

An outcome-based SLO measures whether an entire trace completes a workflow within defined bounds, such as 95% of traces finishing within five steps and under a set token budget per trace, rather than checking individual spans. This matters because token context can grow unnoticed between spans with no compaction strategy, so trace-level completion is a better signal of whether the agent behaved as expected.

_Teams defining SLOs for non-deterministic agent workflows can follow this kind of framing on daily.dev._

### How do you performance test a non-deterministic AI agent before production?

Testing relies on a harness that bounds agent behavior with guardrails like a maximum number of tool calls, reasoning steps, and tokens per trace, combined with shift-left regression testing in lower environments (checking whether the same input still completes in the same number of steps) and shift-right synthetic monitoring in production after deployments, plus resiliency tests that inject faults like a dependency service going down to check for retry storms.

_Engineers building test strategy for non-deterministic agents can dig into this approach on daily.dev._

## Similar posts on daily.dev

- [Building the enterprise environment for agentic AI](https://daily.dev/posts/building-the-enterprise-environment-for-agentic-ai-ymntx9dtl) · MIT Technology Review · 0 upvotes · 0 comments
- [9 Tips for Reducing API Latency in Agentic AI Systems](https://daily.dev/posts/9-tips-for-reducing-api-latency-in-agentic-ai-systems-pc5wpyx5a) · Nordic APIs · 1 upvotes · 0 comments
- [Agentic AI in the enterprise: How to balance autonomy with constraints](https://daily.dev/posts/agentic-ai-in-the-enterprise-how-to-balance-autonomy-with-constraints-siujtbxge) · InfoWorld · 1 upvotes · 1 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#observability](https://daily.dev/tags/observability), [#agentic-ai](https://daily.dev/tags/agentic-ai)

[View this post on daily.dev](https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"AI Performance Testing: How to Scale Agentic AI","url":"https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze"},"datePublished":"2026-09-08T19:55:23.216Z","dateModified":"2026-09-14T07:10:05.712Z","description":"A principal architect and performance engineering veteran discusses how to performance test agentic AI applications at enterprise scale, drawing on his free...","image":"https://i.ytimg.com/vi/v9ULBQYcW80/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/v9ULBQYcW80/sddefault.jpg","isAccessibleForFree":true,"articleSection":"Automation Testing with Joe Colantonio","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Automation Testing with Joe Colantonio","logo":"https://media.daily.dev/image/upload/s--euSb6E7Q--/f_auto,q_auto/v1780213680/logos/joecolantonio?_a=BAMAMiWQ0","url":"https://daily.dev/sources/joecolantonio"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,observability,agentic-ai","timeRequired":"PT43M","video":{"@type":"VideoObject","name":"AI Performance Testing: How to Scale Agentic AI","description":"A principal architect and performance engineering veteran discusses how to performance test agentic AI applications at enterprise scale, drawing on his free...","thumbnailUrl":"https://i.ytimg.com/vi/v9ULBQYcW80/sddefault.jpg","uploadDate":"2026-09-08T19:55:23.216Z","duration":"PT43M","url":"https://api.daily.dev/r/fESlD6Xze","embedUrl":"https://www.youtube.com/embed/v9ULBQYcW80"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Automation Testing with Joe Colantonio","item":"https://daily.dev/sources/joecolantonio"},{"@type":"ListItem","position":3,"name":"AI Performance Testing: How to Scale Agentic AI"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/ai-performance-testing-how-to-scale-agentic-ai-fesld6xze#faq","mainEntity":[{"@type":"Question","name":"How can I reduce LLM token costs for an AI agent in production?","acceptedAnswer":{"@type":"Answer","text":"Token costs can be cut through several layered optimizations: prompt caching (cached tokens cost about 10% versus 100% for regular tokens), summarizing tool call outputs instead of feeding raw JSON data back into the context, removing duplicated system instructions, and routing simple queries to cheaper, smaller models while reserving complex reasoning tasks for premium models, which can cut costs by 30-50%. Anyone tuning agent spend against quality tracks these token optimization patterns on daily.dev."}},{"@type":"Question","name":"What is an outcome-based SLO for an AI agent and how is it different from span-level monitoring?","acceptedAnswer":{"@type":"Answer","text":"An outcome-based SLO measures whether an entire trace completes a workflow within defined bounds, such as 95% of traces finishing within five steps and under a set token budget per trace, rather than checking individual spans. This matters because token context can grow unnoticed between spans with no compaction strategy, so trace-level completion is a better signal of whether the agent behaved as expected. Teams defining SLOs for non-deterministic agent workflows can follow this kind of framing on daily.dev."}},{"@type":"Question","name":"How do you performance test a non-deterministic AI agent before production?","acceptedAnswer":{"@type":"Answer","text":"Testing relies on a harness that bounds agent behavior with guardrails like a maximum number of tool calls, reasoning steps, and tokens per trace, combined with shift-left regression testing in lower environments (checking whether the same input still completes in the same number of steps) and shift-right synthetic monitoring in production after deployments, plus resiliency tests that inject faults like a dependency service going down to check for retry storms. Engineers building test strategy for non-deterministic agents can dig into this approach on daily.dev."}}]}
```

