<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh" -->

---
title: How to Evaluate Voice Agents with LangSmith | daily.dev
description: Evaluating voice agents requires assessing three distinct dimensions: execution (did the agent follow its instructions?), outcome (did the interaction achieve...
canonical: https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How to Evaluate Voice Agents with LangSmith | daily.dev
og:description: Evaluating voice agents requires assessing three distinct dimensions: execution (did the agent follow its instructions?), outcome (did the interaction achieve...
og:url: https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh
og:image: https://api.daily.dev/og/posts/CzwZ0dYOH.png
og:image:alt: How to Evaluate Voice Agents with LangSmith
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How to Evaluate Voice Agents with LangSmith

**[LangChain](https://daily.dev/sources/langchain)** · 11 min read · 1 upvotes · 0 comments

## Summary

Evaluating voice agents requires assessing three distinct dimensions: execution (did the agent follow its instructions?), outcome (did the interaction achieve its goal?), and experience (was the conversation smooth for the caller?). For execution, deterministic code evaluators handle explicit rule checks like tool call order, while LLM judges handle semantic requirements like policy adherence. Outcome evaluation combines LLM judges for qualitative success with downstream business metrics like booking success rate or ticket reopen rate. Experience evaluation covers latency measurement across pipeline components (STT, inference, TTS), naturalness and clarity via audio-aware LLM judges, and conversational friction signals like repeated clarification loops or failed interruption recovery. LangSmith supports all of these through traces, annotation queues, and experiment comparison, enabling a continuous evaluation loop where changes to prompts, models, or workflows can be measured against consistent criteria across the same dataset.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.langchain.com/blog/how-to-evaluate-voice-agents-execution-outcomes-and-experience>

## Similar posts on daily.dev

- [LangSmith Explained: Debugging and Evaluating LLM Agents](https://daily.dev/posts/langsmith-explained-debugging-and-evaluating-llm-agents-wy5ceqqjb) · DigitalOcean Community · 1 upvotes · 0 comments
- [How Similarweb Evaluates Agent Reports with LangSmith](https://daily.dev/posts/how-similarweb-evaluates-agent-reports-with-langsmith-brldcch8t) · LangChain · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#observability](https://daily.dev/tags/observability), [#langsmith](https://daily.dev/tags/langsmith)

[View this post on daily.dev](https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How to Evaluate Voice Agents with LangSmith","url":"https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh"},"datePublished":"2026-08-04T20:48:55.201Z","dateModified":"2026-08-04T20:56:12.197Z","description":"Evaluating voice agents requires assessing three distinct dimensions: execution (did the agent follow its instructions?), outcome (did the interaction achieve...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/95453605c58fd1a02ff58d9040c48252?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/95453605c58fd1a02ff58d9040c48252?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"LangChain","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"LangChain","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/0f4f2629e0724a849823f8cd0d913e13","url":"https://daily.dev/sources/langchain"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-to-evaluate-voice-agents-with-langsmith-czwz0dyoh","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,observability,langsmith","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"LangChain","item":"https://daily.dev/sources/langchain"},{"@type":"ListItem","position":3,"name":"How to Evaluate Voice Agents with LangSmith"}]}
```

