<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k" -->

---
title: LangSmith&#x27;s New Tuned Evaluators: Automatic Error...
description: LangChain launched Tuned Evaluators in LangSmith, starting with a &#x27;Perceived Error&#x27; evaluator that flags production agent conversations where the agent...
canonical: https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: LangSmith&#x27;s New Tuned Evaluators: Automatic Error Detection for Production Agents | daily.dev
og:description: LangChain launched Tuned Evaluators in LangSmith, starting with a &#x27;Perceived Error&#x27; evaluator that flags production agent conversations where the agent...
og:url: https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k
og:image: https://api.daily.dev/og/posts/sM1ci948K.png
og:image:alt: LangSmith&#x27;s New Tuned Evaluators: Automatic Error Detection for Production Agents
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# LangSmith's New Tuned Evaluators: Automatic Error Detection for Production Agents

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

LangChain launched Tuned Evaluators in LangSmith, starting with a 'Perceived Error' evaluator that flags production agent conversations where the agent misunderstood the user or left issues unresolved. It uses a specialized post-trained model rather than a generic frontier LLM-as-judge, claiming better accuracy while cutting evaluation costs by 82-98% in some early-partner cases. Threads need at least two human-AI message pairs, evaluation triggers after an idle period, and results appear within 12 hours. It's now live for Plus and Cloud Enterprise plans in the US, billed per successful evaluation, with Vanta already using it as a safety net alongside their own custom evaluators.

## Content

LangChain has added Tuned Evaluators to LangSmith, giving teams a way to get quality feedback on production agent traces without writing, tuning, or hosting their own LLM-as-judge setups. You just turn it on in a tracing project.

The first evaluator available is called **Perceived Error**. It flags conversations where the agent messed up, misread what the user wanted, or left something unresolved. Instead of relying on a generic frontier model prompted to act as a judge, LangChain built a specialized post-trained model for this specific task. According to their numbers, it beats frontier model accuracy while cutting evaluation costs by up to 82% - and in some early-partner workloads, that number climbs to 98%.

A few practical details on how it works:

- Threads need at least two human-AI message pairs before they're eligible for evaluation
- Evaluation kicks in once the conversation hits an idle period
- Results show up within 12 hours
- It's live now for Plus and Cloud Enterprise plans in the US, billed per successful evaluation

Vanta is using it early on, treating Perceived Error as a safety net while they build out their own custom evaluators on top of it.

The pitch here is straightforward: most teams don't want to spend engineering time building and maintaining judge models just to catch agent mistakes. A versioned, pre-tuned evaluator you can flip on saves that work, and if the accuracy claims hold up in practice, an 82-98% cost reduction on evaluation is hard to ignore for anyone running agents at scale.

## Questions this post answers

### What is the Perceived Error evaluator in LangSmith and how does it work?

Perceived Error is a Tuned Evaluator in LangSmith that automatically flags production agent conversations where the agent misunderstood the user or left something unresolved. It uses a specialized post-trained model instead of a generic frontier LLM-as-judge, requires at least two human-AI message pairs per thread, triggers after an idle period, and returns results within 12 hours.

_daily.dev surfaces releases like this for teams evaluating agent quality tooling._

### How much cheaper is LangSmith's Tuned Evaluator compared to using a frontier model as an LLM judge?

LangSmith's Perceived Error evaluator cuts evaluation costs by up to 82% compared to using a frontier model as an LLM-as-judge, and in some early-partner workloads that reduction reaches as high as 98%, while reportedly matching or beating frontier model accuracy on detecting agent errors.

_Track cost and accuracy trade-offs like these when picking agent evaluation tools on daily.dev._

### Which LangSmith plans support the Perceived Error tuned evaluator and how is it billed?

The Perceived Error evaluator is available now for LangSmith Plus and Cloud Enterprise plans in the US, billed per successful evaluation rather than a flat subscription fee. It only evaluates threads with at least two human-AI message pairs and runs once a conversation goes idle.

_Developers weighing agent observability plans can follow rollout details like these on daily.dev._

## Similar posts on daily.dev

- [Improve agent quality with Insights Agent and Multi-turn Evals, now in LangSmith](https://daily.dev/posts/improve-agent-quality-with-insights-agent-and-multi-turn-evals-now-in-langsmith-rxlr5kvds) · LangChain · 2 upvotes · 0 comments
- [LangSmith Engine Improves Agent Issue Detection by 2x](https://daily.dev/posts/langsmith-engine-improves-agent-issue-detection-by-2x-vgllisrj4) · LangChain · 2 upvotes · 0 comments

---

Tags: [#ai-agents](https://daily.dev/tags/ai-agents), [#observability](https://daily.dev/tags/observability), [#langchain](https://daily.dev/tags/langchain), [#langsmith](https://daily.dev/tags/langsmith)

[View this post on daily.dev](https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"LangSmith's New Tuned Evaluators: Automatic Error Detection for Production Agents","url":"https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k"},"datePublished":"2026-08-18T20:22:00.629Z","dateModified":"2026-08-18T20:22:37.296Z","description":"LangChain launched Tuned Evaluators in LangSmith, starting with a 'Perceived Error' evaluator that flags production agent conversations where the agent...","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai-agents,observability,langchain,langsmith","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"LangSmith's New Tuned Evaluators: Automatic Error Detection for Production Agents"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/langsmith-s-new-tuned-evaluators-automatic-error-detection-for-production-agents-sm1ci948k#faq","mainEntity":[{"@type":"Question","name":"What is the Perceived Error evaluator in LangSmith and how does it work?","acceptedAnswer":{"@type":"Answer","text":"Perceived Error is a Tuned Evaluator in LangSmith that automatically flags production agent conversations where the agent misunderstood the user or left something unresolved. It uses a specialized post-trained model instead of a generic frontier LLM-as-judge, requires at least two human-AI message pairs per thread, triggers after an idle period, and returns results within 12 hours. daily.dev surfaces releases like this for teams evaluating agent quality tooling."}},{"@type":"Question","name":"How much cheaper is LangSmith's Tuned Evaluator compared to using a frontier model as an LLM judge?","acceptedAnswer":{"@type":"Answer","text":"LangSmith's Perceived Error evaluator cuts evaluation costs by up to 82% compared to using a frontier model as an LLM-as-judge, and in some early-partner workloads that reduction reaches as high as 98%, while reportedly matching or beating frontier model accuracy on detecting agent errors. Track cost and accuracy trade-offs like these when picking agent evaluation tools on daily.dev."}},{"@type":"Question","name":"Which LangSmith plans support the Perceived Error tuned evaluator and how is it billed?","acceptedAnswer":{"@type":"Answer","text":"The Perceived Error evaluator is available now for LangSmith Plus and Cloud Enterprise plans in the US, billed per successful evaluation rather than a flat subscription fee. It only evaluates threads with at least two human-AI message pairs and runs once a conversation goes idle. Developers weighing agent observability plans can follow rollout details like these on daily.dev."}}]}
```

