<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y" -->

---
title: How Do You Build an LLM Evaluation Dataset from...
description: A practical guide to building LLM evaluation datasets from production traces using Opik, the open-source LLM observability platform from Comet. Covers which...
canonical: https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How Do You Build an LLM Evaluation Dataset from Production Traces? | daily.dev
og:description: A practical guide to building LLM evaluation datasets from production traces using Opik, the open-source LLM observability platform from Comet. Covers which...
og:url: https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y
og:image: https://api.daily.dev/og/posts/PMIbiNA7Y.png
og:image:alt: How Do You Build an LLM Evaluation Dataset from Production Traces?
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How Do You Build an LLM Evaluation Dataset from Production Traces?

**[HEARTBEAT](https://daily.dev/sources/hrb)** · 9 min read · 2 upvotes · 0 comments

## Summary

A practical guide to building LLM evaluation datasets from production traces using Opik, the open-source LLM observability platform from Comet. Covers which traces to prioritize (high-volume patterns, known failures, edge cases), what fields a dataset item needs (input, expected output, metadata), three approaches for handling missing ground truth (reference-free LLM-as-a-Judge metrics, human annotation, model-generated drafts reviewed by humans), code examples for inserting items via the Opik Python SDK, the distinction between datasets/metrics and Opik Test Suites, and signals that indicate a dataset needs refreshing (prompt changes, distribution shift, metric saturation).

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://heartbeat.comet.ml/how-do-you-build-an-llm-evaluation-dataset-from-production-traces-2cad3ad175b6>

## Questions this post answers

### What fields does a dataset item need to hold for LLM evaluation metrics to score against it?

An Opik dataset item stores three fields: input, an optional expected output, and optional metadata. Input must be the exact production prompt without paraphrasing. Expected output is required for heuristic metrics like Equals or LevenshteinRatio, but optional for LLM-as-a-Judge metrics like Hallucination or AnswerRelevance. For RAG applications, the item also needs retrieved context chunks for ContextRecall and ContextPrecision metrics.

_Teams wiring up eval pipelines can track dataset and metric design patterns like these on daily.dev._

### How do you handle missing ground truth when building an LLM evaluation dataset from production traces?

Three approaches work: reference-free LLM-as-a-Judge metrics like Hallucination, Moderation, and AnswerRelevance that score output without an expected answer; human annotation on a representative sample routed through annotation queues; or model-generated draft answers from a stronger model that a human then accepts, edits, or rejects. A hybrid of a small annotated core plus a larger reference-free set is the most common pattern in practice.

_Developers weighing ground-truth strategies for LLM testing can follow this kind of evaluation guidance on daily.dev._

### How often should you refresh an LLM evaluation dataset built from production traces?

A practical cadence is a monthly refresh of roughly 10 to 20 percent of the dataset, replacing the oldest items with recent traces. Three signals also trigger an earlier refresh: shipping a new prompt version or retrieval strategy, a material distribution shift in query types or user segments, and metric scores that stop moving because the dataset has saturated.

_Anyone maintaining eval datasets against a moving production system can track cadence practices like this on daily.dev._

## Similar posts on daily.dev

- [What Should You Test During an LLM Observability Trial?](https://daily.dev/posts/what-should-you-test-during-an-llm-observability-trial--xsom2haln) · HEARTBEAT · 0 upvotes · 0 comments
- [​Trace and Monitor Any AI/LLM App](https://daily.dev/posts/trace-and-monitor-any-ai-llm-app-1cyxkbjw9) · Daily Dose of Data Science \| Avi Chawla \| Substack · 6 upvotes · 0 comments
- [Step-by-Step Guide to Assessing LLM Observability Tools for AI Teams](https://daily.dev/posts/step-by-step-guide-to-assessing-llm-observability-tools-for-ai-teams-dujbgzbws) · PromptLayer Blog · 1 upvotes · 0 comments
- [AI Evaluation Engineering: Build a Production-Grade LLM Evaluation Platform from Scratch \[Full Handbook\]](https://daily.dev/posts/ai-evaluation-engineering-build-a-production-grade-llm-evaluation-platform-from-scratch-full-handb-ju9v0rogq) · freeCodeCamp · 6 upvotes · 0 comments
- [Using LLM Observability to Cut Costs and Improve Performance](https://daily.dev/posts/using-llm-observability-to-cut-costs-and-improve-performance-np55kgk3t) · finout · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#testing](https://daily.dev/tags/testing), [#observability](https://daily.dev/tags/observability), [#rag](https://daily.dev/tags/rag)

[View this post on daily.dev](https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How Do You Build an LLM Evaluation Dataset from Production Traces?","url":"https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y"},"datePublished":"2026-08-25T22:36:23.492Z","dateModified":"2026-09-14T07:55:59.365Z","description":"A practical guide to building LLM evaluation datasets from production traces using Opik, the open-source LLM observability platform from Comet. Covers which...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3c55c5a9fde9a6b314c2c090ac8f164a?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/3c55c5a9fde9a6b314c2c090ac8f164a?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"HEARTBEAT","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"HEARTBEAT","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/hrb","url":"https://daily.dev/sources/hrb"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,testing,observability,rag","timeRequired":"PT9M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"HEARTBEAT","item":"https://daily.dev/sources/hrb"},{"@type":"ListItem","position":3,"name":"How Do You Build an LLM Evaluation Dataset from Production Traces?"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-do-you-build-an-llm-evaluation-dataset-from-production-traces--pmibina7y#faq","mainEntity":[{"@type":"Question","name":"What fields does a dataset item need to hold for LLM evaluation metrics to score against it?","acceptedAnswer":{"@type":"Answer","text":"An Opik dataset item stores three fields: input, an optional expected output, and optional metadata. Input must be the exact production prompt without paraphrasing. Expected output is required for heuristic metrics like Equals or LevenshteinRatio, but optional for LLM-as-a-Judge metrics like Hallucination or AnswerRelevance. For RAG applications, the item also needs retrieved context chunks for ContextRecall and ContextPrecision metrics. Teams wiring up eval pipelines can track dataset and metric design patterns like these on daily.dev."}},{"@type":"Question","name":"How do you handle missing ground truth when building an LLM evaluation dataset from production traces?","acceptedAnswer":{"@type":"Answer","text":"Three approaches work: reference-free LLM-as-a-Judge metrics like Hallucination, Moderation, and AnswerRelevance that score output without an expected answer; human annotation on a representative sample routed through annotation queues; or model-generated draft answers from a stronger model that a human then accepts, edits, or rejects. A hybrid of a small annotated core plus a larger reference-free set is the most common pattern in practice. Developers weighing ground-truth strategies for LLM testing can follow this kind of evaluation guidance on daily.dev."}},{"@type":"Question","name":"How often should you refresh an LLM evaluation dataset built from production traces?","acceptedAnswer":{"@type":"Answer","text":"A practical cadence is a monthly refresh of roughly 10 to 20 percent of the dataset, replacing the oldest items with recent traces. Three signals also trigger an earlier refresh: shipping a new prompt version or retrieval strategy, a material distribution shift in query types or user segments, and metric scores that stop moving because the dataset has saturated. Anyone maintaining eval datasets against a moving production system can track cadence practices like this on daily.dev."}}]}
```

