<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc" -->

---
title: Testing AI prompts and comparing models with promptfoo
description: Prompts in AI-integrated applications can produce unexpected results, and manual testing doesn&#x27;t scale. Promptfoo automates prompt evaluation by running test...
canonical: https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Testing AI prompts and comparing models with promptfoo | daily.dev
og:description: Prompts in AI-integrated applications can produce unexpected results, and manual testing doesn&#x27;t scale. Promptfoo automates prompt evaluation by running test...
og:url: https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc
og:image: https://api.daily.dev/og/posts/mXhomPHpC.png
og:image:alt: Testing AI prompts and comparing models with promptfoo
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Testing AI prompts and comparing models with promptfoo

**[Tim Deschryver](https://daily.dev/sources/timdeschryver)** · [@timdeschryver](https://daily.dev/timdeschryver) · 11 min read · 1 upvotes · 1 comments

## Summary

Prompts in AI-integrated applications can produce unexpected results, and manual testing doesn't scale. Promptfoo automates prompt evaluation by running test cases against configured models and checking responses with deterministic assertions (equals, contains, JSON schema, cost/latency) and model-graded assertions (tone, relevance, factual consistency). A practical example shows extracting structured JSON from customer messages, comparing two OpenAI GPT tiers side by side. Configurations can be written in YAML or TypeScript, prompts and fixtures can be loaded from files, and tasks can be organized by feature directory. Results are viewable in a terminal table or a local web UI showing outputs, assertion results, token usage, and cost. Evaluations can also run in CI/CD pipelines to catch regressions before deployment. Alternatives like Langfuse and Microsoft.Extensions.AI.Evaluation are briefly noted.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://timdeschryver.dev/blog/testing-ai-prompts-and-comparing-models-with-promptfoo>

## Community discussion

Top comments from developers on daily.dev.

**@agustinbarrientos** · 0 upvotes

> I'll take a boring deterministic assertion over an LLM judge whenever I can.

## Similar posts on daily.dev

- [AI Chatbot Evaluation and Observability with Promptfoo and Langfuse](https://daily.dev/posts/ai-chatbot-evaluation-and-observability-with-promptfoo-and-langfuse-rccwse2wk) · Netguru · 0 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm)

[View this post on daily.dev](https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Testing AI prompts and comparing models with promptfoo","url":"https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc"},"datePublished":"2026-07-29T10:28:04.251Z","dateModified":"2026-07-30T07:00:52.236Z","description":"Prompts in AI-integrated applications can produce unexpected results, and manual testing doesn't scale. Promptfoo automates prompt evaluation by running test...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d5d9e71fb3e8694cc09465ce37f3d8b7?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d5d9e71fb3e8694cc09465ce37f3d8b7?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Tim Deschryver","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Person","name":"Tim Deschryver","url":"https://daily.dev/timdeschryver","image":"https://lh3.googleusercontent.com/a-/AOh14GhKsm-jzL_QCoil31Ak6dis1SFiy_ouaa8qXOPRYA=s100","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":100}},"commentCount":1,"discussionUrl":"https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"llm","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Tim Deschryver","item":"https://daily.dev/sources/timdeschryver"},{"@type":"ListItem","position":3,"name":"Testing AI prompts and comparing models with promptfoo"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/testing-ai-prompts-and-comparing-models-with-promptfoo-mxhomphpc","comment":[{"@type":"Comment","text":"I’ll take a boring deterministic assertion over an LLM judge whenever I can.","datePublished":"2026-08-02T20:55:13.173Z","url":"https://daily.dev/posts/mXhomPHpC#c-D5NhZDprB","author":{"@type":"Person","name":"Agustin Barrientos","url":"https://daily.dev/agustinbarrientos","image":"https://media.daily.dev/image/upload/s--5ayxQnqn--/f_auto/v1788281802/avatars/avatar_wQYYVe5Tbj0NJ7C7qPoa8?_a=BAMAMicg0"}}]}
```

