<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr" -->

---
title: Why You Should Not Use Numeric Evals For LLM As a Judge
description: Using LLMs to conduct numeric evaluations is finicky and unreliable. Small changes in prompt templates and switching between models can lead to vastly...
canonical: https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Why You Should Not Use Numeric Evals For LLM As a Judge | daily.dev
og:description: Using LLMs to conduct numeric evaluations is finicky and unreliable. Small changes in prompt templates and switching between models can lead to vastly...
og:url: https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr
og:image: https://api.daily.dev/og/posts/WggkjApLR.png
og:image:alt: Why You Should Not Use Numeric Evals For LLM As a Judge
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Why You Should Not Use Numeric Evals For LLM As a Judge

**[Towards Data Science](https://daily.dev/sources/tds)** · 7 min read · 0 upvotes · 0 comments

## Summary

Using LLMs to conduct numeric evaluations is finicky and unreliable. Small changes in prompt templates and switching between models can lead to vastly different results. LLMs are often inconsistent in their responses, making it hard to rely on them as reliable arbiters of numeric evaluation criteria.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://towardsdatascience.com/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-bf22424f5379>

---

Tags: [#nlp](https://daily.dev/tags/nlp), [#llm](https://daily.dev/tags/llm), [#text-generation](https://daily.dev/tags/text-generation), [#llmops](https://daily.dev/tags/llmops)

[View this post on daily.dev](https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Why You Should Not Use Numeric Evals For LLM As a Judge","url":"https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr"},"datePublished":"2024-03-08T16:15:23.965Z","dateModified":"2024-05-09T09:28:05.588Z","description":"Using LLMs to conduct numeric evaluations is finicky and unreliable. Small changes in prompt templates and switching between models can lead to vastly...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c4f6c20636982d314642bdbd7126862b?_a=AQAEufR","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/c4f6c20636982d314642bdbd7126862b?_a=AQAEufR","isAccessibleForFree":true,"articleSection":"Towards Data Science","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Towards Data Science","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/tds","url":"https://daily.dev/sources/tds"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/why-you-should-not-use-numeric-evals-for-llm-as-a-judge-wggkjaplr","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"nlp,llm,text-generation,llmops","timeRequired":"PT7M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Towards Data Science","item":"https://daily.dev/sources/tds"},{"@type":"ListItem","position":3,"name":"Why You Should Not Use Numeric Evals For LLM As a Judge"}]}
```

