<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji" -->

---
title: Can Jev replace your LLM judge? Part 2: Testing harder...
description: A follow-up benchmark tests TypeSafe&#x27;s Jev model against GPT-OSS-120B as an LLM judge on 72 harder QA answers (English and Japanese) containing subtle factual...
canonical: https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Can Jev replace your LLM judge? Part 2: Testing harder answers | daily.dev
og:description: A follow-up benchmark tests TypeSafe&#x27;s Jev model against GPT-OSS-120B as an LLM judge on 72 harder QA answers (English and Japanese) containing subtle factual...
og:url: https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji
og:image: https://api.daily.dev/og/posts/Su5BvtmJI.png
og:image:alt: Can Jev replace your LLM judge? Part 2: Testing harder answers
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Can Jev replace your LLM judge? Part 2: Testing harder answers

**[mlflow](https://daily.dev/sources/MLflow)** · 5 min read · 0 upvotes · 0 comments

## Summary

A follow-up benchmark tests TypeSafe's Jev model against GPT-OSS-120B as an LLM judge on 72 harder QA answers (English and Japanese) containing subtle factual errors. Jev matched human labels on 64/72 while GPT-OSS-120B matched all 72, but Jev was about seven times faster (0.20s vs 1.44s median latency) and far cheaper (~$0.020 per 1,000 judgments). Jev's failures were near-correct answers with one material error, such as reversed clause meanings, missing permissions, or a correct explanation paired with a wrong command. Routing Jev's uncertain judgments (probability 0.2-0.8) to GPT-OSS for a second opinion caught all errors while requiring escalation on fewer than 20% of cases. The upcoming MLflow 3.17 release will support Jev via a typesafe:/ model URI in make_judge for building custom and built-in scorers.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://mlflow.org/blog/jev-llm-judge-part-2>

## Questions this post answers

### How accurate is TypeSafe's Jev model compared to GPT-OSS-120B when used as an LLM judge for QA correctness?

Jev (typesafe/jev-1.13) agreed with human labels on 64 of 72 harder QA answers in each of two benchmark runs, while GPT-OSS-120B agreed on all 72. Jev's errors were near-correct answers with one material mistake, such as reversed clause meanings or missing permissions, but it accepted every correct answer including paraphrased ones and ran about seven times faster at roughly $0.020 per 1,000 judgments.

_daily.dev surfaces benchmarks like this for teams weighing judge model speed against accuracy._

### How can I use judge confidence scores to decide when to escalate an LLM judgment to a stronger model?

Route judgments with probability scores between 0.2 and 0.8 to a stronger secondary model for a second opinion. In testing, every false acceptance from Jev scored between roughly 0.5 and 0.8 while every correct answer scored at least 0.91, so escalating that middle range caught all errors while requiring a second check on fewer than 20% of judgments overall.

_daily.dev helps engineers track evaluation techniques like confidence-based routing for LLM judges._

### How will MLflow support the Jev model for LLM evaluation judges?

MLflow 3.17, not yet released, will add support for Jev in built-in scorers and custom judges through a typesafe:/ model URI usable with make_judge. Setting a TYPESAFE_API_KEY environment variable lets you define a boolean correctness judge, with MLflow recording Jev's true/false verdict as jev_correctness and its probability under typesafe.probability metadata.

_daily.dev keeps developers current on MLflow releases before they land in production pipelines._

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#genai](https://daily.dev/tags/genai)

[View this post on daily.dev](https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Can Jev replace your LLM judge? Part 2: Testing harder answers","url":"https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji"},"datePublished":"2026-09-30T15:55:05.859Z","dateModified":"2026-10-01T07:02:27.918Z","description":"A follow-up benchmark tests TypeSafe's Jev model against GPT-OSS-120B as an LLM judge on 72 harder QA answers (English and Japanese) containing subtle factual...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8125d0b467aa95e92a0763f0de50bef9?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/8125d0b467aa95e92a0763f0de50bef9?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"mlflow","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"mlflow","logo":"https://media.daily.dev/image/upload/s--iGp2NwSt--/f_auto/v1743317958/logos/MLflow","url":"https://daily.dev/sources/MLflow"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,llm,genai","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"mlflow","item":"https://daily.dev/sources/MLflow"},{"@type":"ListItem","position":3,"name":"Can Jev replace your LLM judge? Part 2: Testing harder answers"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/can-jev-replace-your-llm-judge-part-2-testing-harder-answers-su5bvtmji#faq","mainEntity":[{"@type":"Question","name":"How accurate is TypeSafe's Jev model compared to GPT-OSS-120B when used as an LLM judge for QA correctness?","acceptedAnswer":{"@type":"Answer","text":"Jev (typesafe/jev-1.13) agreed with human labels on 64 of 72 harder QA answers in each of two benchmark runs, while GPT-OSS-120B agreed on all 72. Jev's errors were near-correct answers with one material mistake, such as reversed clause meanings or missing permissions, but it accepted every correct answer including paraphrased ones and ran about seven times faster at roughly $0.020 per 1,000 judgments. daily.dev surfaces benchmarks like this for teams weighing judge model speed against accuracy."}},{"@type":"Question","name":"How can I use judge confidence scores to decide when to escalate an LLM judgment to a stronger model?","acceptedAnswer":{"@type":"Answer","text":"Route judgments with probability scores between 0.2 and 0.8 to a stronger secondary model for a second opinion. In testing, every false acceptance from Jev scored between roughly 0.5 and 0.8 while every correct answer scored at least 0.91, so escalating that middle range caught all errors while requiring a second check on fewer than 20% of judgments overall. daily.dev helps engineers track evaluation techniques like confidence-based routing for LLM judges."}},{"@type":"Question","name":"How will MLflow support the Jev model for LLM evaluation judges?","acceptedAnswer":{"@type":"Answer","text":"MLflow 3.17, not yet released, will add support for Jev in built-in scorers and custom judges through a typesafe:/ model URI usable with make_judge. Setting a TYPESAFE_API_KEY environment variable lets you define a boolean correctness judge, with MLflow recording Jev's true/false verdict as jev_correctness and its probability under typesafe.probability metadata. daily.dev keeps developers current on MLflow releases before they land in production pipelines."}}]}
```

