<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7" -->

---
title: Jev: System One models for Prod, not God — with Diogo...
description: A long-form podcast interview with Diogo Almeida, CEO of TypeSafe AI and co-author of the InstructGPT paper, covering the launch of Jev, a new class of &#x27;System...
canonical: https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI | daily.dev
og:description: A long-form podcast interview with Diogo Almeida, CEO of TypeSafe AI and co-author of the InstructGPT paper, covering the launch of Jev, a new class of &#x27;System...
og:url: https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7
og:image: https://api.daily.dev/og/posts/1c6lThXm7.png
og:image:alt: Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI

**[Latent Space](https://daily.dev/sources/latentspace)** · 160 min read · 0 upvotes · 0 comments

## Summary

A long-form podcast interview with Diogo Almeida, CEO of TypeSafe AI and co-author of the InstructGPT paper, covering the launch of Jev, a new class of 'System 1' models designed to be consumed by code rather than humans. Almeida argues that RLHF and RLVR are the wrong optimization targets for AI meant to be embedded in software, proposing a new unpublished technique called RLCD (Reinforcement Learning for Calibrated Decisions) that optimizes for epistemically honest probabilities. The conversation ranges across mode collapse in RLHF-tuned models, why safety refusals are a 'type error' for API-consumed models, TypeSafe's rejection of public benchmarks in favor of private evaluation, and the company's self-description as a 'data lab' rather than a model lab, all synthetic data, no user data training.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.latent.space/p/jev>

## Questions this post answers

### What is RLCD and how is it different from RLHF and RLVR?

Reinforcement Learning for Calibrated Decisions (RLCD) is an unpublished training technique from TypeSafe AI that optimizes for epistemically honest probabilities on System One tasks, rather than human-rated feedback (RLHF), which causes hallucinations and sycophancy, or programmatically verifiable rubric-based outputs (RLVR), which solves narrow tasks like Navier Stokes but worsens jagged intelligence and integrates poorly with software.

_Developers weighing post-training approaches for automation-focused models can track emerging techniques like RLCD on daily.dev._

### Why does TypeSafe AI reject public benchmarks for evaluating Jev?

TypeSafe AI's CEO calls public benchmarks extremely gameable, noting labs historically collected data resembling benchmarks like MMLU just to score better, which he calls benchmarking with extra steps. Instead the company favors private benchmarks used as proxies and believes trust in a model's intelligence should come from real workflow evaluation rather than published leaderboard numbers.

_Anyone deciding between AI models for production use can follow debates over benchmarking trust on daily.dev._

### Why do refusals cause problems when AI models are embedded inside software dependencies?

A refusal inside a dependency is described as a type error: if an AI model buried deep in a software stack refuses a request, the failure cascades unpredictably to downstream users who have no visibility into that system, causing software to stochastically break. This is why System One models built for programmatic, code-consumed use avoid safety-alignment-style refusals used in consumer chat products like ChatGPT.

_Teams building automation on top of AI models can follow discussions on refusal behavior and reliability on daily.dev._

---

Tags: [#llm](https://daily.dev/tags/llm), [#ai-agents](https://daily.dev/tags/ai-agents), [#reinforcement-learning](https://daily.dev/tags/reinforcement-learning)

[View this post on daily.dev](https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI","url":"https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7"},"datePublished":"2026-09-21T22:16:55.091Z","dateModified":"2026-09-21T23:16:25.147Z","description":"A long-form podcast interview with Diogo Almeida, CEO of TypeSafe AI and co-author of the InstructGPT paper, covering the launch of Jev, a new class of 'System...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/7a4bc0429fd1053b901c1181581af7ca?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/7a4bc0429fd1053b901c1181581af7ca?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Latent Space","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Latent Space","logo":"https://media.daily.dev/image/upload/s--DnJ9laFj--/f_auto,q_auto/v1780213196/logos/latentspace?_a=BAMAMiWQ0","url":"https://daily.dev/sources/latentspace"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"llm,ai-agents,reinforcement-learning","timeRequired":"PT160M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Latent Space","item":"https://daily.dev/sources/latentspace"},{"@type":"ListItem","position":3,"name":"Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/jev-system-one-models-for-prod-not-god-with-diogo-almeida-ceo-typesafe-ai-1c6lthxm7#faq","mainEntity":[{"@type":"Question","name":"What is RLCD and how is it different from RLHF and RLVR?","acceptedAnswer":{"@type":"Answer","text":"Reinforcement Learning for Calibrated Decisions (RLCD) is an unpublished training technique from TypeSafe AI that optimizes for epistemically honest probabilities on System One tasks, rather than human-rated feedback (RLHF), which causes hallucinations and sycophancy, or programmatically verifiable rubric-based outputs (RLVR), which solves narrow tasks like Navier Stokes but worsens jagged intelligence and integrates poorly with software. Developers weighing post-training approaches for automation-focused models can track emerging techniques like RLCD on daily.dev."}},{"@type":"Question","name":"Why does TypeSafe AI reject public benchmarks for evaluating Jev?","acceptedAnswer":{"@type":"Answer","text":"TypeSafe AI's CEO calls public benchmarks extremely gameable, noting labs historically collected data resembling benchmarks like MMLU just to score better, which he calls benchmarking with extra steps. Instead the company favors private benchmarks used as proxies and believes trust in a model's intelligence should come from real workflow evaluation rather than published leaderboard numbers. Anyone deciding between AI models for production use can follow debates over benchmarking trust on daily.dev."}},{"@type":"Question","name":"Why do refusals cause problems when AI models are embedded inside software dependencies?","acceptedAnswer":{"@type":"Answer","text":"A refusal inside a dependency is described as a type error: if an AI model buried deep in a software stack refuses a request, the failure cascades unpredictably to downstream users who have no visibility into that system, causing software to stochastically break. This is why System One models built for programmatic, code-consumed use avoid safety-alignment-style refusals used in consumer chat products like ChatGPT. Teams building automation on top of AI models can follow discussions on refusal behavior and reliability on daily.dev."}}]}
```

