<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/measuring-llm-lies-oho7myynu" -->

---
title: Measuring LLM Lies | daily.dev
description: A benchmark called &#x27;BS Bench&#x27; tests LLMs by asking them nonsense questions where the premise is logically incoherent (e.g., relating fire safety codes to curry...
canonical: https://daily.dev/posts/measuring-llm-lies-oho7myynu
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Measuring LLM Lies | daily.dev
og:description: A benchmark called &#x27;BS Bench&#x27; tests LLMs by asking them nonsense questions where the premise is logically incoherent (e.g., relating fire safety codes to curry...
og:url: https://daily.dev/posts/measuring-llm-lies-oho7myynu
og:image: https://api.daily.dev/og/posts/oHO7myYnu.png
og:image:alt: Measuring LLM Lies
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Measuring LLM Lies

**[ThePrimeTime](https://daily.dev/sources/primeagen)** · 11 min read · 20 upvotes · 3 comments

## Summary

A benchmark called 'BS Bench' tests LLMs by asking them nonsense questions where the premise is logically incoherent (e.g., relating fire safety codes to curry recipes). Claude models generally refuse to answer such questions, while OpenAI and Google models tend to confidently fabricate detailed answers. Gemini 2.5 (nicknamed 'Kimmy K') surprisingly outperforms OpenAI and Google on pushback. The deeper concern raised is that LLMs act as skill multipliers — meaning engineers with poor judgment who use AI confidently will make bad decisions faster and at greater scale. The real danger isn't obviously nonsense questions but subtly flawed ones that AI answers without pushback.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.youtube.com/watch?v=QTf9RKMGAuI>

## Community discussion

Top comments from developers on daily.dev.

**@baz14** · 3 upvotes

> LLM's biggest problem is its overconfidence. The developers make it so that LLMS tries to give a confident answer even without finding any information.

**@confidentcoding** · 3 upvotes

> This study from back in 2024 was prescient: [ChatGPT is bullshit](https://link.springer.com/article/10.1007/s10676-024-09775-5)
>
> The more we can focus on these tools as a [delegation layer](https://cheewebdevelopment.com/dont-vibe-code-delegate-responsible-development-with-llms/), the better off we'll be when it comes to managing their constant stream of BS.

**@ktg\_one** · 1 upvotes

> [https://medium.com/@ktg.one/all-your-agent-skills-are-broken-8cab4770ccb6](https://medium.com/@ktg.one/all-your-agent-skills-are-broken-8cab4770ccb6)
>
>
> I'm mapping it and will need help passing out the survey soon

## Similar posts on daily.dev

- [Google DeepMind wants to know if chatbots are just virtue signaling](https://daily.dev/posts/google-deepmind-wants-to-know-if-chatbots-are-just-virtue-signaling-y0n7mzotz) · MIT Technology Review · 1 upvotes · 0 comments
- [Researchers discover a shortcoming that makes LLMs less reliable](https://daily.dev/posts/researchers-discover-a-shortcoming-that-makes-llms-less-reliable-wiiiapmcn) · MIT News · 2 upvotes · 0 comments
- [Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence](https://daily.dev/posts/butter-bench-evaluating-llm-controlled-robots-for-practical-intelligence-enmkuqbyc) · Hacker News · 3 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#chatgpt](https://daily.dev/tags/chatgpt), [#claude](https://daily.dev/tags/claude)

[View this post on daily.dev](https://daily.dev/posts/measuring-llm-lies-oho7myynu)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Measuring LLM Lies","url":"https://daily.dev/posts/measuring-llm-lies-oho7myynu","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/measuring-llm-lies-oho7myynu"},"datePublished":"2026-03-03T13:01:55.933Z","dateModified":"2026-03-03T13:02:18.976Z","description":"A benchmark called 'BS Bench' tests LLMs by asking them nonsense questions where the premise is logically incoherent (e.g., relating fire safety codes to curry...","image":"https://i.ytimg.com/vi/QTf9RKMGAuI/sddefault.jpg","thumbnailUrl":"https://i.ytimg.com/vi/QTf9RKMGAuI/sddefault.jpg","isAccessibleForFree":true,"articleSection":"ThePrimeTime","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"ThePrimeTime","logo":"https://media.daily.dev/image/upload/s--J5-GjDeo--/f_auto/v1704628081/logos/primeagen.jpg","url":"https://daily.dev/sources/primeagen"},"commentCount":3,"discussionUrl":"https://daily.dev/posts/measuring-llm-lies-oho7myynu","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":20},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":3}],"keywords":"llm,chatgpt,claude","timeRequired":"PT11M","video":{"@type":"VideoObject","name":"Measuring LLM Lies","description":"A benchmark called 'BS Bench' tests LLMs by asking them nonsense questions where the premise is logically incoherent (e.g., relating fire safety codes to curry...","thumbnailUrl":"https://i.ytimg.com/vi/QTf9RKMGAuI/sddefault.jpg","uploadDate":"2026-03-03T13:01:55.933Z","duration":"PT11M","url":"https://api.daily.dev/r/oHO7myYnu","embedUrl":"https://www.youtube.com/embed/QTf9RKMGAuI"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"ThePrimeTime","item":"https://daily.dev/sources/primeagen"},{"@type":"ListItem","position":3,"name":"Measuring LLM Lies"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/measuring-llm-lies-oho7myynu","comment":[{"@type":"Comment","text":"LLM’s biggest problem is its overconfidence. The developers make it so that LLMS tries to give a confident answer even without finding any information.","datePublished":"2026-03-04T04:42:35.609Z","url":"https://daily.dev/posts/oHO7myYnu#c-ckoTs7X53","author":{"@type":"Person","name":"Baz","url":"https://daily.dev/baz14","image":"https://lh3.googleusercontent.com/a/ACg8ocKd2iFU8YS4sV-rBJptuYSF8EHeQnm23N_dcXE1uON7_GMIlZg=s96-c"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3}},{"@type":"Comment","text":"This study from back in 2024 was prescient: ChatGPT is bullshit\nThe more we can focus on these tools as a delegation layer, the better off we’ll be when it comes to managing their constant stream of BS.","datePublished":"2026-03-03T20:30:04.697Z","dateModified":"2026-03-03T20:30:15.051Z","url":"https://daily.dev/posts/oHO7myYnu#c-ito7w5MSX","author":{"@type":"Person","name":"Lars Faye | Confident Coding","url":"https://daily.dev/confidentcoding","image":"https://media.daily.dev/image/upload/s--OGZu5DEc--/f_auto/v1772569630/avatars/avatar_umWZ9aQAng34qk5aaJl2q?_a=BAMAMiiu0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":3}},{"@type":"Comment","text":"https://medium.com/@ktg.one/all-your-agent-skills-are-broken-8cab4770ccb6\nI’m mapping it and will need help passing out the survey soon","datePublished":"2026-03-26T06:06:22.203Z","url":"https://daily.dev/posts/oHO7myYnu#c-EL7kAq6YM","author":{"@type":"Person","name":"ktg","url":"https://daily.dev/ktg_one","image":"https://media.daily.dev/image/upload/s--jXAtJ1Oj--/f_auto/v1758670788/avatars/avatar_PGHQtneEfnMYow62H8Si2?_a=BAMAK+ZW0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}}]}
```

