<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x" -->

---
title: Mushroom hunting with LLMs: what can go wrong? | daily.dev
description: A benchmark tests several vision-capable LLMs (Gemini 3.6/3.7 Flash, GPT-5.6-Sol, Claude Fable 5.1, GLM-5.3-Flash, DeepSeek v4 Flash Vision) on identifying...
canonical: https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Mushroom hunting with LLMs: what can go wrong? | daily.dev
og:description: A benchmark tests several vision-capable LLMs (Gemini 3.6/3.7 Flash, GPT-5.6-Sol, Claude Fable 5.1, GLM-5.3-Flash, DeepSeek v4 Flash Vision) on identifying...
og:url: https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x
og:image: https://api.daily.dev/og/posts/GYoh1Yw8X.png
og:image:alt: Mushroom hunting with LLMs: what can go wrong?
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Mushroom hunting with LLMs: what can go wrong?

**[Quesma](https://daily.dev/sources/quesma)** · 19 min read · 6 upvotes · 3 comments

## Summary

A benchmark tests several vision-capable LLMs (Gemini 3.6/3.7 Flash, GPT-5.6-Sol, Claude Fable 5.1, GLM-5.3-Flash, DeepSeek v4 Flash Vision) on identifying mushroom species from photos, using the FungiTastic dataset of 55 edible, poisonous, and deadly species common in Poland. Gemini 3.7 Flash leads accuracy while GLM-5.3-Flash wins on cost efficiency; Claude Fable 5.1 and GPT-5.6-Sol underperform despite reputations as top vision models. Critically, several models mistake deadly species (death cap, funeral bell, splendid webcap) for edible or harmless look-alikes at rates of 15-27%, illustrating that AI mushroom identification carries real poisoning risk rather than just minor labeling errors.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://quesma.com/blog/mushroom-llm-vision>

## Questions this post answers

### Can I trust an AI chatbot to tell me if a wild mushroom is safe to eat?

No, current vision-capable LLMs make dangerous misidentification errors on deadly species. In a benchmark across 1,040 photos of 55 species, models like Claude Fable 5.1 and GPT-5.6-Sol sometimes mistook deadly mushrooms such as the death cap, funeral bell, and splendid webcap for edible look-alikes like chanterelles, at rates as high as 15-27%, the same error that kills real foragers.

_Anyone weighing AI tools against real safety risk can track model evaluation results like these on daily.dev._

### Which vision LLM is best at identifying mushroom species from photos?

Gemini 3.6 and 3.7 Flash scored highest on first-guess accuracy in a mushroom identification benchmark, outperforming Claude Fable 5.1 and GPT-5.6-Sol despite being cheaper. GLM-5.3-Flash offered the best cost efficiency. The test used 1,040 photos across 55 mushroom species from the FungiTastic dataset, asking each model for its top five species guesses.

_Developers comparing vision model accuracy and cost can follow benchmarks like this on daily.dev._

## Community discussion

Top comments from developers on daily.dev.

**@agustinbarrientos** · 0 upvotes

> I think dangerous false positives matter more than aggregate accuracy here. I'd report the refusal rate beside poisonous-as-edible errors so a cautious model doesn't look safe merely because it skips harder images.

**@yaireo** · 0 upvotes

> Well, you as a human have extra context LLM might not have been supplied with, which would have increased odds of success:
>
> 1. You know the location
> 2. You know the season (exact date)
> 3. You know the size relative to your hand (not all photos shows this)
> 4. You can view it in multiple angles, in 3D. you gave it a single photo
>
> I assume if LLM had at least access to first 2, it might have gotten it better

**@froggobytes** · 0 upvotes

> why are LLMs used for this task? Shouldn't we use proper already-defined deep neural networks?

## Similar posts on daily.dev

- [Knowledge vs wisdom: asking AI “What mushroom is that?”](https://daily.dev/posts/knowledge-vs-wisdom-asking-ai-what-mushroom-is-that--qbivotact) · Quesma · 0 upvotes · 0 comments
- [Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier](https://daily.dev/posts/gemini-3-7-flash-grok-4-6-glm-5-3-and-deepseek-v4-pro-joined-the-frontier-pfyczsmpb) · Quesma · 0 upvotes · 0 comments
- [AI models get better at math but still get low marks](https://daily.dev/posts/ai-models-get-better-at-math-but-still-get-low-marks-bugdv2bcl) · The Register · 1 upvotes · 0 comments

---

Tags: [#llm](https://daily.dev/tags/llm), [#computer-vision](https://daily.dev/tags/computer-vision), [#google-gemini](https://daily.dev/tags/google-gemini), [#ai-safety](https://daily.dev/tags/ai-safety)

[View this post on daily.dev](https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Mushroom hunting with LLMs: what can go wrong?","url":"https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x"},"datePublished":"2026-09-02T16:53:08.110Z","dateModified":"2026-09-02T18:02:21.057Z","description":"A benchmark tests several vision-capable LLMs (Gemini 3.6/3.7 Flash, GPT-5.6-Sol, Claude Fable 5.1, GLM-5.3-Flash, DeepSeek v4 Flash Vision) on identifying...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/73f31d244a779960d02034278430322c?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/73f31d244a779960d02034278430322c?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Quesma","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Quesma","logo":"https://media.daily.dev/image/upload/s--I-Be0YJY--/f_auto,q_auto/v1774964372/logos/quesma?_a=BAMAMiWQ0","url":"https://daily.dev/sources/quesma"},"commentCount":3,"discussionUrl":"https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":6},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":3}],"keywords":"llm,computer-vision,google-gemini,ai-safety","timeRequired":"PT19M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Quesma","item":"https://daily.dev/sources/quesma"},{"@type":"ListItem","position":3,"name":"Mushroom hunting with LLMs: what can go wrong?"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x","comment":[{"@type":"Comment","text":"I think dangerous false positives matter more than aggregate accuracy here. I’d report the refusal rate beside poisonous-as-edible errors so a cautious model doesn’t look safe merely because it skips harder images.","datePublished":"2026-09-04T20:49:09.952Z","url":"https://daily.dev/posts/GYoh1Yw8X#c-1Y991fEY4","author":{"@type":"Person","name":"Agustin Barrientos","url":"https://daily.dev/agustinbarrientos","image":"https://media.daily.dev/image/upload/s--5ayxQnqn--/f_auto/v1788281802/avatars/avatar_wQYYVe5Tbj0NJ7C7qPoa8?_a=BAMAMicg0"}},{"@type":"Comment","text":"Well, you as a human have extra context LLM might not have been supplied with, which would have increased odds of success:\n\nYou know the location\nYou know the season (exact date)\nYou know the size relative to your hand (not all photos shows this)\nYou can view it in multiple angles, in 3D. you gave it a single photo\n\nI assume if LLM had at least access to first 2, it might have gotten it better","datePublished":"2026-09-07T16:57:40.511Z","url":"https://daily.dev/posts/GYoh1Yw8X#c-TaTQiL3rm","author":{"@type":"Person","name":"Yair Even Or","url":"https://daily.dev/yaireo","image":"https://avatars.githubusercontent.com/u/845031?v=4"}},{"@type":"Comment","text":"why are LLMs used for this task? Shouldn’t we use proper already-defined deep neural networks?","datePublished":"2026-09-10T13:52:34.519Z","url":"https://daily.dev/posts/GYoh1Yw8X#c-gcui4gNqE","author":{"@type":"Person","name":"FroggoBytes","url":"https://daily.dev/froggobytes","image":"https://lh3.googleusercontent.com/a/ACg8ocIhg6NN-TLfaXu8cQTdOSSfKRyxQz9X2NJZBbdTiYjb8BPjg12j=s96-c"}}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/mushroom-hunting-with-llms-what-can-go-wrong--gyoh1yw8x#faq","mainEntity":[{"@type":"Question","name":"Can I trust an AI chatbot to tell me if a wild mushroom is safe to eat?","acceptedAnswer":{"@type":"Answer","text":"No, current vision-capable LLMs make dangerous misidentification errors on deadly species. In a benchmark across 1,040 photos of 55 species, models like Claude Fable 5.1 and GPT-5.6-Sol sometimes mistook deadly mushrooms such as the death cap, funeral bell, and splendid webcap for edible look-alikes like chanterelles, at rates as high as 15-27%, the same error that kills real foragers. Anyone weighing AI tools against real safety risk can track model evaluation results like these on daily.dev."}},{"@type":"Question","name":"Which vision LLM is best at identifying mushroom species from photos?","acceptedAnswer":{"@type":"Answer","text":"Gemini 3.6 and 3.7 Flash scored highest on first-guess accuracy in a mushroom identification benchmark, outperforming Claude Fable 5.1 and GPT-5.6-Sol despite being cheaper. GLM-5.3-Flash offered the best cost efficiency. The test used 1,040 photos across 55 mushroom species from the FungiTastic dataset, asking each model for its top five species guesses. Developers comparing vision model accuracy and cost can follow benchmarks like this on daily.dev."}}]}
```

