<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy" -->

---
title: DeepSeek&#x27;s New Vision Model Beats Anthropic on 3...
description: DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built on V4-Flash that adds image and screenshot understanding, alongside...
canonical: https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: DeepSeek&#x27;s New Vision Model Beats Anthropic on 3 Benchmarks, Loses on 8 | daily.dev
og:description: DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built on V4-Flash that adds image and screenshot understanding, alongside...
og:url: https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy
og:image: https://api.daily.dev/og/posts/CBoNzBxZy.png
og:image:alt: DeepSeek&#x27;s New Vision Model Beats Anthropic on 3 Benchmarks, Loses on 8
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# DeepSeek's New Vision Model Beats Anthropic on 3 Benchmarks, Loses on 8

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 1 upvotes · 0 comments

## Summary

DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built on V4-Flash that adds image and screenshot understanding, alongside agent harness v0.1.1. On DeepSeek's own eleven-benchmark table it beats Anthropic's Opus-4.8 on three benchmarks (DeepSWE, Agents' Last Exam, ZeroBench) but trails on eight, including a 12-point gap on NL2Repo. Notably, Anthropic's Opus 5 and DeepSeek's own newly shipped V4-Pro are both absent from the comparison. Part of the reported multimodal gain comes from the old V4-Flash being scored on image tests it couldn't see, a caveat DeepSeek disclosed itself; even excluding that, V4-Flash-Vision-Exp improves on six of seven text-only benchmarks. The clearest differentiator isn't the benchmarks but price: V4-Flash costs roughly 87 cents per million words versus roughly $50 from Anthropic.

## Content

DeepSeek just put out DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built on top of its text-only V4-Flash, now able to read images and screenshots. It ships alongside agent harness v0.1.1.

The announcement leans hard on beating Anthropic, and the benchmark table DeepSeek published backs that up... selectively. Out of eleven benchmarks, the new model beats Opus 4.8 on three (DeepSWE, Agents' Last Exam, ZeroBench) and loses on the other eight, including a 12-point gap on NL2Repo. That's not exactly a rout.

And here's the part that bugs me: Anthropic's Opus 5, which came out back in July, doesn't appear in the comparison at all. Comparing your brand-new model to the previous generation's flagship isn't the same as comparing it to what's actually current. That's a meaningful omission for a table meant to show competitive standing.

There's also a methodology wrinkle DeepSeek itself flagged: part of the apparent jump over the old V4-Flash comes from the fact that V4-Flash was scored on benchmarks containing images it literally cannot process, since it's text-only. So some of that gain is an artifact of comparing a blind model to a sighted one, not a fair multimodal-to-multimodal test. Credit to DeepSeek for disclosing this rather than burying it, but it does mean the headline numbers need a second look.

What's genuinely solid: the vision model improves on six of seven text benchmarks compared to V4-Flash, meaning this isn't just bolted-on image support at the cost of everything else.

Separately, DeepSeek confirmed V4-Pro has officially shipped, with better agent capabilities, Responses API support, and Codex integration. Oddly, V4-Pro is also missing from the comparison table, which is strange for what should be their strongest model.

Where DeepSeek doesn't need to spin anything is price. V4-Flash runs around 87 cents per million words, versus roughly $50 from Anthropic for comparable usage. That's not a marginal difference, that's the kind of gap that changes what's actually feasible to build. If you're running high-volume agentic workflows, this pricing alone might matter more than a few points on ZeroBench.

A few people online (understandably) got excited about this being a turning point for multimodal agents. I'd hold off on that framing until we see how it performs against Opus 5 directly, and until the image-benchmark comparison gets redone on fairer footing. The pricing story stands on its own, though — no caveats needed there.

## Questions this post answers

### What is DeepSeek-V4-Flash-Vision-Exp and what does it add over V4-Flash?

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal model built on top of the text-only DeepSeek V4-Flash, adding image and screenshot understanding. It ships alongside agent harness v0.1.1 and, excluding a scoring artifact from the old model being tested on images it couldn't see, still improves on six of seven text-only benchmarks compared to V4-Flash.

_Track new multimodal model releases like this one as they ship on daily.dev._

### How does DeepSeek V4-Flash pricing compare to Anthropic Opus for large-scale usage?

DeepSeek V4-Flash costs roughly 87 cents per million words, compared to roughly $50 from Anthropic for the same volume. That price gap is described as mattering more for real-world decisions than the handful of benchmark points separating the two models on any given test.

_Compare model pricing trade-offs like this before picking an LLM provider on daily.dev._

### What new capabilities does DeepSeek V4-Pro add?

DeepSeek V4-Pro has officially shipped with enhanced agent capabilities, Responses API support, and Codex integration. It notably does not appear in DeepSeek's own eleven-benchmark comparison table against Anthropic's Opus-4.8, alongside the also-absent Opus 5.

_Follow model rollouts like V4-Pro to keep coding agent integrations current on daily.dev._

---

Tags: [#ai](https://daily.dev/tags/ai), [#deep-learning](https://daily.dev/tags/deep-learning), [#anthropic](https://daily.dev/tags/anthropic), [#vibe-coding](https://daily.dev/tags/vibe-coding), [#deepseek](https://daily.dev/tags/deepseek)

[View this post on daily.dev](https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"DeepSeek's New Vision Model Beats Anthropic on 3 Benchmarks, Loses on 8","url":"https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy"},"datePublished":"2026-08-21T11:55:49.535Z","dateModified":"2026-08-21T13:53:44.892Z","description":"DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model built on V4-Flash that adds image and screenshot understanding, alongside...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/5f3877acfcc83470b0635c9a93ac9978?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/5f3877acfcc83470b0635c9a93ac9978?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,deep-learning,anthropic,vibe-coding,deepseek","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"DeepSeek's New Vision Model Beats Anthropic on 3 Benchmarks, Loses on 8"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/deepseek-s-new-vision-model-beats-anthropic-on-3-benchmarks-loses-on-8-cbonzbxzy#faq","mainEntity":[{"@type":"Question","name":"What is DeepSeek-V4-Flash-Vision-Exp and what does it add over V4-Flash?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal model built on top of the text-only DeepSeek V4-Flash, adding image and screenshot understanding. It ships alongside agent harness v0.1.1 and, excluding a scoring artifact from the old model being tested on images it couldn't see, still improves on six of seven text-only benchmarks compared to V4-Flash. Track new multimodal model releases like this one as they ship on daily.dev."}},{"@type":"Question","name":"How does DeepSeek V4-Flash pricing compare to Anthropic Opus for large-scale usage?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek V4-Flash costs roughly 87 cents per million words, compared to roughly $50 from Anthropic for the same volume. That price gap is described as mattering more for real-world decisions than the handful of benchmark points separating the two models on any given test. Compare model pricing trade-offs like this before picking an LLM provider on daily.dev."}},{"@type":"Question","name":"What new capabilities does DeepSeek V4-Pro add?","acceptedAnswer":{"@type":"Answer","text":"DeepSeek V4-Pro has officially shipped with enhanced agent capabilities, Responses API support, and Codex integration. It notably does not appear in DeepSeek's own eleven-benchmark comparison table against Anthropic's Opus-4.8, alongside the also-absent Opus 5. Follow model rollouts like V4-Pro to keep coding agent integrations current on daily.dev."}}]}
```

