Snyk
Read post

Snyk VulnBench JS 1.0: LLM Bug Repeatability

Snyk ran 300 repeated vulnerability-finding scans across 10 JavaScript fixtures to measure how repeatable LLM-based security reviews are compared to deterministic SAST. Key findings: LLM reference-matched findings were highly stable (85% consistent across all 5 runs), but extra LLM-only reports were highly inconsistent — nearly 50% appeared in only 1 of 5 identical runs. The best LLM configuration (Claude Opus 4.6 Medium) reached 75.4% F1 against Snyk Code's reference set, leaving a 24.6-point gap. More expensive models (Claude Opus 4.7 Max) cost 5.7x more but scored lower. LLMs excelled at high-signal exploit shapes (command injection, SQLi, SSRF) but missed systematic patterns like repeated path traversal sinks and resource-limit findings. The data supports combining LLM review with SAST rather than replacing one with the other.

    #security#llm#appsec
Jun 29•16m read time•From snyk.io
Post cover image
Table of contents
Why LLM security review needs repeatabilityBenchmark design: 300 repeated security runsResult 1: LLM repeatability varied by model configurationResult 2: LLM agents and SAST found different security gapsResult 3: More expensive LLM runs did not mean better coverageAgreement scores against the Snyk Code reference setWhat this benchmark means for LLM security reviewThe AI Security Crisis in Your Python Environment
346 Impressions
Snyk's image
Snyk

Snyk's blog is a source of information and advice for developers looking to ensure the security of t...

98 Followers

•

619 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard