<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb" -->

---
title: A cheap piece of software erased Booz Allen’s own AI...
description: Booz Allen tested 18 leading AI models, nine American and nine Chinese, on unassisted network intrusion and vulnerability research against a live corporate...
canonical: https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: A cheap piece of software erased Booz Allen’s own AI threat ranking | daily.dev
og:description: Booz Allen tested 18 leading AI models, nine American and nine Chinese, on unassisted network intrusion and vulnerability research against a live corporate...
og:url: https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb
og:image: https://api.daily.dev/og/posts/VIU9tmBQb.png
og:image:alt: A cheap piece of software erased Booz Allen’s own AI threat ranking
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# A cheap piece of software erased Booz Allen’s own AI threat ranking

**[The Next Web](https://daily.dev/sources/tnw)** · 5 min read · 0 upvotes · 0 comments

## Summary

Booz Allen tested 18 leading AI models, nine American and nine Chinese, on unassisted network intrusion and vulnerability research against a live corporate network. Only Anthropic's Claude Mythos completed the full kill chain, scoring 80 and achieving domain admin with and without stolen credentials. Every other model failed the no-credential test, and Alibaba's Qwen3-Coder finished last with a score of 4. The report's headline finding undercuts its own ranking: pairing weaker model Claude Sonnet 5 (ranked 15th, score 13) with an attack harness closed 67 points of the gap to Mythos, leading Booz Allen to conclude that the harness and surrounding system, not the raw model, is the real unit of risk. Guardrails also proved inconsistent across sibling models given identical tasks. Against unseen real-world vulnerabilities in production code, all nine frontier models tested scored zero, though the report muddies this by also crediting Mythos with understanding and exploiting the flaw. Booz Allen published the index alongside its own commercial counter-AI product, Vellox Labs Guile, and calls for regulatory deadlines and a national testing program for foreign and open-weight models. OpenAI's Astra model was excluded from testing despite OpenAI stating it had crossed a critical cybersecurity threshold.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://thenextweb.com/news/booz-allen-cyber-weapon-index-attack-harness>

## Questions this post answers

### Which AI model completed a full autonomous cyberattack from start to finish in Booz Allen's Cyber Weapon Index test?

Claude Mythos, from Anthropic, was the only model among 18 tested that completed the full cyber kill chain, scoring 80 out of the models evaluated. Given a stolen credential it achieved administrator control on every attempt, and on the harder test with no credentials at all it broke in from outside and still took full domain control, something no other model managed.

_Track how frontier models compare on real offensive capability before betting your defenses on daily.dev._

### Does pairing a weaker AI model with an attack harness make it as dangerous as a top-ranked model?

Yes, according to Booz Allen's testing: Claude Sonnet 5 ranked 15th of 18 models with a raw score of 13, but once paired with an attack harness that keeps it on task and chains actions together, its performance rivaled Claude Mythos, closing a 67-point gap. Booz Allen concludes the surrounding system, not the base model, is now the real unit of security risk.

_Developers assessing AI security risk can follow how tooling changes model capability on daily.dev._

### How many frontier AI models found a genuine unseen vulnerability in production code during Booz Allen's cybersecurity benchmark?

None scored on discovering a real, previously unknown vulnerability buried in a large production library; all nine frontier models tested scored zero on that specific task, even though every model performed near the ceiling against a deliberately planted, known flaw. One model even analyzed the vulnerable component correctly but concluded it was safe.

_Anyone weighing AI-assisted vulnerability research against real limitations can follow the results on daily.dev._

---

Tags: [#security](https://daily.dev/tags/security), [#llm](https://daily.dev/tags/llm), [#claude](https://daily.dev/tags/claude), [#anthropic](https://daily.dev/tags/anthropic), [#ai-security](https://daily.dev/tags/ai-security)

[View this post on daily.dev](https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"A cheap piece of software erased Booz Allen’s own AI threat ranking","url":"https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb"},"datePublished":"2026-09-03T08:07:22.182Z","dateModified":"2026-09-09T13:54:10.174Z","description":"Booz Allen tested 18 leading AI models, nine American and nine Chinese, on unassisted network intrusion and vulnerability research against a live corporate...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/90ae77398aa9763754fde92bfa1eab29?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/90ae77398aa9763754fde92bfa1eab29?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"The Next Web","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"The Next Web","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/tnw","url":"https://daily.dev/sources/tnw"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"security,llm,claude,anthropic,ai-security","timeRequired":"PT5M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"The Next Web","item":"https://daily.dev/sources/tnw"},{"@type":"ListItem","position":3,"name":"A cheap piece of software erased Booz Allen’s own AI threat ranking"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/a-cheap-piece-of-software-erased-booz-allen-s-own-ai-threat-ranking-viu9tmbqb#faq","mainEntity":[{"@type":"Question","name":"Which AI model completed a full autonomous cyberattack from start to finish in Booz Allen's Cyber Weapon Index test?","acceptedAnswer":{"@type":"Answer","text":"Claude Mythos, from Anthropic, was the only model among 18 tested that completed the full cyber kill chain, scoring 80 out of the models evaluated. Given a stolen credential it achieved administrator control on every attempt, and on the harder test with no credentials at all it broke in from outside and still took full domain control, something no other model managed. Track how frontier models compare on real offensive capability before betting your defenses on daily.dev."}},{"@type":"Question","name":"Does pairing a weaker AI model with an attack harness make it as dangerous as a top-ranked model?","acceptedAnswer":{"@type":"Answer","text":"Yes, according to Booz Allen's testing: Claude Sonnet 5 ranked 15th of 18 models with a raw score of 13, but once paired with an attack harness that keeps it on task and chains actions together, its performance rivaled Claude Mythos, closing a 67-point gap. Booz Allen concludes the surrounding system, not the base model, is now the real unit of security risk. Developers assessing AI security risk can follow how tooling changes model capability on daily.dev."}},{"@type":"Question","name":"How many frontier AI models found a genuine unseen vulnerability in production code during Booz Allen's cybersecurity benchmark?","acceptedAnswer":{"@type":"Answer","text":"None scored on discovering a real, previously unknown vulnerability buried in a large production library; all nine frontier models tested scored zero on that specific task, even though every model performed near the ceiling against a deliberately planted, known flaw. One model even analyzed the vulnerable component correctly but concluded it was safe. Anyone weighing AI-assisted vulnerability research against real limitations can follow the results on daily.dev."}}]}
```

