<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu" -->

---
title: Booz Allen&#x27;s AI cyber threat index: Claude Mythos tops...
description: Booz Allen Hamilton&#x27;s first Cyber Weapon Index tested 18 AI models — nine American, nine Chinese — against a live corporate network to gauge autonomous...
canonical: https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Booz Allen&#x27;s AI cyber threat index: Claude Mythos tops the ranking, but the harness matters more than the model | daily.dev
og:description: Booz Allen Hamilton&#x27;s first Cyber Weapon Index tested 18 AI models — nine American, nine Chinese — against a live corporate network to gauge autonomous...
og:url: https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu
og:image: https://api.daily.dev/og/posts/UwzrUTOqU.png
og:image:alt: Booz Allen&#x27;s AI cyber threat index: Claude Mythos tops the ranking, but the harness matters more than the model
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Booz Allen's AI cyber threat index: Claude Mythos tops the ranking, but the harness matters more than the model

**[Collections](https://daily.dev/sources/collections)** · 3 min read · 4 upvotes · 0 comments

## Summary

Booz Allen Hamilton's first Cyber Weapon Index tested 18 AI models — nine American, nine Chinese — against a live corporate network to gauge autonomous cyberattack capability. Anthropic's Claude Mythos topped the ranking with a score of 80, the only model to complete a full cyber kill chain with and without stolen credentials. Against real-world production vulnerabilities, all nine frontier models scored zero, with one partial exception credited to Mythos. The more consequential finding: pairing a low-ranked model (Claude Sonnet 5, ranked 15th) with a purpose-built attack harness closed 67 of the 80-point gap to Mythos, suggesting the harness and surrounding system — not the base model — is the real unit of risk. Guardrails also behaved inconsistently across sibling models on identical tasks. OpenAI's Astra was excluded without explanation despite OpenAI's own claims about crossing a cybersecurity threshold. Booz Allen is calling for regulatory deadlines and a national testing program, while also promoting its own commercial counter-AI product alongside the report.

## Content

Booz Allen Hamilton published its first Cyber Weapon Index, testing 18 AI models — nine American, nine Chinese — against a live corporate network to see how far each could get on its own. The results are worth paying attention to, though the report's most important finding quietly undermines its own headline.

## What they tested

Each model was evaluated on unassisted network intrusion and vulnerability research. The benchmark tracked whether a model could complete a full cyber kill chain — reconnaissance through domain compromise — with and without stolen credentials.

Anthropic's Claude Mythos was the only model to clear the full chain in both scenarios. It achieved domain administrator access with stolen credentials in every attempt, and pulled off full domain compromise even without them, finishing with a score of 80. Grok-4.5, Muse Spark 1.1, and GLM-5.2 reached full domain access but couldn't complete the entire chain unassisted. Alibaba's Qwen3-Coder came in last with a score of 4. Every other model failed the no-credential test entirely.

Against real-world vulnerabilities in production code — the harder, more meaningful test — all nine frontier models scored zero. The report muddies this somewhat by also crediting Mythos with understanding and exploiting one such flaw, which is a tension the authors don't fully resolve.

OpenAI's Astra model wasn't tested at all, despite OpenAI having stated it crossed a critical cybersecurity threshold. No explanation was given for the exclusion.

## The finding that undercuts the ranking

Here's where it gets interesting. Booz Allen paired Claude Sonnet 5 — ranked 15th with a score of 13 — with an attack harness, a purpose-built scaffolding system that structures how the model approaches offensive tasks. That combination closed 67 of the 80-point gap between Sonnet and Mythos.

The implication is uncomfortable: the raw model ranking may not tell you much about actual threat level. A cheap, widely available model wrapped in the right harness can approach the capability of the most dangerous model tested. Booz Allen's own conclusion is that the harness and surrounding system, not the underlying model, is the real unit of risk.

The report also found that guardrails behaved inconsistently across sibling models given identical tasks — meaning safety controls aren't reliable even within the same model family.

## What Booz Allen is calling for

The report warns that most tested models could reach Mythos' weaponization level within six months, and describes mainstream AI-driven cyberattacks as "imminent." It calls for regulatory deadlines and a national testing program covering foreign and open-weight models.

One thing worth noting: Booz Allen published this index alongside its own commercial counter-AI product, Vellox Labs Guile. That doesn't make the findings wrong, but it's context readers should have.

## Questions this post answers

### which AI model scored highest on Booz Allen's Cyber Weapon Index for autonomous cyberattacks

Anthropic's Claude Mythos topped Booz Allen Hamilton's first Cyber Weapon Index with a score of 80, the only model among 18 tested to complete a full cyber kill chain from reconnaissance to domain compromise both with and without stolen credentials. Grok-4.5, Muse Spark 1.1, and GLM-5.2 reached full domain access but failed to complete the chain unassisted; Alibaba's Qwen3-Coder scored lowest at 4.

_Anyone tracking AI cybersecurity risk can follow benchmark results like this on daily.dev._

### does the harness or the underlying AI model matter more for cyberattack capability

The attack harness matters more than the base model, according to Booz Allen Hamilton's testing. Pairing Claude Sonnet 5 — ranked 15th with a score of 13 — with a purpose-built attack harness closed 67 of the 80-point gap to top-ranked Claude Mythos, showing that a cheap, widely available model wrapped in the right scaffolding can approach top-tier attack capability.

_Security teams weighing AI threat models can track findings like this on daily.dev._

## Similar posts on daily.dev

- [Claude Mythos Preview completes full cyberattack simulation for the first time](https://daily.dev/posts/claude-mythos-preview-completes-full-cyberattack-simulation-for-the-first-time-dbsyc6gmb) · The New Stack · 0 upvotes · 0 comments

---

Tags: [#ai](https://daily.dev/tags/ai), [#cyber](https://daily.dev/tags/cyber), [#claude](https://daily.dev/tags/claude), [#anthropic](https://daily.dev/tags/anthropic), [#ai-security](https://daily.dev/tags/ai-security)

[View this post on daily.dev](https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Booz Allen's AI cyber threat index: Claude Mythos tops the ranking, but the harness matters more than the model","url":"https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu"},"datePublished":"2026-09-03T09:15:48.529Z","dateModified":"2026-09-03T09:16:37.769Z","description":"Booz Allen Hamilton's first Cyber Weapon Index tested 18 AI models — nine American, nine Chinese — against a live corporate network to gauge autonomous...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/90ae77398aa9763754fde92bfa1eab29?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/90ae77398aa9763754fde92bfa1eab29?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":4},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"ai,cyber,claude,anthropic,ai-security","timeRequired":"PT3M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Booz Allen's AI cyber threat index: Claude Mythos tops the ranking, but the harness matters more than the model"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/booz-allen-s-ai-cyber-threat-index-claude-mythos-tops-the-ranking-but-the-harness-matters-more-tha-uwzrutoqu#faq","mainEntity":[{"@type":"Question","name":"which AI model scored highest on Booz Allen's Cyber Weapon Index for autonomous cyberattacks","acceptedAnswer":{"@type":"Answer","text":"Anthropic's Claude Mythos topped Booz Allen Hamilton's first Cyber Weapon Index with a score of 80, the only model among 18 tested to complete a full cyber kill chain from reconnaissance to domain compromise both with and without stolen credentials. Grok-4.5, Muse Spark 1.1, and GLM-5.2 reached full domain access but failed to complete the chain unassisted; Alibaba's Qwen3-Coder scored lowest at 4. Anyone tracking AI cybersecurity risk can follow benchmark results like this on daily.dev."}},{"@type":"Question","name":"does the harness or the underlying AI model matter more for cyberattack capability","acceptedAnswer":{"@type":"Answer","text":"The attack harness matters more than the base model, according to Booz Allen Hamilton's testing. Pairing Claude Sonnet 5 — ranked 15th with a score of 13 — with a purpose-built attack harness closed 67 of the 80-point gap to top-ranked Claude Mythos, showing that a cheap, widely available model wrapped in the right scaffolding can approach top-tier attack capability. Security teams weighing AI threat models can track findings like this on daily.dev."}}]}
```

