Tenable spent 500+ hours and 40 billion tokens testing Anthropic's Claude Mythos Preview for code security as part of Project Glasswing. Key findings: frontier AI dramatically scales vulnerability discovery but introduces significant noise, non-determinism (up to 30% run-to-run variance), and high costs (estimated ~$500K/FTE/year for continuous use). It works best as a targeted, expert-guided layer alongside deterministic SAST/DAST/SCA tools rather than as a standalone scanning engine. Source code access gives defenders an asymmetric advantage, and human validation remains essential to separate true exposures from raw findings. The model finds the same classes of vulnerabilities as pen testers — just faster and at higher volume — without introducing new vulnerability categories.

9m read timeFrom tenable.com
Post cover image
Table of contents
Key takeawaysHow Tenable is testing Claude Mythos PreviewWhere Mythos and frontier AI fit alongside SAST, DAST, and SCAMatch the tool to the task: determinism for audit, frontier AI for explorationThe real costs of code security with frontier AI go beyond tokensFrontier AI changes the scale of discovery, not the nature of source code flawsSource code access gives defenders a major advantageThe bottom line: frontier AI shifts your edge from finding flaws to proving what matters
1 Impression