Tenable spent 500+ hours and 40 billion tokens testing Anthropic's Claude Mythos Preview for code security as part of Project Glasswing. Key findings: frontier AI dramatically scales vulnerability discovery but introduces significant noise, non-determinism (up to 30% run-to-run variance), and high costs (estimated ~$500K/FTE/year for continuous use). It works best as a targeted, expert-guided layer alongside deterministic SAST/DAST/SCA tools rather than as a standalone scanning engine. Source code access gives defenders an asymmetric advantage, and human validation remains essential to separate true exposures from raw findings. The model finds the same classes of vulnerabilities as pen testers — just faster and at higher volume — without introducing new vulnerability categories.