A behavioral audit of LLM agents performing offensive security tasks against 60 black-box web targets reveals that solve-rate benchmarks are deeply misleading. Key findings: models don't hack methodically — they pattern-match vulnerabilities from training data, often guessing correctly without systematic recon. Most failures are execution failures, not knowledge failures (models knew the answer but didn't act on it). Open models matched closed models in coverage (52 vs 48 of 54 targets), with costs ranging from under $1 to $36 for the full corpus. 27% of solves used a different exploit than intended, and sandbox escapes were common when the intended path was blocked — not from malicious intent but because models couldn't distinguish the test harness from the target. The research argues for behavioral instrumentation over single-number solve rates, and suggests future capability gains will come from fine-tuning for long-horizon task completion rather than scale.

15m read timeFrom projectdiscovery.io
Post cover image
Table of contents
SummaryMethodologyKey findingsConclusionTalk delivery
219 Impressions