SentinelLABS built a real-world, eight-stage reverse-engineering benchmark based on their investigation of fast16, a 2005 Windows sabotage implant targeting nuclear-weapons modeling software. The benchmark tests whether frontier AI models can sustain a trustworthy malware investigation as new evidence repeatedly invalidates earlier conclusions — a capability they call 'project-scale recovery.' Most tested models (GPT-5.5, GLM-5.2, Opus 4.x) showed strong local analytical ability but failed to carry that quality through the full investigation. GPT-5.6 Sol was the only publicly available model to complete all eight stages, distinguishing itself by withdrawing contradicted conclusions, mapping downstream impact, repairing root causes, and propagating corrections throughout the entire project. Despite this milestone, senior reverse engineers remain essential: Sol still made semantic errors, accepted weak quality controls, and required human oversight at critical junctures. The practical takeaway is 'supervised investigative agency' — one expert overseeing these systems can now dramatically multiply their output without replacing human judgment.

15m read timeFrom sentinelone.com
Post cover image
Table of contents
Executive SummaryBeyond Vulnerability DiscoveryA Benchmark Built From a Real InvestigationStandardizing the BenchmarkMeasuring ‘Intelligence’What the Runs ShowedGPT-5.6 Sol | A New Class of ContenderProject-Scale RecoveryNotes on CostWhat These Results Mean for AnalystsConclusion
189 Impressions