SentinelLABS built a real-world, eight-stage reverse-engineering benchmark based on their investigation of fast16, a 2005 Windows sabotage implant targeting nuclear-weapons modeling software. The benchmark tests whether frontier AI models can sustain a trustworthy malware investigation as new evidence repeatedly invalidates earlier conclusions — a capability they call 'project-scale recovery.' Most tested models (GPT-5.5, GLM-5.2, Opus 4.x) showed strong local analytical ability but failed to carry that quality through the full investigation. GPT-5.6 Sol was the only publicly available model to complete all eight stages, distinguishing itself by withdrawing contradicted conclusions, mapping downstream impact, repairing root causes, and propagating corrections throughout the entire project. Despite this milestone, senior reverse engineers remain essential: Sol still made semantic errors, accepted weak quality controls, and required human oversight at critical junctures. The practical takeaway is 'supervised investigative agency' — one expert overseeing these systems can now dramatically multiply their output without replacing human judgment.