Anthropic's second Risk Report reveals that bioweapon-blocking classifiers were disabled on its human feedback platforms for eleven months (May 2025 to April 2026), covering roughly 133 million contractor exchanges from about 50,000 people whose vetting Anthropic outsourced to vendors, some of which lacked screening capable of stopping serious threat actors. An internal-only flag silently disabled both the blocking and the logging, so incidents could not be reviewed after the fact. A review using Claude Sonnet 5 flagged 1,197 high-risk transcripts; manual review found no clear misuse but flagged some dual-use conversations, and Anthropic now says it believes similar undiscovered issues are more likely. The company also retroactively downgraded its February safety verdict from 'very low' to 'low' risk, disclosed a second incident where contractors misused a leaked API key to access an unreleased model called Mythos Preview for about two weeks, and included a self-critical review of the report written by a Claude instance, which flagged sections as overly reassuring and criticized a fully redacted incident. Anthropic also disclosed an unreleased internal model, Model 2, and rewrote its novel-weapons risk threshold.
Table of contents
Eleven months with the filters offWhat the review foundFebruary’s report has been correctedA second incident, in the same placeWhy the misalignment number actually movedAnthropic asked Claude to mark the homeworkModel 2, and the thresholds that movedWhat would settle itQuestions this post answers
Why did Anthropic disable its bioweapon safety classifiers for contractor chats?
An internal-use-only flag was mistakenly left in a state that disabled both the blocking classifiers and the logging on Anthropic's human feedback platforms from May 2025 until April 2026. This affected roughly 133 million exchanges from about 50,000 contractors vetted by outside vendors, some of which lacked screening capable of stopping serious biological threat actors, and flagged traffic was not recorded for later review. Track how AI labs disclose and patch safety gaps like this one on daily.dev.
What did Anthropic's review find after the bioweapon classifier gap was discovered?
Running Claude Sonnet 5 over every human turn sent during the affected period flagged 1,197 transcripts as high risk, of which 757 came from Anthropic's own teams and all but 62 of the remainder came from deliberate red-teaming. Manual review of those 62 plus 30 random red-teaming transcripts found no clearly concerning misuse, though some potentially dual-use conversations were identified, and Anthropic now believes similar undiscovered issues are more likely. Follow how AI safety incident reviews get conducted and disclosed on daily.dev.
What happened when contractors got unauthorized access to Anthropic's Mythos Preview model?
A few contractors at data-labelling vendors exploited a flaw in April 2026 to obtain an API key and used models outside their assigned work, including Mythos Preview, one of Anthropic's most capable models at the time. The access path stayed open for several weeks, with Mythos Preview exposed for roughly two weeks without blocking biological classifiers, before Anthropic contained it within 90 minutes of learning about it. Keep up with AI vendor security incidents like this one on daily.dev.