<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g" -->

---
title: You picked Claude Sonnet 5.5 — but Anthropic may send...
description: Claude Sonnet 5.5, released with the same three-stage cyber classifier and model-fallback system previously reserved for Anthropic&#x27;s Opus tier, brings...
canonical: https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5 | daily.dev
og:description: Claude Sonnet 5.5, released with the same three-stage cyber classifier and model-fallback system previously reserved for Anthropic&#x27;s Opus tier, brings...
og:url: https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g
og:image: https://api.daily.dev/og/posts/9AYi4nC5g.png
og:image:alt: You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5

**[The New Stack](https://daily.dev/sources/newstack)** · 11 min read · 0 upvotes · 0 comments

## Summary

Claude Sonnet 5.5, released with the same three-stage cyber classifier and model-fallback system previously reserved for Anthropic's Opus tier, brings routing-based safety enforcement to its cheaper, production-tier model. When cyber-related requests trigger a block, Anthropic's own apps automatically fall back to the older Sonnet 5, but API developers must opt in, meaning a straight swap from Sonnet 5 to 5.5 isn't guaranteed. Benchmarks show Sonnet 5.5 has dramatically improved offensive-security capability (46.1% on CyScenarioBench versus 0.7% for Sonnet 5), prompting stricter jailbreak checks and more refusals, even on legitimate cybersecurity work. Anthropic's own testing found requests rerouted to Sonnet 5 after a cyber block were compromised by prompt injection 12.01% of the time, versus near-zero for Sonnet 5.5 itself, showing fallback introduces a real security trade-off for agentic coding workflows.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://thenewstack.io/claude-sonnet-cyber-safeguards>

## Questions this post answers

### Does Claude Sonnet 5.5 automatically fall back to Sonnet 5 when a cyber-related request is blocked?

It depends on the platform. Anthropic's own apps automatically route blocked cyber requests to Sonnet 5, but developers using the API must explicitly enable that fallback themselves, and other platforms may handle blocked requests differently, so a blocked request without fallback enabled simply stops instead of continuing.

_Developers wiring up Claude API fallback behavior can track these safety-routing changes on daily.dev._

### Is it safe to rely on Sonnet 5 fallback for coding agents when Sonnet 5.5 hits a cyber block?

Not entirely risk-free: in Anthropic's own testing of coding environments, 25% of requests to Sonnet 5.5 got rerouted to Sonnet 5 after triggering a cyber block, often due to injected instructions like disk-wipe commands, and 12.01% of those rerouted requests were successfully compromised, compared to just four compromises out of 5,901 requests Sonnet 5.5 handled itself.

_Teams building agentic coding workflows can weigh model-fallback security trade-offs like this via daily.dev._

### How much better is Claude Sonnet 5.5 at offensive cybersecurity tasks compared to Sonnet 5?

Sonnet 5.5 shows a large jump in offensive security capability: it completed 46.1% of challenges on Irregular's CyScenarioBench versus just 0.7% for Sonnet 5, achieved full arbitrary code execution in 178 of 410 ExploitBench runs with safeguards off, and managed 50 control-flow hijacks on an OSS-Fuzz-based binary exploitation benchmark versus three for Sonnet 5.

_Security researchers evaluating model capability jumps like this can follow model safety benchmarks on daily.dev._

## Similar posts on daily.dev

- [Anthropic Sonnet 5: It closes the gap with Opus 4.8, and is cheap until August](https://daily.dev/posts/anthropic-sonnet-5-it-closes-the-gap-with-opus-4-8-and-is-cheap-until-august-o0toggvty) · The New Stack · 0 upvotes · 0 comments
- [Claude Sonnet 5: A Security Deep Dive for AI Agent Deployments](https://daily.dev/posts/claude-sonnet-5-a-security-deep-dive-for-ai-agent-deployments-niacrikpi) · Medium · 1 upvotes · 0 comments
- [Anthropic releases Sonnet 4.6](https://daily.dev/posts/anthropic-releases-sonnet-4-6-aufw9ozvi) · TechCrunch · 1 upvotes · 0 comments

---

Tags: [#claude](https://daily.dev/tags/claude), [#anthropic](https://daily.dev/tags/anthropic), [#ai-safety](https://daily.dev/tags/ai-safety), [#prompt-injection](https://daily.dev/tags/prompt-injection), [#ai-gateway](https://daily.dev/tags/ai-gateway)

[View this post on daily.dev](https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5","url":"https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g"},"datePublished":"2026-09-28T21:36:18.110Z","dateModified":"2026-09-28T23:06:23.052Z","description":"Claude Sonnet 5.5, released with the same three-stage cyber classifier and model-fallback system previously reserved for Anthropic's Opus tier, brings...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/de430354a1badf2e103ab9e9afef56ad?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/de430354a1badf2e103ab9e9afef56ad?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"The New Stack","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"The New Stack","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/newstack","url":"https://daily.dev/sources/newstack"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"claude,anthropic,ai-safety,prompt-injection,ai-gateway","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"The New Stack","item":"https://daily.dev/sources/newstack"},{"@type":"ListItem","position":3,"name":"You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-9ayi4nc5g#faq","mainEntity":[{"@type":"Question","name":"Does Claude Sonnet 5.5 automatically fall back to Sonnet 5 when a cyber-related request is blocked?","acceptedAnswer":{"@type":"Answer","text":"It depends on the platform. Anthropic's own apps automatically route blocked cyber requests to Sonnet 5, but developers using the API must explicitly enable that fallback themselves, and other platforms may handle blocked requests differently, so a blocked request without fallback enabled simply stops instead of continuing. Developers wiring up Claude API fallback behavior can track these safety-routing changes on daily.dev."}},{"@type":"Question","name":"Is it safe to rely on Sonnet 5 fallback for coding agents when Sonnet 5.5 hits a cyber block?","acceptedAnswer":{"@type":"Answer","text":"Not entirely risk-free: in Anthropic's own testing of coding environments, 25% of requests to Sonnet 5.5 got rerouted to Sonnet 5 after triggering a cyber block, often due to injected instructions like disk-wipe commands, and 12.01% of those rerouted requests were successfully compromised, compared to just four compromises out of 5,901 requests Sonnet 5.5 handled itself. Teams building agentic coding workflows can weigh model-fallback security trade-offs like this via daily.dev."}},{"@type":"Question","name":"How much better is Claude Sonnet 5.5 at offensive cybersecurity tasks compared to Sonnet 5?","acceptedAnswer":{"@type":"Answer","text":"Sonnet 5.5 shows a large jump in offensive security capability: it completed 46.1% of challenges on Irregular's CyScenarioBench versus just 0.7% for Sonnet 5, achieved full arbitrary code execution in 178 of 410 ExploitBench runs with safeguards off, and managed 50 control-flow hijacks on an OSS-Fuzz-based binary exploitation benchmark versus three for Sonnet 5. Security researchers evaluating model capability jumps like this can follow model safety benchmarks on daily.dev."}}]}
```

