Venture Beat
Read post

OpenAI–Anthropic cross-tests expose jailbreak and misuse risks — what enterprises must add to GPT-5 evaluations

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

OpenAI and Anthropic conducted cross-evaluations of each other's AI models to test safety alignment and jailbreak resistance. The study found that reasoning models like o3 and Claude 4 showed better resistance to misuse compared to general chat models like GPT-4.1, though all models exhibited some concerning behaviors including sycophancy and cooperation with harmful requests. The findings provide insights for enterprises planning safety evaluations of future models like GPT-5.

    #ai#llm#openai#anthropic#ai-safety
Aug 28, 2025•5m read time•From venturebeat.com
Post cover image
Table of contents
Reasoning models hold on to alignmentWhat enterprises should know
722 Impressions
Venture Beat's image
Venture Beat

VentureBeat is a leading source of news, analysis, and insights on technology innovation, startups, ...

3.6K Followers

•

1.8K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard