Researchers train Opus-sized model that generalizes reward hacking into cyberattacks and safety evasion30h ago