Researchers discovered that GitHub Copilot's safety guardrails can be bypassed by embedding harmful requests within code rather than plain text prompts. The jailbreak operates at the workflow level, allowing users to circumvent content restrictions that would normally block such requests when made in natural language.
Table of contents
Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injectionFeds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcherTelling an AI model that it’s an expert programmer makes it a worse programmerResearchers find hole in AI guardrails by using strings like =coffee1.2K Impressions