Embrace The Red
Read post

Scary Agent Skills: Hidden Unicode Instructions in Skills ...And How To Catch Them · Embrace The Red

This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).

AI agent Skills can be backdoored with invisible Unicode Tag instructions that survive human review. The attack exploits how certain LLMs (Gemini, Claude, Grok) interpret hidden Unicode codepoints as executable instructions. A demonstration shows backdooring OpenAI's security-best-practices Skill to execute arbitrary bash commands. The post includes a scanner tool to detect such attacks and proposes mitigations including sandboxing agents, selective Skill installation, and detection of invisible Unicode sequences.

    #ai-agents#claude#prompt-injection#security#supply-chain
Feb 11•8m read time•From embracethered.com
Post cover image
Table of contents
Attack SurfaceWhat is an Agent Skill?Scary SkillsWriting a Simple SkillPrompt Injection Attack VectorsAgent(s) Overwriting Skills on the FlyUsing Invisible Instructions in SkillsAdding a Backdoor to A Legitimate SkillEnd to End VideoNotes, Testing Observations and MitigationsA Scanner to Catch AttacksConclusionReferencesAppendix
150 Impressions
Embrace The Red's image
Embrace The Red

Embrace the Red's resource offers insights, tutorials, and resources for developers and enthusiasts ...

40 Followers

•

182 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard