Simon Willison
Read post

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

The paper discusses training Language Model Manners (LLMs) to prioritize privileged instructions and proposes a hierarchy for models to consider conflict or alignment with higher-level instructions. The authors claim improved performance against prompt injection benchmarks but acknowledge vulnerability to powerful adversarial attacks.

    #llm#openai#prompt-injection
Apr 23, 2024•1m read time•From simonwillison.net
Post cover image
26 Impressions
Simon Willison's image
Simon Willison

Simon Willison's blog offers a mix of technical tutorials, data analysis projects, and reflections o...

215 Followers

•

1.4K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard