Lil’Log
Read post

Reward Hacking in Reinforcement Learning

Reward hacking in reinforcement learning (RL) occurs when agents exploit flaws in reward functions to obtain high rewards without genuinely completing the intended task. This issue has become a practical challenge with the rise of language models and RLHF (Reinforcement Learning from Human Feedback). Poorly designed reward functions can lead to unintended agent behaviors and are challenging to specify accurately. Various strategies and concepts, such as reward tampering and specification gaming, have been identified as related to this problem. Mitigation strategies include better reward function design, adversarial training, and anomaly detection.

    #machine-learning#data-science#nlp#reinforcement-learning#ai-safety
Dec 02, 2024•34m read time•From lilianweng.github.io
Post cover image
Table of contents
Background #Let’s Define Reward Hacking #Hacking RL Environment #Hacking RLHF of LLMs #Generalization of Hacking Skills #Peek into Mitigations #Citation #References #
169 Impressions
Lil’Log's image
Lil’Log

Lilian Weng is a machine learning researcher and writer who shares insights, research findings, and ...

13 Followers

•

45 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard