Researchers from the New England RLHF Hackers (NERH) group held a hackathon at Brown University in September 2023, producing three research directions: (1) a heuristic framework for directly evaluating reward models using an ensemble of classifiers for undesirable attributes like harmfulness and unhelpfulness; (2) a fact-based reward model that uses document retrieval from sources like Wikipedia and Stack Exchange combined with a logician LLM to score factual accuracy; and (3) an RLAIF approach to the 'pink elephant problem,' using ILQL to train language models to avoid generating responses about specific topics (demonstrated with Ubuntu as the restricted topic).
Table of contents
Introduction #On the Evaluation of Reward Models #A Fact-Based Reward Model for Language Models #The Pink Elephant Problem in Language Models #How to cite #3 Impressions