OpenAI's path to ChatGPT began not with language models but with reinforcement learning. Ilya Sutskever, Alec Radford, and colleagues pivoted to Google's transformer architecture and devised a two-stage training approach: self-supervised pre-training on a 7,000-book dataset followed by supervised fine-tuning on specific language tasks. This produced GPT-1, which set state-of-the-art results on nine language benchmarks. Though largely unnoticed publicly, GPT-1 broke the dependency on human-labeled data and enabled massive scaling — leading to GPT-2 (2019), GPT-3 (2020), and ultimately ChatGPT (2022).

2m watch time
28 Impressions