Skip to main content

Key Papers & Reading List

Link copied!

A chronological spine of the field. The ones tagged Start here or Landmark are the priority reads; the rest fill in context. Links go to arXiv (or the canonical source). "Why it matters" in one line each.

Foundations (2013–2020)

  • Word2Vec / "Efficient Estimation of Word Representations in Vector Space" (2013): arxiv.org/abs/1301.3781. Learned word embeddings; the vector-space intuition everything downstream inherits.
  • Attention Is All You Need (2017): arxiv.org/abs/1706.03762. Introduced the Transformer. Every modern LLM descends from it. Start here.
  • Deep RL from Human Preferences (Christiano et al., 2017): arxiv.org/abs/1706.03741. The origin of RLHF: learning a reward from human comparisons. The recipe InstructGPT later scaled to language. Landmark.
  • BERT (2018): arxiv.org/abs/1810.04805. Bidirectional pretraining that made "fine-tune a big pretrained model" the default paradigm.
  • GPT-2 / "Language Models are Unsupervised Multitask Learners" (2019): cdn.openai.com/better-language-models. Scaling generative pretraining. The "too dangerous to release" moment.
  • Scaling Laws for Neural Language Models (2020): arxiv.org/abs/2001.08361. Showed loss falls predictably with compute, data, and parameters; the economic logic of the whole race.
  • GPT-3 / "Language Models are Few-Shot Learners" (2020): arxiv.org/abs/2005.14165. In-context learning at scale. The start of prompting as a discipline. Landmark.

The alignment & reasoning era (2021–2023)

  • LoRA (2021): arxiv.org/abs/2106.09685. Parameter-efficient fine-tuning; why you can adapt huge models on one GPU.
  • Chain-of-Thought Prompting (2022): arxiv.org/abs/2201.11903. "Let's think step by step" unlocked reasoning. The seed of today's thinking models.
  • InstructGPT / "Training LMs to follow instructions with human feedback" (2022): arxiv.org/abs/2203.02155. RLHF; the recipe that turned GPT-3 into something usable and became ChatGPT. Landmark.
  • Chinchilla / "Training Compute-Optimal LLMs" (2022): arxiv.org/abs/2203.15556. Rewrote the scaling rules: most models were undertrained on data. Reshaped every training budget after it.
  • ReAct (2022): arxiv.org/abs/2210.03629. Interleaves reasoning and tool use. The conceptual blueprint for agents.
  • Constitutional AI (2022): arxiv.org/abs/2212.08073. RLAIF: replace human harm labels with a model critiquing against a written constitution. The basis of Anthropic's alignment approach.
  • LLaMA (2023): arxiv.org/abs/2302.13971. Efficient open weights that ignited the open-source LLM ecosystem. Landmark.
  • GPT-4 Technical Report (2023): arxiv.org/abs/2303.08774. The capability jump that mainstreamed AI; notable also for what it does not disclose.
  • Toolformer (2023): arxiv.org/abs/2302.04761. Models teaching themselves to call APIs. Foundational for tool use.
  • DPO / Direct Preference Optimization (2023): arxiv.org/abs/2305.18290. Preference tuning without the RL; simpler and now used everywhere.
  • QLoRA (2023): arxiv.org/abs/2305.14314. 4-bit fine-tuning of 65B models on a single GPU; democratized adaptation.

Efficiency, agents & multimodality (2021–2025)

  • CLIP / "Learning Transferable Visual Models From Natural Language Supervision" (2021): arxiv.org/abs/2103.00020. Contrastive image-text pretraining; the backbone of modern multimodal models. Landmark.
  • FlashAttention (2022): arxiv.org/abs/2205.14135. IO-aware exact attention; the kernel that made long context practical.
  • Mistral 7B (2023): arxiv.org/abs/2310.06825. Small-model efficiency (GQA, sliding-window attention) that performed well above its size.
  • Llama 2 (2023): arxiv.org/abs/2307.09288. Open weights plus an RLHF-tuned chat model; the release that seeded the local-model ecosystem (the list has LLaMA 1 above).
  • Mixtral / Mixture-of-Experts (2024): arxiv.org/abs/2401.04088. Sparse MoE at open scale; the architecture behind much frontier-model efficiency.
  • Mamba / State Space Models (2023): arxiv.org/abs/2312.00752. A sub-quadratic challenger to attention; a leading candidate for what comes after Transformers.
  • DeepSeek-V3 (2024): arxiv.org/abs/2412.19437. The 671B open-weight MoE base that R1 is built on; efficient training at frontier scale.
  • DeepSeek-R1 (2025): arxiv.org/abs/2501.12948. Showed strong reasoning can emerge from RL alone, open-weight; a watershed for open reasoning models.
  • RAG / "Retrieval-Augmented Generation" (2020, read it now): arxiv.org/abs/2005.11401. Grounding generation in retrieved documents. The backbone of most production LLM apps.

How to read the frontier: Papers now ship faster than you can read them. Follow People to Follow (X/Twitter) researchers for the "read this one" signal, use alphaXiv or Hugging Face's Daily Papers to triage, and skim abstracts + figures before committing to a full read. 🔄

Link copied!