principles.fyi · Transformers, ELI5

Sources

Every factual claim in this topic traces back to one of these 38 sources — the textbook spine plus the primary papers.

  1. Jurafsky, Daniel, and James H. Martin. Speech and Language Processing (3rd ed. draft, 2026), Chapter 8: "Transformers." pt. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
  2. Vaswani, Ashish, et al. "Attention Is All You Need." NeurIPS, 2017. arXiv:1706.03762. pt. 2, 3, 4, 9
  3. Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. "Efficient Estimation of Word Representations in Vector Space." 2013. arXiv:1301.3781. pt. 2
  4. Mikolov, Tomas, Wen-tau Yih, and Geoffrey Zweig. "Linguistic Regularities in Continuous Space Word Representations." NAACL-HLT, 2013. pt. 2
  5. Sennrich, Rico, Barry Haddow, and Alexandra Birch. "Neural Machine Translation of Rare Words with Subword Units." ACL, 2016. arXiv:1508.07909. pt. 2
  6. Su, Jianlin, et al. "RoFormer: Enhanced Transformer with Rotary Position Embedding." 2021. arXiv:2104.09864. pt. 2
  7. Levesque, Hector J., Ernest Davis, and Leora Morgenstern. "The Winograd Schema Challenge." KR, 2012. pt. 3
  8. Hendrycks, Dan, and Kevin Gimpel. "Gaussian Error Linear Units (GELUs)." 2016. arXiv:1606.08415. pt. 4
  9. Shazeer, Noam. "GLU Variants Improve Transformer." 2020. arXiv:2002.05202. pt. 4
  10. Touvron, Hugo, et al. "LLaMA: Open and Efficient Foundation Language Models." 2023. arXiv:2302.13971. pt. 4, 6
  11. Geva, Mor, Roei Schuster, Jonathan Berant, and Omer Levy. "Transformer Feed-Forward Layers Are Key-Value Memories." EMNLP, 2021. arXiv:2012.14913. pt. 4
  12. Elhage, Nelson, et al. "Toy Models of Superposition." Transformer Circuits Thread, 2022. pt. 4, 10
  13. He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. "Deep Residual Learning for Image Recognition." CVPR, 2016. arXiv:1512.03385. pt. 5
  14. Elhage, Nelson, et al. "A Mathematical Framework for Transformer Circuits." Transformer Circuits Thread, 2021. pt. 5
  15. Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. "Layer Normalization." 2016. arXiv:1607.06450. pt. 5
  16. Xiong, Ruibin, et al. "On Layer Normalization in the Transformer Architecture." ICML, 2020. arXiv:2002.04745. pt. 5
  17. Zhang, Biao, and Rico Sennrich. "Root Mean Square Layer Normalization." NeurIPS, 2019. arXiv:1910.07467. pt. 5
  18. Brown, Tom B., et al. "Language Models are Few-Shot Learners." NeurIPS, 2020. arXiv:2005.14165. pt. 5
  19. Radford, Alec, et al. "Language Models are Unsupervised Multitask Learners." OpenAI Technical Report, 2019. pt. 6
  20. Press, Ofir, and Lior Wolf. "Using the Output Embedding to Improve Language Models." EACL, 2017. arXiv:1608.05859. pt. 6
  21. Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. "The Curious Case of Neural Text Degeneration." ICLR, 2020. arXiv:1904.09751. pt. 7
  22. Nguyen, Minh Nhat, et al. "Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs." 2024. arXiv:2407.01082. pt. 7
  23. Keskar, Nitish Shirish, et al. "CTRL: A Conditional Transformer Language Model for Controllable Generation." 2019. arXiv:1909.05858. pt. 7
  24. Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. "Learning representations by back-propagating errors." Nature 323 (1986): 533–536. pt. 8
  25. Ouyang, Long, et al. "Training Language Models to Follow Instructions with Human Feedback." NeurIPS, 2022. arXiv:2203.02155. pt. 8
  26. Rafailov, Rafael, et al. "Direct Preference Optimization: Your Language Model is Secretly a Reward Model." NeurIPS, 2023. arXiv:2305.18290. pt. 8
  27. DeepSeek-AI. "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning." 2025. arXiv:2501.12948. pt. 8
  28. OpenAI. "Learning to Reason with LLMs." 2024. pt. 8
  29. Kaplan, Jared, et al. "Scaling Laws for Neural Language Models." 2020. arXiv:2001.08361. pt. 9
  30. Hoffmann, Jordan, et al. "Training Compute-Optimal Large Language Models." NeurIPS, 2022. arXiv:2203.15556. pt. 9
  31. Ainslie, Joshua, et al. "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints." EMNLP, 2023. arXiv:2305.13245. pt. 9
  32. Gemini Team, Google. "Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context." 2024. arXiv:2403.05530. pt. 9
  33. Hu, Edward J., et al. "LoRA: Low-Rank Adaptation of Large Language Models." 2021. arXiv:2106.09685. pt. 9
  34. Dettmers, Tim, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. "LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale." NeurIPS, 2022. arXiv:2208.07339. pt. 9
  35. Fedus, William, Barret Zoph, and Noam Shazeer. "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity." 2021. arXiv:2101.03961. pt. 9
  36. Olsson, Catherine, et al. "In-context Learning and Induction Heads." Transformer Circuits Thread, 2022. pt. 10
  37. nostalgebraist. "interpreting GPT: the logit lens." LessWrong, 2020. pt. 10
  38. Bricken, Trenton, et al. "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning." Transformer Circuits Thread, 2023. pt. 10