Sources
Every factual claim in Transformers, ELI5 traces back to one of these 38 sources.
- Jurafsky, Daniel, and James H. Martin. Speech and Language Processing (3rd ed. draft, 2026), Chapter 8: “Transformers.” — pt. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10
- Vaswani, Ashish, et al. “Attention Is All You Need.” NeurIPS, 2017. arXiv:1706.03762. — pt. 2, 3, 4, 9
- Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. “Efficient Estimation of Word Representations in Vector Space.” 2013. arXiv:1301.3781. — pt. 2
- Mikolov, Tomas, Wen-tau Yih, and Geoffrey Zweig. “Linguistic Regularities in Continuous Space Word Representations.” NAACL-HLT, 2013. — pt. 2
- Sennrich, Rico, Barry Haddow, and Alexandra Birch. “Neural Machine Translation of Rare Words with Subword Units.” ACL, 2016. arXiv:1508.07909. — pt. 2
- Su, Jianlin, et al. “RoFormer: Enhanced Transformer with Rotary Position Embedding.” 2021. arXiv:2104.09864. — pt. 2
- Levesque, Hector J., Ernest Davis, and Leora Morgenstern. “The Winograd Schema Challenge.” KR, 2012. — pt. 3
- Hendrycks, Dan, and Kevin Gimpel. “Gaussian Error Linear Units (GELUs).” 2016. arXiv:1606.08415. — pt. 4
- Shazeer, Noam. “GLU Variants Improve Transformer.” 2020. arXiv:2002.05202. — pt. 4
- Touvron, Hugo, et al. “LLaMA: Open and Efficient Foundation Language Models.” 2023. arXiv:2302.13971. — pt. 4, 6
- Geva, Mor, Roei Schuster, Jonathan Berant, and Omer Levy. “Transformer Feed-Forward Layers Are Key-Value Memories.” EMNLP, 2021. arXiv:2012.14913. — pt. 4
- Elhage, Nelson, et al. “Toy Models of Superposition.” Transformer Circuits Thread, 2022. — pt. 4, 10
- He, Kaiming, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition.” CVPR, 2016. arXiv:1512.03385. — pt. 5
- Elhage, Nelson, et al. “A Mathematical Framework for Transformer Circuits.” Transformer Circuits Thread, 2021. — pt. 5
- Ba, Jimmy Lei, Jamie Ryan Kiros, and Geoffrey E. Hinton. “Layer Normalization.” 2016. arXiv:1607.06450. — pt. 5
- Xiong, Ruibin, et al. “On Layer Normalization in the Transformer Architecture.” ICML, 2020. arXiv:2002.04745. — pt. 5
- Zhang, Biao, and Rico Sennrich. “Root Mean Square Layer Normalization.” NeurIPS, 2019. arXiv:1910.07467. — pt. 5
- Brown, Tom B., et al. “Language Models are Few-Shot Learners.” NeurIPS, 2020. arXiv:2005.14165. — pts. 5, 9
- Radford, Alec, et al. “Language Models are Unsupervised Multitask Learners.” OpenAI Technical Report, 2019. — pt. 6
- Press, Ofir, and Lior Wolf. “Using the Output Embedding to Improve Language Models.” EACL, 2017. arXiv:1608.05859. — pt. 6
- Holtzman, Ari, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. “The Curious Case of Neural Text Degeneration.” ICLR, 2020. arXiv:1904.09751. — pt. 7
- Nguyen, Minh Nhat, et al. “Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs.” 2024. arXiv:2407.01082. — pt. 7
- Keskar, Nitish Shirish, et al. “CTRL: A Conditional Transformer Language Model for Controllable Generation.” 2019. arXiv:1909.05858. — pt. 7
- Rumelhart, David E., Geoffrey E. Hinton, and Ronald J. Williams. “Learning representations by back-propagating errors.” Nature 323 (1986): 533–536. — pt. 8
- Ouyang, Long, et al. “Training Language Models to Follow Instructions with Human Feedback.” NeurIPS, 2022. arXiv:2203.02155. — pt. 8
- Rafailov, Rafael, et al. “Direct Preference Optimization: Your Language Model is Secretly a Reward Model.” NeurIPS, 2023. arXiv:2305.18290. — pt. 8
- DeepSeek-AI. “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” 2025. arXiv:2501.12948. — pt. 8
- OpenAI. “Learning to Reason with LLMs.” 2024. — pt. 8
- Kaplan, Jared, et al. “Scaling Laws for Neural Language Models.” 2020. arXiv:2001.08361. — pt. 9
- Hoffmann, Jordan, et al. “Training Compute-Optimal Large Language Models.” NeurIPS, 2022. arXiv:2203.15556. — pt. 9
- Ainslie, Joshua, et al. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.” EMNLP, 2023. arXiv:2305.13245. — pt. 9
- Gemini Team, Google. “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.” 2024. arXiv:2403.05530. — pt. 9
- Hu, Edward J., et al. “LoRA: Low-Rank Adaptation of Large Language Models.” 2021. arXiv:2106.09685. — pt. 9
- Dettmers, Tim, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.” NeurIPS, 2022. arXiv:2208.07339. — pt. 9
- Fedus, William, Barret Zoph, and Noam Shazeer. “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.” 2021. arXiv:2101.03961. — pt. 9
- Wei, Jason, et al. “Emergent Abilities of Large Language Models.” Transactions on Machine Learning Research, 2022. arXiv:2206.07682. — pt. 9
- Schaeffer, Rylan, Brando Miranda, and Sanmi Koyejo. “Are Emergent Abilities of Large Language Models a Mirage?” NeurIPS, 2023. arXiv:2304.15004. — pt. 9
- Olsson, Catherine, et al. “In-context Learning and Induction Heads.” Transformer Circuits Thread, 2022. — pt. 10
- nostalgebraist. “interpreting GPT: the logit lens.” LessWrong, 2020. — pt. 10
- Bricken, Trenton, et al. “Towards Monosemanticity: Decomposing Language Models With Dictionary Learning.” Transformer Circuits Thread, 2023. — pt. 10
- Dehghani, Mostafa, et al. “Universal Transformers.” ICLR, 2019. arXiv:1807.03819. — recurrence article
- Geiping, Jonas, et al. “Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.” 2025. arXiv:2502.05171. — recurrence article
- OpenAI. “What Parameter Golf taught us.” May 12, 2026. — recurrence article
- Pachocki, Jakub. “An Alien Mind.” OpenAI, September 6, 2026. — recurrence article