[1hr Talk] Intro to Large Language Models
Andrej Karpathy
Andrej Karpathy's hour-long talk explaining LLMs as a compressed model of internet text, covering pretraining, finetuning into an assistant, and emerging capabilities plus security issues like jailbreaks and prompt injection.
More resources on Large Language Models
The TWIML AI Podcast
Long-running interview podcast hosted by Sam Charrington in which ML and AI researchers and practitioners discuss their work on deep learning, natural language processing, neural networks and data science, helping listeners follow current research and how it is applied in industry.
The Illustrated Transformer by Jay Alammar
A brilliant visual explanation of the Transformer architecture, making complex concepts much easier to grasp. This is often cited as a first step for many learners.
"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (Devlin et al., 2018)
Introduced BERT, a breakthrough in pre-training language representations using a Transformer encoder, significantly impacting the development of LLMs.
"Attention Is All You Need" (Vaswani et al., 2017)
The seminal paper that introduced the Transformer architecture, which is the foundation of almost all modern Large Language Models (LLMs). Essential reading for anyone serious about understanding LLMs.
Hugging Face NLP Course
A free, online course by Hugging Face, a leading company in NLP tools and models. It covers transformers, fine-tuning, and using the Hugging Face ecosystem, which is crucial for practical LLM work.
Speech and Language Processing
Free draft third edition of the standard NLP textbook, covering n-gram and neural language models, word embeddings, transformers, sequence labeling, parsing, machine translation and speech recognition. Readers gain the theory needed to build and evaluate modern language-processing systems.