---
title: Large Language Models
description: Explore the world of large-scale language models like GPT, BERT, and their variants. Learn about pre-training, fine-tuning, prompt engineering, and deploying LLMs for various applications.
category: programming-tech
subcategory: artificial-intelligence
difficulty: intermediate, advanced
url: /subject/large-language-models
---

# Large Language Models

Explore the world of large-scale language models like GPT, BERT, and their variants. Learn about pre-training, fine-tuning, prompt engineering, and deploying LLMs for various applications.

## Where to start

Start with Andrej Karpathy's free hour-long talk Intro to Large Language Models, which explains pretraining, fine-tuning into an assistant and security issues such as prompt injection. If you only use one resource, make it Sebastian Raschka's Build a Large Language Model (From Scratch), a Manning book that codes a GPT-style model in PyTorch step by step. For fine-tuning, quantisation, RAG and deployment, continue with Maxime Labonne's LLM Course on GitHub.

## Available Resources

2 Books • 3 Courses • 7 Websites • 3 Papers

## Papers

### 1. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

**Author:** Jason Wei et al.

Shows that prompting a large model with a few worked examples containing intermediate reasoning steps substantially improves arithmetic, commonsense, and symbolic reasoning, and that the effect emerges only at sufficient model scale. The origin of chain-of-thought prompting.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://lnkd.in/gaK5CXzD

**Tags:** prompting, reasoning, llm, few-shot, emergent-abilities

### 2. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" (Devlin et al., 2018)

**Author:** Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova

Introduced BERT, a breakthrough in pre-training language representations using a Transformer encoder, significantly impacting the development of LLMs.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/1810.04805

**Tags:** bert, transformers, pretraining, transfer-learning, nlp

### 3. "Attention Is All You Need" (Vaswani et al., 2017)

**Author:** Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin

The seminal paper that introduced the Transformer architecture, which is the foundation of almost all modern Large Language Models (LLMs). Essential reading for anyone serious about understanding LLMs.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/1706.03762

**Tags:** transformers, attention-mechanism, sequence-modeling, foundational-paper, nlp

## Websites

### 1. Deep (Learning) Focus

**Author:** Cameron R. Wolfe

Cameron R. Wolfe's newsletter, which takes one deep learning research topic at a time and walks through several papers on it in sequence, supplying the background needed to read the primary literature yourself.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://cameronrwolfe.substack.com/

**Tags:** deep-learning, llm, research-summaries, transformers, language-models

### 2. LLM Course

Maxime Labonne's open GitHub curriculum covers three tracks: LLM fundamentals, building models, and deploying them, with Colab notebooks on fine-tuning, quantization, RAG, and evaluation. Learners finish able to train and ship their own models.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://github.com/mlabonne/llm-course

**Tags:** large-language-models, fine-tuning, rag, quantization, notebooks

### 3. Awesome Generative AI Guide

**Author:** Aishwarya Naresh Reganti

Aggregates generative AI learning paths, free courses such as an agentic AI crash course and applied LLM series, interview question banks, code notebooks, and monthly research paper roundups into one continuously updated GitHub hub.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://github.com/aishwaryanr/awesome-generative-ai-guide

**Tags:** generative-ai, llm, llm-courses, interview-prep, learning-path

### 4. Hands-On Large Language Models

Visual, code-first material from Jay Alammar and Maarten Grootendorst covering tokenization, embeddings, semantic search, text classification, clustering, prompt engineering, and fine-tuning, aimed at practitioners who want intuition alongside working implementations.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://lnkd.in/dxaVF86w

**Tags:** large-language-models, embeddings, fine-tuning, prompt-engineering, nlp

### 5. Prompt Engineering Guide

Reference covering prompting techniques with worked examples: zero-shot and few-shot, chain-of-thought, self-consistency, ReAct, and tree of thoughts, plus model-specific notes and adversarial prompting risks. Readers finish able to choose techniques deliberately.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://lnkd.in/gJjGbxQr

**Tags:** prompt-engineering, chain-of-thought, few-shot, llm-techniques, react

### 6. The Illustrated Transformer by Jay Alammar

A brilliant visual explanation of the Transformer architecture, making complex concepts much easier to grasp. This is often cited as a first step for many learners.

**Difficulty:** Beginner | **Language:** English | **Price:** Free

**Link:** https://jalammar.github.io/illustrated-transformer/

**Tags:** transformers, attention-mechanism, neural-networks, visual-explainer, nlp

### 7. GenAI Handbook: A Roadmap for Learning Resources

A comprehensive open-source handbook acting as a roadmap for understanding the core concepts of modern artificial intelligence systems, with a strong focus on Generative AI.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://genai-handbook.github.io/

**Tags:** generative-ai, llms, transformers, learning-roadmap, deep-learning

## Youtubes

### 1. Stanford CS229 I Machine Learning I Building Large Language Models (LLMs)

A Stanford CS229 lecture surveying the full LLM pipeline: tokenization, transformer architecture, training data and scaling laws, evaluation, and post-training alignment. It explains the engineering decisions and costs behind a production model.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://www.youtube.com/watch?v=9vM4p9NN0Ts

**Tags:** llm, transformers, pretraining, scaling-laws, rlhf

### 2. [1hr Talk] Intro to Large Language Models

Andrej Karpathy's hour-long talk explaining LLMs as a compressed model of internet text, covering pretraining, finetuning into an assistant, and emerging capabilities plus security issues like jailbreaks and prompt injection.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://www.youtube.com/watch?v=zjkBMFhNj_g

**Tags:** llm, pretraining, fine-tuning, prompt-injection, ai-fundamentals

## Books

### 1. Build a Large Language Model (From Scratch)

**Author:** Sebastian Raschka

Sebastian Raschka's Manning book walks through coding a GPT-style model in PyTorch step by step: tokenization, attention, the transformer block, pretraining on unlabeled text, then fine-tuning for classification and instruction following.

**Difficulty:** Intermediate | **Language:** English | **Price:** Paid

**Link:** https://www.amazon.com/dp/1633437167?tag=edmonddante07-20

**Tags:** llm, transformers, pytorch, fine-tuning, from-scratch

### 2. Speech and Language Processing

**Author:** Daniel Jurafsky, James H. Martin

Free draft third edition of the standard NLP textbook, covering n-gram and neural language models, word embeddings, transformers, sequence labeling, parsing, machine translation and speech recognition. Readers gain the theory needed to build and evaluate modern language-processing systems.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://web.stanford.edu/~jurafsky/slp3/

**Tags:** nlp, computational-linguistics, language-models, transformers, speech-recognition

## Courses

### 1. CS224N: Natural Language Processing with Deep Learning

Learn cutting-edge Natural Language Processing with Deep Learning in Stanford's CS224N course by Christopher Manning.

**Difficulty:** Advanced | **Price:** Free

**Link:** https://web.stanford.edu/class/cs224n/

**Tags:** deep-learning, word-embeddings, transformers, stanford, neural-machine-translation

### 2. ARENA: Alignment Research Engineer Accelerator Curriculum

**Author:** Callum McDougall

Five chapters of PyTorch exercises with solutions: deep learning foundations, transformer interpretability with TransformerLens, reinforcement learning, LLM evaluations and alignment science. Working through them leaves you able to replicate interpretability papers and build evaluation harnesses yourself.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://learn.arena.education/

**Tags:** ai-alignment, mechanistic-interpretability, transformerlens, pytorch, reinforcement-learning, llm-evaluations

### 3. Hugging Face NLP Course

A free, online course by Hugging Face, a leading company in NLP tools and models. It covers transformers, fine-tuning, and using the Hugging Face ecosystem, which is crucial for practical LLM work.

**Difficulty:** Beginner | **Language:** English | **Price:** Free

**Link:** https://huggingface.co/course

**Tags:** nlp, transformers, fine-tuning, tokenization, deep-learning

## Podcasts

### 1. The TWIML AI Podcast

**Author:** Sam Charrington

Long-running interview podcast hosted by Sam Charrington in which ML and AI researchers and practitioners discuss their work on deep learning, natural language processing, neural networks and data science, helping listeners follow current research and how it is applied in industry.

**Difficulty:** Intermediate | **Language:** en | **Price:** Free

**Link:** https://twimlai.com

**Tags:** machine-learning, artificial-intelligence, deep-learning, ai-research

---

*This content is part of Dantes.io - Your Treasure Map to Knowledge*

*Curated by humans at Dantes.io. Personal study use welcome; republishing this curation requires permission (team@dantes.io).*

View this page online: https://dantes.io/subject/large-language-models