---
title: Generative Image & Video Models
description: Diffusion and related models generate images and video from text. You will learn how they work, how to run and prompt them, control techniques like ControlNet and LoRA, and the creative and legal issues around them.
category: programming-tech
subcategory: artificial-intelligence
difficulty: beginner, intermediate, advanced
url: /subject/generative-image-models
---

# Generative Image & Video Models

Diffusion and related models generate images and video from text. You will learn how they work, how to run and prompt them, control techniques like ControlNet and LoRA, and the creative and legal issues around them.

## Available Resources

1 Books • 1 Courses • 2 Websites • 4 Papers

## Papers

### 1. Extracting Training Data from Diffusion Models

**Author:** Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, Eric Wallace

A USENIX Security 2023 study that recovers over a thousand memorized training images from Stable Diffusion, Imagen and DALL-E 2, including photographs of identifiable people and trademarked logos, and measures which design choices worsen memorization. It quantifies how often models reproduce training images and reports the copyright status of what it extracted.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2301.13188

**Tags:** diffusion-models, memorization, privacy, ml-security, generative-models

### 2. Adding Conditional Control to Text-to-Image Diffusion Models (ControlNet)

**Author:** Lvmin Zhang, Anyi Rao, Maneesh Agrawala

The paper introducing ControlNet, which adds spatial conditioning such as edges, depth, segmentation and human pose to a frozen text-to-image diffusion model using trainable copies of its encoder joined by zero-initialized convolutions.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2302.05543

**Tags:** controlnet, diffusion-models, text-to-image, conditional-generation, computer-vision

### 3. High-Resolution Image Synthesis with Latent Diffusion Models

**Author:** Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer

The CVPR 2022 paper behind Stable Diffusion. It runs the diffusion process in the latent space of a pretrained autoencoder and adds cross-attention conditioning, cutting compute enough for consumer hardware while handling text, layout, inpainting and super-resolution. Every practical tool in this topic — SD 1.5, SDXL, and their descendants — is an implementation of this paper.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2112.10752

**Tags:** latent-diffusion, stable-diffusion, text-to-image, generative-models, computer-vision

### 4. Step-by-Step Diffusion: An Elementary Tutorial

**Author:** Preetum Nakkiran, Arwen Bradley, Hattie Zhou, Madhu Advani

A 35-page tutorial deriving diffusion models and flow matching from first principles, deliberately avoiding SDEs, ELBOs and score functions. Assumes probability, calculus and linear algebra, and gives pseudocode for the sampling algorithms it derives.
The single best free entry point to the mathematics. Needs no stochastic calculus or variational-bound machinery, so a reader with undergraduate probability can follow a complete derivation of DDPM and DDIM sampling and then implement it.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2406.08929

**Tags:** diffusion-models, flow-matching, ddpm, tutorial, mathematics

## Books

### 1. Hands-On Generative AI with Transformers and Diffusion Models

**Author:** Omar Sanseviero, Pedro Cuenca, Apolinário Passos, Jonathan Whitaker

A practitioner's book from the Hugging Face Diffusers team covering how diffusion models work, Stable Diffusion internals, sampling and guidance, fine-tuning with LoRA and DreamBooth, ControlNet and inpainting workflows, and chapters on audio and video generation.

**Difficulty:** Intermediate | **Language:** English | **Price:** Paid

**Link:** https://www.amazon.com/dp/1098149246?tag=edmonddante07-20

**Tags:** diffusion-models, generative-ai, stable-diffusion, transformers, fine-tuning

## Courses

### 1. Practical Deep Learning for Coders, Part 2: From Deep Learning Foundations to Stable Diffusion

**Author:** Jeremy Howard, Jonathan Whitaker, Tanishq Abraham, Wasim Lorgat

Over thirty hours of free video lessons rebuilding Stable Diffusion from scratch, starting at matrix multiplication and autograd and working up through DDPM, DDIM, conditional sampling, textual inversion and DreamBooth, with paper-reading practice throughout.
The only free course taking a reader from tensor primitives to a working diffusion implementation without hand-waving, and the depth means the understanding survives model churn. Long and demanding.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://course.fast.ai/Lessons/part2.html

**Tags:** stable-diffusion, diffusion-models, deep-learning, pytorch, implementation-from-scratch

## Websites

### 1. Hugging Face Diffusers Documentation

Official documentation for the Diffusers library: pipelines for image, video and audio generation, scheduler choices, LoRA and adapter loading, ControlNet usage, DreamBooth and textual-inversion training scripts, quantization and memory-offloading guides, and conceptual explanations of each component.
The practical half of the topic — running models, choosing schedulers, applying LoRA and ControlNet, fine-tuning — has no better free source, and the maintainers' own docs avoid the huge volume of blog tutorials written against long-dead API versions.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://huggingface.co/docs/diffusers/index

**Tags:** diffusers, stable-diffusion, controlnet, hugging-face, documentation

### 2. What are Diffusion Models?

**Author:** Lilian Weng

A continuously revised technical survey covering forward and reverse diffusion processes, the DDPM training objective, score-based formulations, classifier and classifier-free guidance, latent diffusion, ControlNet, diffusion transformers, and acceleration methods including DDIM and consistency models.
The map of the whole field in one page, with consistent notation across papers that each use their own. Kept updated since 2021 rather than abandoned, so it now covers guidance, latent diffusion, ControlNet and DiT in the same framework. The reference you return to after the Nakkiran tutorial teaches the core derivation.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://lilianweng.github.io/posts/2021-07-11-diffusion-models/

**Tags:** diffusion-models, score-based-models, classifier-free-guidance, generative-models, survey

---

*This content is part of Dantes.io - Your Treasure Map to Knowledge*

*Curated by humans at Dantes.io. Personal study use welcome; republishing this curation requires permission (team@dantes.io).*

View this page online: https://dantes.io/subject/generative-image-models