---
title: Fine-Tuning LLMs
description: Fine-tuning adapts a pretrained language model to a task or style. You will learn supervised fine-tuning, LoRA and parameter-efficient methods, preference optimisation, datasets and evaluation.
category: programming-tech
subcategory: artificial-intelligence
difficulty: beginner, intermediate, advanced
url: /subject/llm-fine-tuning
---

# Fine-Tuning LLMs

Fine-tuning adapts a pretrained language model to a task or style. You will learn supervised fine-tuning, LoRA and parameter-efficient methods, preference optimisation, datasets and evaluation.

## Available Resources

2 Courses • 2 Websites • 3 Papers

## Courses

### 1. Improving Accuracy of LLM Applications

**Author:** Sharon Zhou, Amit Sangani

Walks through a systematic loop for making LLM applications reliable: build evaluation metrics, apply prompting and self-reflection, then fine-tune with LoRA and memory tuning. The running example is a text-to-SQL agent that hallucinates less each iteration.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://www.deeplearning.ai/courses/improving-accuracy-of-llm-applications

**Tags:** llm-evaluation, fine-tuning, lora, hallucination, text-to-sql

### 2. Hugging Face smol-course: Fine-Tuning Language Models

Hands-on post-training course built around SmolLM3. Units cover instruction tuning, evaluation, preference alignment, vision-language models and reinforcement learning, each with runnable TRL notebooks sized for a single consumer GPU or free Colab.
The only free, structured, start-to-finish course dedicated specifically to fine-tuning rather than to LLMs generally. Actively maintained with new units, uses current TRL APIs, and every exercise runs on hardware a learner already has.

**Difficulty:** Beginner | **Language:** English | **Price:** Free

**Link:** https://huggingface.co/learn/smol-course/unit0/1

**Tags:** supervised-fine-tuning, preference-alignment, trl, smollm, llm-fine-tuning

## Papers

### 1. Direct Preference Optimization: Your Language Model is Secretly a Reward Model

**Author:** Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, Chelsea Finn

Derives a closed-form mapping between reward models and optimal policies, replacing the reward-model-plus-PPO pipeline of RLHF with a single classification loss over preference pairs. The reference behind TRL's DPOTrainer and its ORPO and KTO successors.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2305.18290

**Tags:** dpo, preference-optimization, rlhf, alignment, llm-fine-tuning

### 2. QLoRA: Efficient Finetuning of Quantized LLMs

**Author:** Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer

Dettmers and colleagues combine 4-bit NormalFloat quantisation, double quantisation and paged optimisers with LoRA adapters to fine-tune a 65B model on one 48GB GPU. The technical basis for most consumer-hardware fine-tuning stacks in use today. 
Explains why a 4-bit base model can be fine-tuned without quality collapse, which is the assumption every consumer-GPU tutorial silently relies on. Also contains a still-useful data-quality result: small curated datasets beat large scraped ones.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2305.14314

**Tags:** qlora, quantization, lora, parameter-efficient-fine-tuning, llm-fine-tuning

### 3. LoRA: Low-Rank Adaptation of Large Language Models

**Author:** Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen

Foundational and short. Reading it once removes the guesswork behind rank, alpha and target-module settings that the tutorials present as folklore. Age is irrelevant here — the field's own notation is downstream of this paper. The paper introducing low-rank adaptation: freezing pretrained weights and training rank-decomposition matrices injected into each layer. Reduces GPT-3 175B trainable parameters roughly ten-thousand-fold and adds no inference latency, unlike earlier adapter methods.

**Difficulty:** Advanced | **Language:** English | **Price:** Free

**Link:** https://arxiv.org/abs/2106.09685

**Tags:** lora, low-rank-adaptation, parameter-efficient-fine-tuning, transformers, llm-fine-tuning

## Websites

### 1. TRL: Transformer Reinforcement Learning Documentation

Reference documentation for the library implementing SFT, DPO, KTO, ORPO, GRPO, reward modelling and distillation trainers. Includes dataset format specifications, memory-reduction and distributed-training how-tos, and PEFT, vLLM and DeepSpeed integration guides.

**Difficulty:** Intermediate | **Language:** English | **Price:** Free

**Link:** https://huggingface.co/docs/trl/index

**Tags:** trl, hugging-face, supervised-fine-tuning, dpo, rlhf, llm-fine-tuning

### 2. Fine-tuning LLMs Guide (Unsloth Documentation)

End-to-end documentation for fine-tuning open models: choosing a base model, LoRA versus QLoRA, dataset formatting, hyperparameter selection, running training in free Colab or Kaggle notebooks, evaluating results, and exporting to GGUF or vLLM.

**Difficulty:** Beginner | **Language:** English | **Price:** Free

**Link:** https://docs.unsloth.ai/get-started/fine-tuning-llms-guide

**Tags:** unsloth, lora, qlora, gguf, llm-fine-tuning

---

*This content is part of Dantes.io - Your Treasure Map to Knowledge*

*Curated by humans at Dantes.io. Personal study use welcome; republishing this curation requires permission (team@dantes.io).*

View this page online: https://dantes.io/subject/llm-fine-tuning