---
title: "Preference tuning, PPO, DPO, and RLHF intuition"
lesson_id: "08"
---

# Preference tuning, PPO, DPO, and RLHF intuition

- **Lesson ID:** 08
- **Goal:** Preference data is a ranking, not a vibe. Walk RLHF enough to know why DPO exists, then run DPO on a small pair set.
- **Human lesson:** [08-preference-tuning-and-rlhf.html](08-preference-tuning-and-rlhf.html)
- **Phase:** Phase 4 Post-training

## Prerequisites

- The previous week's artifact, or a written note if this is week 1.
- Python 3.11+ and a laptop. CUDA is useful from week 5 onward and not required to read.

## Inputs, outputs, and artifacts

- **Inputs:** The replica from the prior week.
- **Outputs:** Core project 6: a DPO step you can explain on a whiteboard.
- **Artifacts:** Lesson notes plus the code named in the human page.

## Agent build steps

1. State the week 1 principle: you do not need to pretrain an 800B model to understand the shape of the problem; you do need progressively realistic replicas.
2. Teach the topics on the human page without inventing paper URLs or news claims.
3. Keep starter code in fenced blocks that match the human lesson.
4. End with the assignment and the opinion checkpoint.
5. Use Edge FDE tone: ship, do not sightseeing. Opinion checkpoints matter.

## Constraints

Plain spoken English. No quizzes. No em dashes. No invented metrics, dates, or citations. Brand Edge FDE only. Name well-known papers by title only: Attention Is All You Need, InstructGPT, Direct Preference Optimization, ZeRO, PagedAttention/vLLM, FlashAttention, Scaling Laws for Neural Language Models.

## Key concepts

- RLHF is SFT, then a reward model, then PPO with a KL penalty.
- DPO trains on chosen/rejected pairs with a reference model.
- Pair quality is the product. The loss is the mechanism.
- Reward hacking is why you keep a KL or a reference.

## Takeaways

- Run DPO on a small pair set you can read.
- Name InstructGPT and DPO as paper titles, not as brands.
- Describe the behaviour change with examples, not invented metrics.

## Acceptance checks

- Human HTML has Key concepts and Takeaways.
- Opinion checkpoint is present.
- No em dash and no fake URL.
- [ ] Proceed to [lesson 09 brief](09-vllm-internals-and-serving.llms.md).
