Module 2: Fine-Tuning Techniques
Preference optimization: RLHF and DPO
Refining behavior toward human preferences — and why DPO simplified it.
Loading lesson…
Discussion (0)
Ask a question or share what worked for you. Comments are reviewed before they appear.
Log in to join the discussion and ask questions about this lesson.
No comments yet. Be the first to start the discussion!