RLHF: Reinforcement Learning from Human Feedback
The gap between 'predicts text well' and 'is actually helpful'
A raw pre-trained LLM is extremely good at one thing: predicting the next token in a way that's statistically consistent with its training corpus. That's not the same as being a good assistant. A pre-trained model will happily continue a question with more questions, ramble, generate plausible-sounding misinformation, or produce toxic content, because none of that was penalized during pre-training — the model was only ever rewarded for matching the statistics of internet text.
RLHF (Reinforcement Learning from Human Feedback) is the technique that closes this gap. It's the process that turned raw pre-trained models like GPT into ChatGPT, and is a core part of how Claude was trained to be helpful, honest, and harmless. RLHF doesn't teach the model new facts — it reshapes *how* the model uses what it already knows.
Step 1: Supervised fine-tuning (SFT)
Before RLHF proper begins, the pre-trained model typically goes through supervised fine-tuning: human labelers write high-quality example responses to a range of prompts ("Explain photosynthesis to a 10-year-old", "Write a professional email declining a meeting"), and the model is fine-tuned on these demonstrations using standard supervised learning. This gives the model a reasonable starting point — it now roughly knows what a helpful, well-formatted response looks like — before the more nuanced reward-based refinement begins.
Step 2: Training a reward model from human preferences
Writing a mathematical formula for "good response" is essentially impossible — helpfulness, tone, safety, and correctness are all fuzzy, contextual, and often in tension. RLHF's key trick is to not write that formula by hand, but to *learn* it from human judgment.
Labelers are shown a prompt along with multiple candidate responses from the model, and asked to rank them from best to worst (or simply pick which of two is better — pairwise comparison is easier and more reliable for humans than absolute scoring). This produces a large dataset of preference pairs. A separate neural network, the reward model, is then trained to predict these human preferences: given a prompt and a response, output a scalar score such that preferred responses score higher than rejected ones. Once trained, the reward model becomes a stand-in for human judgment that can score responses automatically, at scale, without a human in the loop for every single evaluation.
Step 3: Optimizing the policy with PPO
With a trained reward model in hand, the SFT model is now treated as an RL policy and further trained to maximize the reward model's score, using PPO (Proximal Policy Optimization) — the same class of policy-based RL algorithm introduced in the previous lesson.
The loop: the policy generates a response to a prompt, the reward model scores that response, and PPO nudges the policy's parameters to make higher-scoring responses more likely in the future. Critically, PPO includes a constraint (implemented as a KL-divergence penalty against the original SFT model) that stops the policy from drifting too far from its starting point in any single update. Without this constraint, the policy could learn to exploit quirks in the reward model — generating outputs that score artificially high without actually being good responses, a failure mode called reward hacking.
Note
Note
Beyond classic RLHF: DPO and other variants
Classic RLHF with PPO is complex to implement correctly and computationally expensive — it requires running and coordinating four separate models (policy, reference/SFT model, reward model, and a value function). Newer techniques like DPO (Direct Preference Optimization) achieve similar alignment goals by reformulating the preference learning problem so it can be optimized directly with a simpler supervised-style loss on preference pairs, skipping the separate reward model and the RL loop entirely. DPO and its variants have become popular for their simplicity, though PPO-based RLHF pipelines remain in use, particularly at labs with the infrastructure to run them well.
SFT vs. RLHF vs. DPO
| Method | What it needs | Complexity | Learns from |
|---|---|---|---|
| Supervised fine-tuning (SFT) | Human-written demonstrations | Low | Direct imitation of good examples |
| RLHF (reward model + PPO) | Preference rankings + RL infrastructure | High — 4 coordinated models | Human preferences, via reward model + RL |
| DPO | Preference pairs only | Medium — single supervised-style loss | Human preferences, directly |
What's next
RLHF sits on top of the reinforcement learning fundamentals from the previous lesson and the fine-tuning techniques covered earlier in this track. To see how alignment quality gets measured after training, revisit Evaluating LLMs — human preference evaluation and LLM-as-judge are directly downstream of the same ideas covered here.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.