AI & MLDeep
Advanced

RLHF: Reinforcement Learning from Human Feedback

12 min read

Learn
Deep Reading
Estimated 12 mins
Prereq
Advanced
Requires advanced math/coding
Interactive
Static Playbook
Static guide & reference tables

The gap between 'predicts text well' and 'is actually helpful'

A raw pre-trained LLM is extremely good at one thing: predicting the next token in a way that's statistically consistent with its training corpus. That's not the same as being a good assistant. A pre-trained model will happily continue a question with more questions, ramble, generate plausible-sounding misinformation, or produce toxic content, because none of that was penalized during pre-training — the model was only ever rewarded for matching the statistics of internet text.

RLHF (Reinforcement Learning from Human Feedback) is the technique that closes this gap. It's the process that turned raw pre-trained models like GPT into ChatGPT, and is a core part of how Claude was trained to be helpful, honest, and harmless. RLHF doesn't teach the model new facts — it reshapes *how* the model uses what it already knows.

Step 1: Supervised fine-tuning (SFT)

Before RLHF proper begins, the pre-trained model typically goes through supervised fine-tuning: human labelers write high-quality example responses to a range of prompts ("Explain photosynthesis to a 10-year-old", "Write a professional email declining a meeting"), and the model is fine-tuned on these demonstrations using standard supervised learning. This gives the model a reasonable starting point — it now roughly knows what a helpful, well-formatted response looks like — before the more nuanced reward-based refinement begins.

Step 2: Training a reward model from human preferences

Writing a mathematical formula for "good response" is essentially impossible — helpfulness, tone, safety, and correctness are all fuzzy, contextual, and often in tension. RLHF's key trick is to not write that formula by hand, but to *learn* it from human judgment.

Labelers are shown a prompt along with multiple candidate responses from the model, and asked to rank them from best to worst (or simply pick which of two is better — pairwise comparison is easier and more reliable for humans than absolute scoring). This produces a large dataset of preference pairs. A separate neural network, the reward model, is then trained to predict these human preferences: given a prompt and a response, output a scalar score such that preferred responses score higher than rejected ones. Once trained, the reward model becomes a stand-in for human judgment that can score responses automatically, at scale, without a human in the loop for every single evaluation.

python

Step 3: Optimizing the policy with PPO

With a trained reward model in hand, the SFT model is now treated as an RL policy and further trained to maximize the reward model's score, using PPO (Proximal Policy Optimization) — the same class of policy-based RL algorithm introduced in the previous lesson.

The loop: the policy generates a response to a prompt, the reward model scores that response, and PPO nudges the policy's parameters to make higher-scoring responses more likely in the future. Critically, PPO includes a constraint (implemented as a KL-divergence penalty against the original SFT model) that stops the policy from drifting too far from its starting point in any single update. Without this constraint, the policy could learn to exploit quirks in the reward model — generating outputs that score artificially high without actually being good responses, a failure mode called reward hacking.

text

Note

Because the reward model is an imperfect proxy for true human judgment, a policy optimized aggressively against it can learn to game that proxy rather than genuinely improve — for example, producing overly long, hedge-everything responses because the reward model happens to correlate length with thoroughness. This is why the KL constraint against the SFT model matters, and why reward models themselves need continual evaluation against fresh human judgments rather than being trusted indefinitely.

Note

When people say a model was 'aligned to be helpful and harmless,' RLHF (or close variants of it) is usually the mechanism they mean. It's the stage where a model goes from a raw text predictor to something that refuses harmful requests, admits uncertainty, follows formatting instructions, and generally behaves like an assistant rather than a document-completion engine.

Beyond classic RLHF: DPO and other variants

Classic RLHF with PPO is complex to implement correctly and computationally expensive — it requires running and coordinating four separate models (policy, reference/SFT model, reward model, and a value function). Newer techniques like DPO (Direct Preference Optimization) achieve similar alignment goals by reformulating the preference learning problem so it can be optimized directly with a simpler supervised-style loss on preference pairs, skipping the separate reward model and the RL loop entirely. DPO and its variants have become popular for their simplicity, though PPO-based RLHF pipelines remain in use, particularly at labs with the infrastructure to run them well.

SFT vs. RLHF vs. DPO

MethodWhat it needsComplexityLearns from
Supervised fine-tuning (SFT)Human-written demonstrationsLowDirect imitation of good examples
RLHF (reward model + PPO)Preference rankings + RL infrastructureHigh — 4 coordinated modelsHuman preferences, via reward model + RL
DPOPreference pairs onlyMedium — single supervised-style lossHuman preferences, directly

What's next

RLHF sits on top of the reinforcement learning fundamentals from the previous lesson and the fine-tuning techniques covered earlier in this track. To see how alignment quality gets measured after training, revisit Evaluating LLMs — human preference evaluation and LLM-as-judge are directly downstream of the same ideas covered here.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.