Reinforcement Learning Basics
Learning by trial and error
Reinforcement learning (RL) is a fundamentally different paradigm from the supervised learning covered elsewhere in this track. There's no labeled dataset of correct answers. Instead, an agent interacts with an environment, takes actions, and receives rewards — numeric feedback signals indicating how good or bad an outcome was. The agent's entire goal is to learn a strategy, called a policy, that maximizes the total reward it accumulates over time.
This is the framework behind game-playing systems like AlphaGo, robotic control, and — critically for modern AI — RLHF, the technique used to align chatbots like ChatGPT and Claude with human preferences (covered in the next lesson).
The core loop: agent, environment, state, action, reward
RL formalizes the problem as a loop. At each timestep, the agent observes the current state $s$ of the environment, chooses an action $a$ according to its policy $\pi$, the environment transitions to a new state $s'$ and returns a reward $r$, and the cycle repeats. This is typically modeled as a Markov Decision Process (MDP): the assumption that the next state and reward depend only on the current state and action, not on the full history that led there.
The agent's objective is to maximize the expected cumulative reward, often discounted so that rewards received sooner count more than rewards received later: $G_t = r_t + \gamma r_{t+1} + \gamma^2 r_{t+2} + \dots$, where $\gamma \in [0,1]$ is the discount factor.
Note
Value-based methods: learning how good states and actions are
Value-based methods learn a function that estimates expected future reward, then derive a policy from it implicitly — always pick the action with the highest estimated value. The Q-function, $Q(s, a)$, estimates the expected cumulative reward of taking action $a$ in state $s$ and then acting optimally afterward.
Q-learning learns this function iteratively using the Bellman equation as an update rule: after observing a transition $(s, a, r, s')$, update the estimate toward $r + \gamma \max_{a'} Q(s', a')$ — the immediate reward plus the discounted value of the best action available from the next state. Repeated over many episodes, these estimates converge toward accurate values, and the resulting policy (always take the highest-Q action) converges toward optimal behavior.
Policy-based methods: learning the strategy directly
Policy-based methods skip value estimation and instead directly parameterize and optimize the policy $\pi(a \mid s)$ — a function (often a neural network) that outputs a probability distribution over actions given a state. Training adjusts the policy's parameters to increase the probability of actions that led to high reward and decrease the probability of actions that led to low reward, typically via a gradient estimator like REINFORCE or more stable modern variants like PPO (Proximal Policy Optimization).
Policy-based methods handle continuous action spaces more naturally than Q-learning (which requires discretizing or maximizing over actions) and tend to be more stable in high-dimensional problems — which is exactly why PPO became the workhorse algorithm for RLHF, the topic of the next lesson.
Value-based vs. policy-based RL
| Approach | Learns | Action selection | Works well with |
|---|---|---|---|
| Value-based (Q-learning) | Expected reward of each state-action pair | Pick the highest-value action | Discrete, small action spaces |
| Policy-based (PPO, REINFORCE) | A direct mapping from state to action probabilities | Sample from the learned distribution | Continuous or huge action spaces, LLM alignment |
| Actor-critic (hybrid) | Both a policy and a value function | Policy guided by value estimates | Most modern deep RL systems |
Optimization intuition
Both value-based and policy-based RL ultimately rely on gradient-based optimization to update their parameters toward better expected reward, just like supervised training. Use the gradient descent demo to build intuition for how iterative updates navigate toward better solutions — RL applies the same underlying mechanics to a much noisier, delayed reward signal.
📉 Gradient Descent Optimization
InteractiveWhat's next
Reinforcement learning becomes especially important in modern AI through RLHF — Reinforcement Learning from Human Feedback — the technique that takes a raw pre-trained language model and aligns it with human preferences using exactly the policy-based methods introduced here. Continue to the RLHF lesson to see how this works in practice.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.