AI & MLDeep
Intermediate

Reinforcement Learning Basics

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Live Simulator
Contains interactive viz

Learning by trial and error

Reinforcement learning (RL) is a fundamentally different paradigm from the supervised learning covered elsewhere in this track. There's no labeled dataset of correct answers. Instead, an agent interacts with an environment, takes actions, and receives rewards — numeric feedback signals indicating how good or bad an outcome was. The agent's entire goal is to learn a strategy, called a policy, that maximizes the total reward it accumulates over time.

This is the framework behind game-playing systems like AlphaGo, robotic control, and — critically for modern AI — RLHF, the technique used to align chatbots like ChatGPT and Claude with human preferences (covered in the next lesson).

The core loop: agent, environment, state, action, reward

RL formalizes the problem as a loop. At each timestep, the agent observes the current state $s$ of the environment, chooses an action $a$ according to its policy $\pi$, the environment transitions to a new state $s'$ and returns a reward $r$, and the cycle repeats. This is typically modeled as a Markov Decision Process (MDP): the assumption that the next state and reward depend only on the current state and action, not on the full history that led there.

The agent's objective is to maximize the expected cumulative reward, often discounted so that rewards received sooner count more than rewards received later: $G_t = r_t + \gamma r_{t+1} + \gamma^2 r_{t+2} + \dots$, where $\gamma \in [0,1]$ is the discount factor.

text

Note

A core tension in RL: should the agent exploit what it already knows works well, or explore untried actions that might work even better? Pure exploitation risks getting stuck at a mediocre strategy; pure exploration wastes time on actions already known to be bad. Common strategies to balance this include epsilon-greedy (mostly exploit, occasionally take a random action) and more sophisticated approaches like upper-confidence-bound methods that explore actions proportional to uncertainty about their value.

Value-based methods: learning how good states and actions are

Value-based methods learn a function that estimates expected future reward, then derive a policy from it implicitly — always pick the action with the highest estimated value. The Q-function, $Q(s, a)$, estimates the expected cumulative reward of taking action $a$ in state $s$ and then acting optimally afterward.

Q-learning learns this function iteratively using the Bellman equation as an update rule: after observing a transition $(s, a, r, s')$, update the estimate toward $r + \gamma \max_{a'} Q(s', a')$ — the immediate reward plus the discounted value of the best action available from the next state. Repeated over many episodes, these estimates converge toward accurate values, and the resulting policy (always take the highest-Q action) converges toward optimal behavior.

python

Policy-based methods: learning the strategy directly

Policy-based methods skip value estimation and instead directly parameterize and optimize the policy $\pi(a \mid s)$ — a function (often a neural network) that outputs a probability distribution over actions given a state. Training adjusts the policy's parameters to increase the probability of actions that led to high reward and decrease the probability of actions that led to low reward, typically via a gradient estimator like REINFORCE or more stable modern variants like PPO (Proximal Policy Optimization).

Policy-based methods handle continuous action spaces more naturally than Q-learning (which requires discretizing or maximizing over actions) and tend to be more stable in high-dimensional problems — which is exactly why PPO became the workhorse algorithm for RLHF, the topic of the next lesson.

Value-based vs. policy-based RL

ApproachLearnsAction selectionWorks well with
Value-based (Q-learning)Expected reward of each state-action pairPick the highest-value actionDiscrete, small action spaces
Policy-based (PPO, REINFORCE)A direct mapping from state to action probabilitiesSample from the learned distributionContinuous or huge action spaces, LLM alignment
Actor-critic (hybrid)Both a policy and a value functionPolicy guided by value estimatesMost modern deep RL systems

Optimization intuition

Both value-based and policy-based RL ultimately rely on gradient-based optimization to update their parameters toward better expected reward, just like supervised training. Use the gradient descent demo to build intuition for how iterative updates navigate toward better solutions — RL applies the same underlying mechanics to a much noisier, delayed reward signal.

📉 Gradient Descent Optimization

Interactive
Parameter (w)0.03.0 (Min)6.0Loss
Learning Rate (η)0.10
Steps taken:0
Current Weight (w):5.4000
Loss Value:6.2600
Gradient (dL/dw):+4.8000

What's next

Reinforcement learning becomes especially important in modern AI through RLHF — Reinforcement Learning from Human Feedback — the technique that takes a raw pre-trained language model and aligns it with human preferences using exactly the policy-based methods introduced here. Continue to the RLHF lesson to see how this works in practice.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.