AI & MLDeep
Intermediate

RNNs and LSTMs

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Live Simulator
Contains interactive viz

Processing sequences one step at a time

Before transformers, Recurrent Neural Networks (RNNs) were the standard architecture for sequential data — text, speech, time series, anything where order matters and the input length varies. The core idea: instead of processing an entire input at once, an RNN processes a sequence one element at a time, maintaining a hidden state — a vector that acts as the network's memory of everything it has seen so far — and updating that hidden state at every step.

At each timestep $t$, the RNN takes the current input $x_t$ and the previous hidden state $h_{t-1}$, and computes a new hidden state: $h_t = \tanh(W_h h_{t-1} + W_x x_t + b)$. The same weights $W_h$ and $W_x$ are reused at every timestep, which is what lets an RNN handle sequences of any length with a fixed number of parameters.

python

The vanishing gradient problem

Training an RNN uses backpropagation through time: the network is effectively unrolled into one layer per timestep, and gradients flow backward through every one of those steps to update the shared weights. This is where plain RNNs run into a serious problem.

Because the same weight matrix is multiplied repeatedly at every timestep during backpropagation, gradients tend to either shrink toward zero (vanishing gradients) or grow explosively (exploding gradients) as they propagate backward across many steps — the same way repeatedly multiplying a number less than 1 by itself drives it toward zero, and repeatedly multiplying a number greater than 1 drives it toward infinity. In practice, vanishing gradients were the more common and damaging issue: information and error signal from early in a long sequence would effectively disappear by the time it reached the earliest timesteps, meaning a plain RNN struggled to learn dependencies spanning more than roughly 10-20 steps — a serious limitation for tasks like understanding a paragraph where an early sentence sets context for a much later one.

Note

Vanishing gradients affect any very deep network where signal has to pass through many sequential transformations — it's the same underlying issue that residual connections address in very deep CNNs and transformers. RNNs simply hit this wall earlier and harder, because a long sequence effectively makes the network as "deep" as it is long.

LSTMs: gates that control what to remember and forget

The Long Short-Term Memory (LSTM) network, introduced in 1997, was designed specifically to fix the vanishing gradient problem. Its key innovation: alongside the regular hidden state, an LSTM maintains a separate cell state — a kind of long-term memory conveyor belt that information can flow along with minimal transformation, protected from the repeated multiplication that causes gradients to vanish.

Controlling what goes onto and comes off that conveyor belt is the job of three learned gates, each a small neural network producing values between 0 and 1 (via a sigmoid) that act like dimmer switches:

Forget gate — decides what fraction of the existing cell state to discard.

Input gate — decides what new information from the current input to add to the cell state.

Output gate — decides what fraction of the (updated) cell state to expose as the hidden state used for output at this timestep.

Because the cell state update is largely additive rather than a repeated matrix multiplication, gradients can flow backward across many more timesteps without vanishing — allowing LSTMs to learn dependencies spanning hundreds of steps, far beyond what plain RNNs could manage.

text

GRUs: a simpler alternative

The Gated Recurrent Unit (GRU), introduced later, simplifies the LSTM design by merging the cell state and hidden state into one, and combining the forget and input gates into a single "update gate." GRUs have fewer parameters and are faster to train, and in many tasks perform comparably to LSTMs — making the choice between them often a matter of empirical testing on the specific task rather than a clear-cut rule.

RNN vs. LSTM vs. Transformer

PropertyVanilla RNNLSTMTransformer
Long-range dependenciesPoor — vanishing gradientsGood — gated cell stateExcellent — direct attention to any token
Parallelizable trainingNo — strictly sequentialNo — strictly sequentialYes — all tokens processed simultaneously
Parameters per layerLow~4x an RNN (four gate computations)Higher, but highly parallel-friendly
Typical use todayRare — mostly supersededNiche — small sequence tasks, streamingStandard for nearly all modern NLP

Why transformers replaced RNNs and LSTMs

Two compounding advantages explain why transformers displaced recurrent architectures for the vast majority of sequence modeling tasks, as covered in depth in the transformers lesson. First, attention directly connects any two tokens regardless of distance, without needing gradient signal to survive a long chain of sequential steps — solving the long-range dependency problem more thoroughly than even LSTM gating does. Second, and just as important operationally, transformers process an entire sequence in parallel rather than one step at a time, which RNNs and LSTMs fundamentally cannot do because each timestep depends on the previous hidden state. That parallelism is what let transformer training scale efficiently on modern GPU hardware — the same scaling that made today's large language models possible.

RNNs and LSTMs haven't disappeared entirely — they remain useful in some low-resource, streaming, or strictly causal real-time settings — but for nearly all mainstream NLP and sequence modeling work, transformers are now the default.

Layered network structure

Like any neural network, RNNs and LSTMs are trained end-to-end with backpropagation — the recurrence just adds a time dimension to how signal flows between layers. Use the visualization below to build general intuition for how information propagates through a layered network before diving into the sequential, gated version described above.

🧠 Neural Network Builder

Interactive
Hidden Layers (Max 5)
L14
L24
Total Layers: 4Neurons: 12Weights: 32Output Value: —

What's next

This closes the architectural arc from CNNs through RNNs/LSTMs to the attention-based approach that superseded both for sequence tasks. If you haven't already, read Attention and Transformers to see exactly how self-attention solves the long-range dependency and parallelism problems described here.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.