RNNs and LSTMs
Processing sequences one step at a time
Before transformers, Recurrent Neural Networks (RNNs) were the standard architecture for sequential data — text, speech, time series, anything where order matters and the input length varies. The core idea: instead of processing an entire input at once, an RNN processes a sequence one element at a time, maintaining a hidden state — a vector that acts as the network's memory of everything it has seen so far — and updating that hidden state at every step.
At each timestep $t$, the RNN takes the current input $x_t$ and the previous hidden state $h_{t-1}$, and computes a new hidden state: $h_t = \tanh(W_h h_{t-1} + W_x x_t + b)$. The same weights $W_h$ and $W_x$ are reused at every timestep, which is what lets an RNN handle sequences of any length with a fixed number of parameters.
The vanishing gradient problem
Training an RNN uses backpropagation through time: the network is effectively unrolled into one layer per timestep, and gradients flow backward through every one of those steps to update the shared weights. This is where plain RNNs run into a serious problem.
Because the same weight matrix is multiplied repeatedly at every timestep during backpropagation, gradients tend to either shrink toward zero (vanishing gradients) or grow explosively (exploding gradients) as they propagate backward across many steps — the same way repeatedly multiplying a number less than 1 by itself drives it toward zero, and repeatedly multiplying a number greater than 1 drives it toward infinity. In practice, vanishing gradients were the more common and damaging issue: information and error signal from early in a long sequence would effectively disappear by the time it reached the earliest timesteps, meaning a plain RNN struggled to learn dependencies spanning more than roughly 10-20 steps — a serious limitation for tasks like understanding a paragraph where an early sentence sets context for a much later one.
Note
LSTMs: gates that control what to remember and forget
The Long Short-Term Memory (LSTM) network, introduced in 1997, was designed specifically to fix the vanishing gradient problem. Its key innovation: alongside the regular hidden state, an LSTM maintains a separate cell state — a kind of long-term memory conveyor belt that information can flow along with minimal transformation, protected from the repeated multiplication that causes gradients to vanish.
Controlling what goes onto and comes off that conveyor belt is the job of three learned gates, each a small neural network producing values between 0 and 1 (via a sigmoid) that act like dimmer switches:
Forget gate — decides what fraction of the existing cell state to discard.
Input gate — decides what new information from the current input to add to the cell state.
Output gate — decides what fraction of the (updated) cell state to expose as the hidden state used for output at this timestep.
Because the cell state update is largely additive rather than a repeated matrix multiplication, gradients can flow backward across many more timesteps without vanishing — allowing LSTMs to learn dependencies spanning hundreds of steps, far beyond what plain RNNs could manage.
GRUs: a simpler alternative
The Gated Recurrent Unit (GRU), introduced later, simplifies the LSTM design by merging the cell state and hidden state into one, and combining the forget and input gates into a single "update gate." GRUs have fewer parameters and are faster to train, and in many tasks perform comparably to LSTMs — making the choice between them often a matter of empirical testing on the specific task rather than a clear-cut rule.
RNN vs. LSTM vs. Transformer
| Property | Vanilla RNN | LSTM | Transformer |
|---|---|---|---|
| Long-range dependencies | Poor — vanishing gradients | Good — gated cell state | Excellent — direct attention to any token |
| Parallelizable training | No — strictly sequential | No — strictly sequential | Yes — all tokens processed simultaneously |
| Parameters per layer | Low | ~4x an RNN (four gate computations) | Higher, but highly parallel-friendly |
| Typical use today | Rare — mostly superseded | Niche — small sequence tasks, streaming | Standard for nearly all modern NLP |
Why transformers replaced RNNs and LSTMs
Two compounding advantages explain why transformers displaced recurrent architectures for the vast majority of sequence modeling tasks, as covered in depth in the transformers lesson. First, attention directly connects any two tokens regardless of distance, without needing gradient signal to survive a long chain of sequential steps — solving the long-range dependency problem more thoroughly than even LSTM gating does. Second, and just as important operationally, transformers process an entire sequence in parallel rather than one step at a time, which RNNs and LSTMs fundamentally cannot do because each timestep depends on the previous hidden state. That parallelism is what let transformer training scale efficiently on modern GPU hardware — the same scaling that made today's large language models possible.
RNNs and LSTMs haven't disappeared entirely — they remain useful in some low-resource, streaming, or strictly causal real-time settings — but for nearly all mainstream NLP and sequence modeling work, transformers are now the default.
Layered network structure
Like any neural network, RNNs and LSTMs are trained end-to-end with backpropagation — the recurrence just adds a time dimension to how signal flows between layers. Use the visualization below to build general intuition for how information propagates through a layered network before diving into the sequential, gated version described above.
🧠 Neural Network Builder
InteractiveWhat's next
This closes the architectural arc from CNNs through RNNs/LSTMs to the attention-based approach that superseded both for sequence tasks. If you haven't already, read Attention and Transformers to see exactly how self-attention solves the long-range dependency and parallelism problems described here.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.