Convolutional Neural Networks (CNNs)
Why not just use a regular neural network on images?
A standard fully connected neural network treats every input as a flat vector of independent numbers. Feed it an image and it has to flatten pixels into a giant list — a modest 224x224 color image already becomes over 150,000 input values. A fully connected first layer with even a few hundred neurons would need tens of millions of weights, before the network has learned anything useful, and critically, it would treat a pixel in the top-left corner as having no special relationship to its neighbor just one pixel away.
Convolutional Neural Networks (CNNs) were built specifically to exploit the structure that images actually have: nearby pixels are related, and useful visual patterns (edges, textures, shapes) tend to look the same no matter where in the image they appear. CNNs bake both of these assumptions directly into the architecture, which is why they dominated computer vision for over a decade and remain heavily used today, often alongside or inside transformer-based vision models.
The convolution operation
The core building block is the convolution: a small matrix of learnable weights, called a filter or kernel (commonly 3x3 or 5x5), slides across the input image. At each position, it computes a weighted sum of the pixel values it currently covers, producing a single output value. Sliding the same filter across the entire image produces a 2D grid of outputs called a feature map.
Because the same filter is reused at every position (called weight sharing), a CNN can detect a pattern — say, a vertical edge — no matter where it appears in the image, using a tiny number of parameters compared to a fully connected layer. A convolutional layer typically applies many different filters in parallel, each learning to detect a different pattern, producing a stack of feature maps as output.
Pooling: shrinking while keeping what matters
After a convolutional layer, CNNs typically apply a pooling layer to reduce the spatial size of the feature maps. Max pooling, the most common variant, slides a small window (e.g. 2x2) across the feature map and keeps only the maximum value in each window, discarding the rest. This reduces the amount of computation needed downstream, makes the network somewhat robust to small translations of the input (a feature detected slightly to the left still survives pooling), and progressively summarizes fine-grained detail into coarser, more abstract representations as the network gets deeper.
Note
Landmark CNN architectures
LeNet-5 (1998) was one of the earliest practical CNNs, used for handwritten digit recognition. AlexNet (2012) is widely credited with kicking off the deep learning boom in computer vision — it dramatically outperformed non-neural approaches on the ImageNet competition, demonstrating that deep CNNs trained on GPUs with large datasets could achieve breakthrough accuracy. VGGNet showed that simply stacking more small (3x3) convolutional layers improved performance. ResNet introduced residual (skip) connections, allowing networks with over a hundred layers to be trained successfully by giving gradients a direct path around each block — the same residual connection idea later became a core component of the transformer architecture as well.
Why CNNs suit images specifically
| Property | Fully connected network | CNN |
|---|---|---|
| Parameters for a 224x224 image | Tens of millions in the first layer alone | A few thousand per filter, reused everywhere |
| Translation invariance | None — position matters completely | Built in, via weight sharing |
| Exploits local pixel structure | No — treats all inputs as independent | Yes — filters see local neighborhoods |
| Scales to high-resolution images | Poorly | Well, with pooling to control growth |
Seeing the network structure
A CNN is still fundamentally a layered neural network — convolution and pooling are specialized layer types stacked in sequence, trained end-to-end with backpropagation just like the network in the visualization below. Use it to build intuition for how information flows and transforms through successive layers.
🧠 Neural Network Builder
InteractiveWhat's next
CNNs excel at spatial data like images; RNNs and LSTMs were the corresponding workhorse for sequential data like text and time series before transformers took over both domains. Continue to RNNs and LSTMs to see that architecture and why transformers eventually replaced it.
I build these systems professionally.
Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.