AI & MLDeep
Intermediate

Convolutional Neural Networks (CNNs)

11 min read

Learn
Deep Reading
Estimated 11 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Live Simulator
Contains interactive viz

Why not just use a regular neural network on images?

A standard fully connected neural network treats every input as a flat vector of independent numbers. Feed it an image and it has to flatten pixels into a giant list — a modest 224x224 color image already becomes over 150,000 input values. A fully connected first layer with even a few hundred neurons would need tens of millions of weights, before the network has learned anything useful, and critically, it would treat a pixel in the top-left corner as having no special relationship to its neighbor just one pixel away.

Convolutional Neural Networks (CNNs) were built specifically to exploit the structure that images actually have: nearby pixels are related, and useful visual patterns (edges, textures, shapes) tend to look the same no matter where in the image they appear. CNNs bake both of these assumptions directly into the architecture, which is why they dominated computer vision for over a decade and remain heavily used today, often alongside or inside transformer-based vision models.

The convolution operation

The core building block is the convolution: a small matrix of learnable weights, called a filter or kernel (commonly 3x3 or 5x5), slides across the input image. At each position, it computes a weighted sum of the pixel values it currently covers, producing a single output value. Sliding the same filter across the entire image produces a 2D grid of outputs called a feature map.

Because the same filter is reused at every position (called weight sharing), a CNN can detect a pattern — say, a vertical edge — no matter where it appears in the image, using a tiny number of parameters compared to a fully connected layer. A convolutional layer typically applies many different filters in parallel, each learning to detect a different pattern, producing a stack of feature maps as output.

python

Pooling: shrinking while keeping what matters

After a convolutional layer, CNNs typically apply a pooling layer to reduce the spatial size of the feature maps. Max pooling, the most common variant, slides a small window (e.g. 2x2) across the feature map and keeps only the maximum value in each window, discarding the rest. This reduces the amount of computation needed downstream, makes the network somewhat robust to small translations of the input (a feature detected slightly to the left still survives pooling), and progressively summarizes fine-grained detail into coarser, more abstract representations as the network gets deeper.

Note

Nobody hand-designs what each filter should detect — the filters are learned end-to-end via backpropagation, just like any other weight. But a consistent pattern emerges: early convolutional layers learn to detect simple, low-level patterns like edges and color gradients. Middle layers combine those into textures and simple shapes. Later layers combine those into complex, semantically meaningful parts — eyes, wheels, wings — that are directly useful for classification. This hierarchical structure is a big part of why CNNs generalize so well to new images.
text

Landmark CNN architectures

LeNet-5 (1998) was one of the earliest practical CNNs, used for handwritten digit recognition. AlexNet (2012) is widely credited with kicking off the deep learning boom in computer vision — it dramatically outperformed non-neural approaches on the ImageNet competition, demonstrating that deep CNNs trained on GPUs with large datasets could achieve breakthrough accuracy. VGGNet showed that simply stacking more small (3x3) convolutional layers improved performance. ResNet introduced residual (skip) connections, allowing networks with over a hundred layers to be trained successfully by giving gradients a direct path around each block — the same residual connection idea later became a core component of the transformer architecture as well.

Why CNNs suit images specifically

PropertyFully connected networkCNN
Parameters for a 224x224 imageTens of millions in the first layer aloneA few thousand per filter, reused everywhere
Translation invarianceNone — position matters completelyBuilt in, via weight sharing
Exploits local pixel structureNo — treats all inputs as independentYes — filters see local neighborhoods
Scales to high-resolution imagesPoorlyWell, with pooling to control growth

Seeing the network structure

A CNN is still fundamentally a layered neural network — convolution and pooling are specialized layer types stacked in sequence, trained end-to-end with backpropagation just like the network in the visualization below. Use it to build intuition for how information flows and transforms through successive layers.

🧠 Neural Network Builder

Interactive
Hidden Layers (Max 5)
L14
L24
Total Layers: 4Neurons: 12Weights: 32Output Value: —

What's next

CNNs excel at spatial data like images; RNNs and LSTMs were the corresponding workhorse for sequential data like text and time series before transformers took over both domains. Continue to RNNs and LSTMs to see that architecture and why transformers eventually replaced it.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.