By the end of this lesson, you will understand how Convolutional Neural Networks (CNNs) extract spatial features from images using filters and pooling, and be able to implement a basic CNN layer in Python.
What it is
A Convolutional Neural Network (CNN) is a deep learning architecture specifically designed for processing grid-like data, such as images. Unlike standard neural networks that treat input pixels as independent vectors, CNNs preserve the spatial structure of the image. The core mental model involves three key operations: Convolution, where small matrices called filters or kernels slide over the image to detect local patterns like edges or textures; Activation, typically using ReLU to introduce non-linearity; and Pooling, which reduces the spatial dimensions while retaining important information. Related terms include feature maps, stride, padding, and receptive fields.
Why it matters
- Translation Invariance: CNNs can recognize objects regardless of their position in the image because filters scan the entire space.
- Parameter Efficiency: By sharing weights across the image, CNNs require far fewer parameters than fully connected layers for high-resolution inputs.
- Hierarchical Feature Learning: Early layers learn simple features (edges), while deeper layers combine them into complex structures (shapes, objects).
- State-of-the-Art Performance: They are the dominant architecture for computer vision tasks including classification, detection, and segmentation.
Syntax or steps
The fundamental operation is the convolution. A filter of size $k \times k$ slides over an input matrix. At each position, element-wise multiplication occurs between the filter and the underlying input patch, followed by summation to produce a single value in the output feature map. This process repeats across the width and height of the input. Pooling usually follows convolution, taking the maximum (Max Pooling) or average value within a smaller window to downsample the feature map.
Example
Below is a minimal implementation of a 2D convolution and max pooling step using NumPy to demonstrate the mechanics without heavy framework overhead.
import numpy as np
def conv2d(input_img, kernel):
# Get dimensions
h_in, w_in = input_img.shape
k_h, k_w = kernel.shape
# Calculate output dimensions (assuming valid padding, stride=1)
h_out = h_in - k_h + 1
w_out = w_in - k_w + 1
# Initialize output array
output = np.zeros((h_out, w_out))
# Slide kernel over input
for i in range(h_out):
for j in range(w_out):
# Extract region of interest
region = input_img[i:i+k_h, j:j+k_w]
# Element-wise multiply and sum
output[i, j] = np.sum(region * kernel)
return output
def max_pool2d(input_map, pool_size=2):
h, w = input_map.shape
h_out = h // pool_size
w_out = w // pool_size
output = np.zeros((h_out, w_out))
for i in range(h_out):
for j in range(w_out):
# Define region
region = input_map[i*pool_size:(i+1)*pool_size,
j*pool_size:(j+1)*pool_size]
# Take maximum
output[i, j] = np.max(region)
return output
# Create a simple 5x5 grayscale image (0-255)
image = np.array([
[10, 20, 30, 40, 50],
[60, 70, 80, 90, 100],
[110, 120, 130, 140, 150],
[160, 170, 180, 190, 200],
[210, 220, 230, 240, 250]
])
# Define a 3x3 edge detection kernel
kernel = np.array([
[-1, -1, -1],
[-1, 8, -1],
[-1, -1, -1]
])
# Apply convolution
feature_map = conv2d(image, kernel)
# Apply max pooling
pooled_output = max_pool2d(feature_map, pool_size=2)
print("Feature Map:\n", feature_map)
print("Pooled Output:\n", pooled_output)
In this example, conv2d applies the kernel to every possible position in the 5x5 image, resulting in a 3x3 feature map. The max_pool2d function then reduces this 3x3 map (conceptually, though strictly it requires even dimensions for perfect division, here it truncates) by taking the highest activation in 2x2 blocks, emphasizing strong edge responses.
Common mistakes
- Ignoring Channel Depth: Real images have RGB channels. Filters must match the depth of the input (e.g., a 3x3x3 filter for RGB). Beginners often forget to sum across all input channels.
- Incorrect Padding/Stride Calculation: Miscalculating output dimensions leads to shape mismatches in subsequent layers. Always verify: $Output = \frac{Input - Filter + 2 \times Padding}{Stride} + 1$.
- Forgetting Non-Linearity: Stacking linear convolutions without activation functions (like ReLU) collapses the network into a single linear transformation, losing expressive power.
- Over-Pooling: Using too large a pool size or too many pooling layers early on destroys spatial information needed for precise localization tasks.
When to use it
| Scenario | CNN | Fully Connected (MLP) |
|---|---|---|
| Data Type | Grid-like (Images, Audio spectrograms) | Tabular, unstructured vectors |
| Parameters | Low (Weight Sharing) | High (All-to-all connections) |
| Performance | Superior for spatial patterns | Poor for high-res images |
Use CNNs when spatial relationships matter. Use MLPs only for very small images or when spatial structure is irrelevant.
Practice
Guided Exercise: Modify the code above to use a "blur" kernel (all values 1/9) instead of an edge detector. Observe how the feature map changes.
Challenge: Implement a version of conv2d that handles multi-channel input (e.g., 3 channels) and outputs multiple feature maps (e.g., 10 filters). Hint: You will need nested loops for channels and filters.
Quick check
Question: Why do we use pooling layers after convolution?
Answer: To reduce spatial dimensions (downsampling), decrease computational cost, and provide some degree of translation invariance by summarizing local regions.
Summary
CNNs excel at image analysis by leveraging weight sharing through filters and dimensionality reduction via pooling. Understanding the manual mechanics of convolution helps demystify how these networks automatically learn hierarchical visual features.