Back to Data Science Notes
Topic #79

Convolutional Neural Networks (CNNs)

By the end of this lesson, you will understand how Convolutional Neural Networks (CNNs) extract spatial features from images using filters and pooling, and be able to implement a basic CNN layer in Python.

What it is

A Convolutional Neural Network (CNN) is a deep learning architecture specifically designed for processing grid-like data, such as images. Unlike standard neural networks that treat input pixels as independent vectors, CNNs preserve the spatial structure of the image. The core mental model involves three key operations: Convolution, where small matrices called filters or kernels slide over the image to detect local patterns like edges or textures; Activation, typically using ReLU to introduce non-linearity; and Pooling, which reduces the spatial dimensions while retaining important information. Related terms include feature maps, stride, padding, and receptive fields.

Why it matters

  • Translation Invariance: CNNs can recognize objects regardless of their position in the image because filters scan the entire space.
  • Parameter Efficiency: By sharing weights across the image, CNNs require far fewer parameters than fully connected layers for high-resolution inputs.
  • Hierarchical Feature Learning: Early layers learn simple features (edges), while deeper layers combine them into complex structures (shapes, objects).
  • State-of-the-Art Performance: They are the dominant architecture for computer vision tasks including classification, detection, and segmentation.

Syntax or steps

The fundamental operation is the convolution. A filter of size $k \times k$ slides over an input matrix. At each position, element-wise multiplication occurs between the filter and the underlying input patch, followed by summation to produce a single value in the output feature map. This process repeats across the width and height of the input. Pooling usually follows convolution, taking the maximum (Max Pooling) or average value within a smaller window to downsample the feature map.

Example

Below is a minimal implementation of a 2D convolution and max pooling step using NumPy to demonstrate the mechanics without heavy framework overhead.

import numpy as np

def conv2d(input_img, kernel):
    # Get dimensions
    h_in, w_in = input_img.shape
    k_h, k_w = kernel.shape
    
    # Calculate output dimensions (assuming valid padding, stride=1)
    h_out = h_in - k_h + 1
    w_out = w_in - k_w + 1
    
    # Initialize output array
    output = np.zeros((h_out, w_out))
    
    # Slide kernel over input
    for i in range(h_out):
        for j in range(w_out):
            # Extract region of interest
            region = input_img[i:i+k_h, j:j+k_w]
            # Element-wise multiply and sum
            output[i, j] = np.sum(region * kernel)
            
    return output

def max_pool2d(input_map, pool_size=2):
    h, w = input_map.shape
    h_out = h // pool_size
    w_out = w // pool_size
    output = np.zeros((h_out, w_out))
    
    for i in range(h_out):
        for j in range(w_out):
            # Define region
            region = input_map[i*pool_size:(i+1)*pool_size, 
                               j*pool_size:(j+1)*pool_size]
            # Take maximum
            output[i, j] = np.max(region)
            
    return output

# Create a simple 5x5 grayscale image (0-255)
image = np.array([
    [10, 20, 30, 40, 50],
    [60, 70, 80, 90, 100],
    [110, 120, 130, 140, 150],
    [160, 170, 180, 190, 200],
    [210, 220, 230, 240, 250]
])

# Define a 3x3 edge detection kernel
kernel = np.array([
    [-1, -1, -1],
    [-1,  8, -1],
    [-1, -1, -1]
])

# Apply convolution
feature_map = conv2d(image, kernel)

# Apply max pooling
pooled_output = max_pool2d(feature_map, pool_size=2)

print("Feature Map:\n", feature_map)
print("Pooled Output:\n", pooled_output)

In this example, conv2d applies the kernel to every possible position in the 5x5 image, resulting in a 3x3 feature map. The max_pool2d function then reduces this 3x3 map (conceptually, though strictly it requires even dimensions for perfect division, here it truncates) by taking the highest activation in 2x2 blocks, emphasizing strong edge responses.

Common mistakes

  • Ignoring Channel Depth: Real images have RGB channels. Filters must match the depth of the input (e.g., a 3x3x3 filter for RGB). Beginners often forget to sum across all input channels.
  • Incorrect Padding/Stride Calculation: Miscalculating output dimensions leads to shape mismatches in subsequent layers. Always verify: $Output = \frac{Input - Filter + 2 \times Padding}{Stride} + 1$.
  • Forgetting Non-Linearity: Stacking linear convolutions without activation functions (like ReLU) collapses the network into a single linear transformation, losing expressive power.
  • Over-Pooling: Using too large a pool size or too many pooling layers early on destroys spatial information needed for precise localization tasks.

When to use it

ScenarioCNNFully Connected (MLP)
Data TypeGrid-like (Images, Audio spectrograms)Tabular, unstructured vectors
ParametersLow (Weight Sharing)High (All-to-all connections)
PerformanceSuperior for spatial patternsPoor for high-res images

Use CNNs when spatial relationships matter. Use MLPs only for very small images or when spatial structure is irrelevant.

Practice

Guided Exercise: Modify the code above to use a "blur" kernel (all values 1/9) instead of an edge detector. Observe how the feature map changes.

Challenge: Implement a version of conv2d that handles multi-channel input (e.g., 3 channels) and outputs multiple feature maps (e.g., 10 filters). Hint: You will need nested loops for channels and filters.

Quick check

Question: Why do we use pooling layers after convolution?

Answer: To reduce spatial dimensions (downsampling), decrease computational cost, and provide some degree of translation invariance by summarizing local regions.

Summary

CNNs excel at image analysis by leveraging weight sharing through filters and dimensionality reduction via pooling. Understanding the manual mechanics of convolution helps demystify how these networks automatically learn hierarchical visual features.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Convolutional Neural Networks (CNNs) – FAQs

Quick answers about learning Convolutional Neural Networks (CNNs) in Data Science.

This free note from Coding Now Tech Institute explains Convolutional Neural Networks (CNNs) in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Convolutional Neural Networks (CNNs), is 100% free with no signup required.
With focused practice, most students grasp Convolutional Neural Networks (CNNs) in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now