Back to Data Science Notes
Topic #77

Forward & Backward Propagation

Understand how neural networks update weights by calculating gradients through forward and backward propagation.

What it is

Forward propagation computes the output of a neural network given an input. It passes data through layers, applying weights, biases, and activation functions until a prediction is made. Backward propagation calculates the gradient of the loss function with respect to each weight using the chain rule. These gradients indicate how much each weight contributed to the error, allowing the optimizer to adjust them in the opposite direction of the gradient to minimize loss.

Key terms include activation function (e.g., ReLU, Sigmoid), loss function (e.g., Mean Squared Error), and learning rate (step size for updates).

Why it matters

  • Automated Learning: Enables networks to learn complex patterns without manual feature engineering.
  • Efficiency: Computes gradients for millions of parameters simultaneously via matrix operations.
  • Scalability: Forms the backbone of deep learning frameworks like TensorFlow and PyTorch.
  • Optimization: Provides the mathematical basis for algorithms like Stochastic Gradient Descent (SGD).

Syntax or steps

  1. Forward Pass: Compute layer outputs $z = Wx + b$ and activations $a = \sigma(z)$.
  2. Loss Calculation: Compare final prediction $\hat{y}$ to true label $y$ using a loss function $L$.
  3. Backward Pass: Start from the output layer and compute partial derivatives $\frac{\partial L}{\partial W}$ moving backward.
  4. Weight Update: Adjust weights: $W_{new} = W_{old} - \eta \cdot \frac{\partial L}{\partial W}$.

Example

import numpy as np

# Simple 1-layer network: y_pred = sigmoid(W * x + b)
def sigmoid(x):
    return 1 / (1 + np.exp(-x))

def sigmoid_derivative(x):
    s = sigmoid(x)
    return s * (1 - s)

# Data
X = np.array([[0.5], [1.0]]) # Inputs
Y = np.array([[1], [0]])     # Targets

# Initialization
W = np.random.randn(1, 1)
b = np.zeros((1, 1))
lr = 0.1

for epoch in range(1000):
    # Forward Propagation
    z = np.dot(X, W) + b
    A = sigmoid(z)
    
    # Loss (MSE)
    loss = np.mean((A - Y)**2)
    
    # Backward Propagation
    dL_dA = 2 * (A - Y) / len(Y)       # Derivative of Loss w.r.t Output
    dA_dZ = sigmoid_derivative(z)      # Derivative of Activation w.r.t Z
    dZ_dW = X                          # Derivative of Z w.r.t Weight
    
    dL_dW = np.dot(dL_dA * dA_dZ).T @ dZ_dW # Chain Rule
    dL_db = np.sum(dL_dA * dA_dZ, axis=0)   # Bias derivative
    
    # Update Weights
    W -= lr * dL_dW
    b -= lr * dL_db

print(f"Final Prediction: {sigmoid(np.dot(X, W) + b)}")

This code manually implements one training step. The forward pass generates predictions. The backward pass uses the chain rule to propagate errors back to the weight W. Finally, the weight is updated to reduce future error.

Common mistakes

  • Vanishing Gradients: Using Sigmoid/Tanh in deep networks can cause gradients to shrink to zero. Fix: Use ReLU or Leaky ReLU.
  • Exploding Gradients: Gradients become too large, causing instability. Fix: Clip gradients or use Batch Normalization.
  • Incorrect Shape Broadcasting: Matrix dimensions mismatch during dot products. Fix: Ensure consistent shapes for inputs, weights, and biases.
  • Learning Rate Too High: Causes oscillation around minima. Fix: Reduce learning rate or use adaptive optimizers like Adam.

When to use it

MethodUse CasePros/Cons
Manual ImplementationLearning concepts, small modelsHigh control; slow, error-prone
Autograd (PyTorch/TF)Production, complex architecturesFast, accurate; less transparent

Use manual implementation for educational purposes to understand the math. Use automatic differentiation libraries for real-world applications to handle complexity and efficiency.

Practice

Guided Exercise: Modify the example above to use a ReLU activation function instead of Sigmoid. Note that ReLU's derivative is 1 if $x > 0$, else 0.

Challenge: Add a second hidden layer to the network. You will need to implement two sets of weights and biases and extend the backward pass to calculate gradients for both layers.

Quick check

Q: Why do we multiply by the derivative of the activation function during backward propagation?

A: Because of the chain rule. The loss depends on the activation, which depends on the weighted sum ($z$). To find how the loss changes with respect to weights, we must account for how sensitive the activation is to changes in $z$.

Summary

Forward propagation predicts outcomes, while backward propagation calculates how wrong those predictions were relative to each parameter. Together, they enable iterative optimization, allowing neural networks to learn from data by minimizing loss functions.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Forward & Backward Propagation – FAQs

Quick answers about learning Forward & Backward Propagation in Data Science.

This free note from Coding Now Tech Institute explains Forward & Backward Propagation in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Forward & Backward Propagation, is 100% free with no signup required.
With focused practice, most students grasp Forward & Backward Propagation in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now