Understand how neural networks update weights by calculating gradients through forward and backward propagation.
What it is
Forward propagation computes the output of a neural network given an input. It passes data through layers, applying weights, biases, and activation functions until a prediction is made. Backward propagation calculates the gradient of the loss function with respect to each weight using the chain rule. These gradients indicate how much each weight contributed to the error, allowing the optimizer to adjust them in the opposite direction of the gradient to minimize loss.
Key terms include activation function (e.g., ReLU, Sigmoid), loss function (e.g., Mean Squared Error), and learning rate (step size for updates).
Why it matters
- Automated Learning: Enables networks to learn complex patterns without manual feature engineering.
- Efficiency: Computes gradients for millions of parameters simultaneously via matrix operations.
- Scalability: Forms the backbone of deep learning frameworks like TensorFlow and PyTorch.
- Optimization: Provides the mathematical basis for algorithms like Stochastic Gradient Descent (SGD).
Syntax or steps
- Forward Pass: Compute layer outputs $z = Wx + b$ and activations $a = \sigma(z)$.
- Loss Calculation: Compare final prediction $\hat{y}$ to true label $y$ using a loss function $L$.
- Backward Pass: Start from the output layer and compute partial derivatives $\frac{\partial L}{\partial W}$ moving backward.
- Weight Update: Adjust weights: $W_{new} = W_{old} - \eta \cdot \frac{\partial L}{\partial W}$.
Example
import numpy as np
# Simple 1-layer network: y_pred = sigmoid(W * x + b)
def sigmoid(x):
return 1 / (1 + np.exp(-x))
def sigmoid_derivative(x):
s = sigmoid(x)
return s * (1 - s)
# Data
X = np.array([[0.5], [1.0]]) # Inputs
Y = np.array([[1], [0]]) # Targets
# Initialization
W = np.random.randn(1, 1)
b = np.zeros((1, 1))
lr = 0.1
for epoch in range(1000):
# Forward Propagation
z = np.dot(X, W) + b
A = sigmoid(z)
# Loss (MSE)
loss = np.mean((A - Y)**2)
# Backward Propagation
dL_dA = 2 * (A - Y) / len(Y) # Derivative of Loss w.r.t Output
dA_dZ = sigmoid_derivative(z) # Derivative of Activation w.r.t Z
dZ_dW = X # Derivative of Z w.r.t Weight
dL_dW = np.dot(dL_dA * dA_dZ).T @ dZ_dW # Chain Rule
dL_db = np.sum(dL_dA * dA_dZ, axis=0) # Bias derivative
# Update Weights
W -= lr * dL_dW
b -= lr * dL_db
print(f"Final Prediction: {sigmoid(np.dot(X, W) + b)}")
This code manually implements one training step. The forward pass generates predictions. The backward pass uses the chain rule to propagate errors back to the weight W. Finally, the weight is updated to reduce future error.
Common mistakes
- Vanishing Gradients: Using Sigmoid/Tanh in deep networks can cause gradients to shrink to zero. Fix: Use ReLU or Leaky ReLU.
- Exploding Gradients: Gradients become too large, causing instability. Fix: Clip gradients or use Batch Normalization.
- Incorrect Shape Broadcasting: Matrix dimensions mismatch during dot products. Fix: Ensure consistent shapes for inputs, weights, and biases.
- Learning Rate Too High: Causes oscillation around minima. Fix: Reduce learning rate or use adaptive optimizers like Adam.
When to use it
| Method | Use Case | Pros/Cons |
|---|---|---|
| Manual Implementation | Learning concepts, small models | High control; slow, error-prone |
| Autograd (PyTorch/TF) | Production, complex architectures | Fast, accurate; less transparent |
Use manual implementation for educational purposes to understand the math. Use automatic differentiation libraries for real-world applications to handle complexity and efficiency.
Practice
Guided Exercise: Modify the example above to use a ReLU activation function instead of Sigmoid. Note that ReLU's derivative is 1 if $x > 0$, else 0.
Challenge: Add a second hidden layer to the network. You will need to implement two sets of weights and biases and extend the backward pass to calculate gradients for both layers.
Quick check
Q: Why do we multiply by the derivative of the activation function during backward propagation?
A: Because of the chain rule. The loss depends on the activation, which depends on the weighted sum ($z$). To find how the loss changes with respect to weights, we must account for how sensitive the activation is to changes in $z$.
Summary
Forward propagation predicts outcomes, while backward propagation calculates how wrong those predictions were relative to each parameter. Together, they enable iterative optimization, allowing neural networks to learn from data by minimizing loss functions.