By the end of this lesson, you will understand how artificial neurons process data through weighted sums and activation functions to form a basic neural network architecture.
What it is
A neural network is a computational model inspired by biological brains. It consists of layers of interconnected nodes called neurons. Each neuron receives inputs, multiplies them by weights, adds a bias, and passes the result through an activation function. This structure allows the network to learn complex patterns in data.
The mental model is a pipeline: Input Layer → Hidden Layers (processing) → Output Layer (prediction). Key terms include weights (importance of input), bias (threshold adjustment), and activation (non-linear decision boundary).
Why it matters
- Pattern Recognition: Excels at identifying features in images, audio, and text that traditional algorithms miss.
- Non-Linearity: Activation functions allow networks to model complex, non-linear relationships between variables.
- Adaptability: Networks can be trained on new data without rewriting the underlying logic.
- Versatility: Used in everything from spam filtering to medical diagnosis and autonomous driving.
Syntax or steps
To build a minimal neural network layer, follow these steps:
- Define input vector $x$.
- Initialize weight matrix $W$ and bias vector $b$.
- Compute linear combination: $z = W \cdot x + b$.
- Apply activation function: $a = \sigma(z)$.
Example
import numpy as np
# Define sigmoid activation function
def sigmoid(x):
return 1 / (1 + np.exp(-x))
# Inputs: 3 samples, 4 features each
X = np.array([
[0.5, 0.8, 0.2, 0.9],
[0.1, 0.3, 0.7, 0.4],
[0.6, 0.2, 0.5, 0.8]
])
# Weights: 4 inputs connected to 3 hidden neurons
W = np.random.rand(4, 3) * 0.1
# Bias: one per hidden neuron
b = np.zeros((1, 3))
# Forward pass
linear_output = np.dot(X, W) + b
activated_output = sigmoid(linear_output)
print("Input Shape:", X.shape)
print("Output Shape:", activated_output.shape)
print("Sample Output:\n", activated_output[0])
This code creates a single hidden layer. The input matrix $X$ is multiplied by weights $W$, added to bias $b$, and passed through the sigmoid function. The output represents the "firing rate" of each neuron.
Common mistakes
- Forgetting Non-Linearity: Using only linear activations collapses deep networks into simple linear models. Always use functions like Sigmoid, ReLU, or Tanh.
- Incorrect Weight Initialization: Initializing all weights to zero causes symmetry problems where neurons learn identical features. Use small random values instead.
- Mismatched Dimensions: The number of columns in the weight matrix must match the number of rows in the input matrix for dot product operations.
- Ignoring Normalization: Raw inputs with large ranges can cause gradient explosion. Normalize data to mean 0 and variance 1 before training.
When to use it
| Scenario | Neural Network | Traditional ML (e.g., Linear Regression) |
|---|---|---|
| Data Volume | Large datasets (thousands+) | Small to medium datasets |
| Relationship | Complex, non-linear | Simple, linear |
| Interpretability | Low ("Black Box") | High |
| Compute Cost | High | Low |
Use neural networks when accuracy on complex tasks outweighs interpretability and computational cost.
Practice
Guided Exercise: Modify the example above to use the ReLU activation function ($max(0, x)$) instead of Sigmoid. Observe how negative values become zero.
Challenge: Add a second hidden layer. Create a new weight matrix connecting the first layer's output to a third layer, then apply another activation. Hint: Chain the forward pass calculations.
Quick check
Question: Why is the bias term necessary in a neuron?
Answer: The bias allows the activation function to shift left or right, enabling the neuron to fire even if all inputs are zero, thus fitting the data better.
Summary
Neural networks transform inputs through weighted sums and non-linear activations to learn complex patterns. Understanding the flow from input to output via weights and biases is the foundation for building deeper architectures.