By the end of this lesson, you will be able to construct, compile, and train a simple neural network using TensorFlow's Keras API to solve a basic regression problem.
What it is
TensorFlow is an open-source machine learning framework developed by Google. Keras is a high-level neural networks API, now integrated into TensorFlow as tf.keras. It allows developers to build deep learning models with minimal code. The mental model involves defining a "graph" or sequence of layers (the architecture), compiling it with an optimizer and loss function (the training strategy), and fitting it to data (the learning process).
Related terms include tensors (multi-dimensional arrays), epochs (one complete pass through the training dataset), and batch size (number of samples processed before updating weights).
Why it matters
- Rapid Prototyping: Keras provides a user-friendly interface that lets you test ideas quickly without managing low-level computational graphs.
- Scalability: Models built in Keras can easily scale from CPU to GPU or TPU clusters for faster training on large datasets.
- Industry Standard: Understanding
tf.kerasis essential for most modern ML engineering roles and research projects. - Modularity: You can swap out layers, optimizers, or loss functions independently, making experimentation systematic.
Syntax or steps
The standard workflow follows three main steps:
1. Build: Define the model architecture using Sequential or functional APIs.
2. Compile: Configure the learning process by specifying the optimizer, loss function, and metrics.
3. Fit: Train the model by passing input data (x) and target labels (y).
Example
import tensorflow as tf
from tensorflow import keras
import numpy as np
# 1. Prepare Data
# Simple linear relationship: y = 2x - 1
x_train = np.array([-1.0, 0.0, 1.0, 2.0, 3.0], dtype=float)
y_train = np.array([-3.0, -1.0, 1.0, 3.0, 5.0], dtype=float)
# 2. Build Model
model = keras.Sequential([
keras.layers.Dense(1, input_shape=[1]) # Single neuron, single feature
])
# 3. Compile Model
model.compile(optimizer='sgd', loss='mean_squared_error')
# 4. Train Model
model.fit(x_train, y_train, epochs=500, verbose=0)
# 5. Predict
prediction = model.predict(np.array([4.0]))
print(f"Prediction for x=4: {prediction[0][0]:.2f}")
# Expected output close to 7.0 (since 2*4 - 1 = 7)
This example creates a model with one layer containing one neuron. It learns the slope and intercept of a line. After 500 iterations, it should approximate the underlying mathematical rule.
Common mistakes
- Mismatched Input Shapes: Ensure your data dimensions match the
input_shapedefined in the first layer. For tabular data, this is often[num_features]. - Forgetting to Normalize: Neural networks perform poorly if features have vastly different scales. Always normalize or standardize inputs before training.
- Incorrect Loss Function: Using
categorical_crossentropyfor regression tasks will fail. Usemean_squared_errorfor continuous values. - Too Few Epochs: If the loss doesn't decrease, the model may not have trained long enough. Conversely, too many epochs can lead to overfitting.
When to use it
Keras is ideal for most standard deep learning tasks. Compare it with PyTorch below:
| Feature | TensorFlow/Keras | PyTorch |
|---|---|---|
| Learning Curve | Easier for beginners; high-level abstraction. | Steeper; requires more manual control. |
| Deployment | Strong support for production (TF Serving, Lite). | Growing support, but historically research-focused. |
| Debugging | Graph execution can be harder to debug initially. | Dynamic graph makes debugging intuitive. |
Use Keras when you need quick results and production readiness. Use PyTorch when you need fine-grained control over the training loop or are conducting novel research.
Practice
Guided Exercise: Modify the example above to predict y = 3x + 2. Change the y_train array accordingly and observe how the prediction changes for x=5.
Challenge: Add a second Dense layer with 10 neurons and activation 'relu' between the input and output layers. Does the model still learn the linear relationship? Why might adding depth help or hurt here?
Quick check
Q: What does the compile step do in Keras?
A: It configures the model for training by setting up the optimizer (how weights update), the loss function (what error to minimize), and metrics (what to track during training).
Summary
TensorFlow's Keras API simplifies deep learning by abstracting complex tensor operations into readable layer definitions. Mastering the Build-Compile-Fit cycle is the foundational skill for creating any neural network model.