By the end of this lesson, you will understand how MLOps bridges the gap between experimental data science and reliable production systems through versioning, automated pipelines, and continuous monitoring.
What it is
MLOps (Machine Learning Operations) is a set of practices that combines Machine Learning (ML), DevOps, and Data Engineering. It aims to deploy and maintain ML models in production reliably and efficiently. The core mental model shifts from "building a model" to "managing a lifecycle." Key components include Data Versioning (tracking datasets like code), Experiment Tracking (logging metrics and parameters), Pipelines (automating training and deployment), and Monitoring (detecting drift or performance decay).
Why it matters
- Reproducibility: Ensures that any result can be recreated exactly by tracking code, data, and environment versions.
- Automation: Reduces manual errors by automating retraining and deployment when new data arrives.
- Reliability: Continuous monitoring detects issues like data drift before they impact business outcomes.
- Collaboration: Provides a shared framework for data scientists and engineers to work together seamlessly.
Syntax or steps
A minimal MLOps workflow involves three distinct stages: 1. Track: Log experiments using a tool like MLflow or Weights & Biases. 2. Build: Create a pipeline script that loads data, trains a model, and saves artifacts. 3. Deploy/Monitor: Serve the model via an API and log inference statistics.
Example
This example uses Python with scikit-learn and mlflow to demonstrate basic experiment tracking and model logging.
import mlflow
import mlflow.sklearn
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score
# 1. Load and split data
data = load_iris()
X_train, X_test, y_train, y_test = train_test_split(
data.data, data.target, test_size=0.2, random_state=42
)
# 2. Start an MLflow run to track the experiment
with mlflow.start_run(run_name="iris_rf_experiment"):
# Log parameters
mlflow.log_param("n_estimators", 100)
mlflow.log_param("max_depth", 5)
# Train model
clf = RandomForestClassifier(n_estimators=100, max_depth=5, random_state=42)
clf.fit(X_train, y_train)
# Evaluate
predictions = clf.predict(X_test)
acc = accuracy_score(y_test, predictions)
# Log metrics
mlflow.log_metric("accuracy", acc)
# Log model artifact
mlflow.sklearn.log_model(clf, "model")
print(f"Model trained with accuracy: {acc:.2f}")
Explanation:
mlflow.start_run()initializes a unique ID for this specific training session.log_paramrecords hyperparameters, allowing comparison between different runs.log_metricstores evaluation scores for visualization and alerting.log_modelsaves the serialized model object, making it retrievable for deployment later.
Common mistakes
- Ignoring Data Versioning: Training on unversioned data makes debugging impossible if results change unexpectedly. Always link dataset hashes to model versions.
- Lack of Monitoring: Deploying a model without checking for input distribution changes leads to silent failures. Implement alerts for accuracy drops or schema mismatches.
- Manual Pipelines: Copy-pasting scripts for deployment introduces human error. Use orchestration tools (like Airflow or Kubeflow) to automate the flow from data ingestion to serving.
When to use it
| Approach | Best For | Limitations |
|---|---|---|
| Ad-hoc Scripting | Quick prototypes, one-off analyses, small teams. | Not reproducible, hard to scale, no audit trail. |
| MLOps Pipeline | Production systems, regulated industries, frequent updates. | Higher initial setup cost, requires infrastructure knowledge. |
Practice
Guided Exercise: Modify the example above to log two additional metrics: precision and recall. Hint: Import precision_score and recall_score from sklearn.metrics.
Challenge: Write a simple function that loads the logged model using mlflow.sklearn.load_model and predicts on a single sample from the Iris dataset. Verify the prediction matches the original model's output.
Quick check
Question: Why is logging parameters as important as logging metrics in MLOps?
Answer: Parameters define the configuration that produced the metric. Without them, you cannot reproduce the result or understand which settings led to high or low performance.
Summary
MLOps transforms machine learning from an art into an engineering discipline by enforcing reproducibility, automation, and observability. By integrating versioning and monitoring into your workflow, you ensure that models remain reliable and valuable throughout their lifecycle.