Back to Data Science Notes
Topic #107

Introduction to MLOps

By the end of this lesson, you will understand how MLOps bridges the gap between experimental data science and reliable production systems through versioning, automated pipelines, and continuous monitoring.

What it is

MLOps (Machine Learning Operations) is a set of practices that combines Machine Learning (ML), DevOps, and Data Engineering. It aims to deploy and maintain ML models in production reliably and efficiently. The core mental model shifts from "building a model" to "managing a lifecycle." Key components include Data Versioning (tracking datasets like code), Experiment Tracking (logging metrics and parameters), Pipelines (automating training and deployment), and Monitoring (detecting drift or performance decay).

Why it matters

  • Reproducibility: Ensures that any result can be recreated exactly by tracking code, data, and environment versions.
  • Automation: Reduces manual errors by automating retraining and deployment when new data arrives.
  • Reliability: Continuous monitoring detects issues like data drift before they impact business outcomes.
  • Collaboration: Provides a shared framework for data scientists and engineers to work together seamlessly.

Syntax or steps

A minimal MLOps workflow involves three distinct stages: 1. Track: Log experiments using a tool like MLflow or Weights & Biases. 2. Build: Create a pipeline script that loads data, trains a model, and saves artifacts. 3. Deploy/Monitor: Serve the model via an API and log inference statistics.

Example

This example uses Python with scikit-learn and mlflow to demonstrate basic experiment tracking and model logging.

import mlflow
import mlflow.sklearn
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score

# 1. Load and split data
data = load_iris()
X_train, X_test, y_train, y_test = train_test_split(
    data.data, data.target, test_size=0.2, random_state=42
)

# 2. Start an MLflow run to track the experiment
with mlflow.start_run(run_name="iris_rf_experiment"):
    # Log parameters
    mlflow.log_param("n_estimators", 100)
    mlflow.log_param("max_depth", 5)

    # Train model
    clf = RandomForestClassifier(n_estimators=100, max_depth=5, random_state=42)
    clf.fit(X_train, y_train)

    # Evaluate
    predictions = clf.predict(X_test)
    acc = accuracy_score(y_test, predictions)
    
    # Log metrics
    mlflow.log_metric("accuracy", acc)

    # Log model artifact
    mlflow.sklearn.log_model(clf, "model")

print(f"Model trained with accuracy: {acc:.2f}")

Explanation:

  • mlflow.start_run() initializes a unique ID for this specific training session.
  • log_param records hyperparameters, allowing comparison between different runs.
  • log_metric stores evaluation scores for visualization and alerting.
  • log_model saves the serialized model object, making it retrievable for deployment later.

Common mistakes

  • Ignoring Data Versioning: Training on unversioned data makes debugging impossible if results change unexpectedly. Always link dataset hashes to model versions.
  • Lack of Monitoring: Deploying a model without checking for input distribution changes leads to silent failures. Implement alerts for accuracy drops or schema mismatches.
  • Manual Pipelines: Copy-pasting scripts for deployment introduces human error. Use orchestration tools (like Airflow or Kubeflow) to automate the flow from data ingestion to serving.

When to use it

ApproachBest ForLimitations
Ad-hoc Scripting Quick prototypes, one-off analyses, small teams. Not reproducible, hard to scale, no audit trail.
MLOps Pipeline Production systems, regulated industries, frequent updates. Higher initial setup cost, requires infrastructure knowledge.

Practice

Guided Exercise: Modify the example above to log two additional metrics: precision and recall. Hint: Import precision_score and recall_score from sklearn.metrics.

Challenge: Write a simple function that loads the logged model using mlflow.sklearn.load_model and predicts on a single sample from the Iris dataset. Verify the prediction matches the original model's output.

Quick check

Question: Why is logging parameters as important as logging metrics in MLOps?

Answer: Parameters define the configuration that produced the metric. Without them, you cannot reproduce the result or understand which settings led to high or low performance.

Summary

MLOps transforms machine learning from an art into an engineering discipline by enforcing reproducibility, automation, and observability. By integrating versioning and monitoring into your workflow, you ensure that models remain reliable and valuable throughout their lifecycle.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Introduction to MLOps – FAQs

Quick answers about learning Introduction to MLOps in Data Science.

This free note from Coding Now Tech Institute explains Introduction to MLOps in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Introduction to MLOps, is 100% free with no signup required.
With focused practice, most students grasp Introduction to MLOps in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now