Back to Data Science Notes
Topic #64

Introduction to Machine Learning

By the end of this lesson, you will understand the three main types of machine learning, the standard workflow for building models, and why splitting data into training and testing sets is critical for evaluating performance.

What it is

Machine Learning (ML) is a subset of artificial intelligence where computers learn patterns from data rather than following explicit instructions. The core mental model is "learning by example." Instead of writing rules like if temperature > 30 then predict_hot, you provide examples of temperatures and whether it was hot, letting the algorithm discover the threshold itself. The three primary types are:
  • Supervised Learning: Predicting a known output (label) from input features (e.g., predicting house prices).
  • Unsupervised Learning: Finding hidden structures in unlabeled data (e.g., customer segmentation).
  • Reinforcement Learning: Learning optimal actions through trial and error with rewards (e.g., game playing).

Why it matters

  • Automation: Handles complex tasks like fraud detection or recommendation engines at scale.
  • Pattern Discovery: Reveals insights humans might miss in large datasets.
  • Adaptability: Models can improve as new data becomes available.
  • Prediction: Enables forecasting future outcomes based on historical trends.

Syntax or steps

The standard ML workflow follows these steps:
  1. Data Collection: Gather relevant data.
  2. Preprocessing: Clean data, handle missing values, and encode categorical variables.
  3. Splitting: Divide data into Training and Testing sets.
  4. Model Selection: Choose an algorithm (e.g., Linear Regression, Decision Tree).
  5. Training: Fit the model to the training data.
  6. Evaluation: Test the model on unseen test data.

Example

Here is a minimal Python example using scikit-learn to demonstrate supervised learning with a train-test split.
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error
import numpy as np

# 1. Create synthetic data
X = np.array([[1], [2], [3], [4], [5]]) # Features
y = np.array([2, 4, 6, 8, 10])          # Labels (y = 2x)

# 2. Split data: 80% training, 20% testing
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# 3. Initialize and train the model
model = LinearRegression()
model.fit(X_train, y_train)

# 4. Make predictions on test set
predictions = model.predict(X_test)

# 5. Evaluate performance
mse = mean_squared_error(y_test, predictions)
print(f"Test MSE: {mse:.4f}")
print(f"Coefficient: {model.coef_[0]:.2f}")
Part-by-part explanation:
  • train_test_split: Randomly divides data so the model never sees the test data during training.
  • LinearRegression(): Creates an empty model object.
  • .fit(): Trains the model by finding the best line that fits X_train and y_train.
  • .predict(): Uses the learned pattern to guess values for X_test.
  • mean_squared_error: Calculates how far off the predictions were from actual values.

Common mistakes

  • Data Leakage: Preprocessing the entire dataset before splitting. Always split first, then preprocess each part separately.
  • Overfitting: A model performs perfectly on training data but poorly on test data. Fix by simplifying the model or adding more data.
  • Ignoring Scale: Some algorithms require normalized features. Use StandardScaler if needed.
  • Small Test Sets: If the test set is too small, evaluation metrics become unreliable. Ensure sufficient samples remain after splitting.

When to use it

Compare Supervised vs. Unsupervised learning based on your goal.
ScenarioTypeReason
Predicting sales revenueSupervisedYou have historical labeled data (past sales).
Grouping similar usersUnsupervisedNo predefined groups exist; find natural clusters.
Detecting anomaliesUnsupervisedNormal behavior defines the baseline; outliers are rare.

Practice

Guided Exercise: Modify the code above to change test_size to 0.5. Observe how the Mean Squared Error changes. Why might it increase? Challenge: Add a second feature to X (e.g., [1, 10], [2, 20]...) and retrain. Does the coefficient change? Hint: Check if the relationship remains linear.

Quick check

Question: Why do we evaluate a model on a test set instead of the training set? Answer: To measure how well the model generalizes to unseen data. Evaluating on training data leads to overestimation of performance because the model has already memorized those examples.

Summary

Machine Learning automates pattern recognition by training models on data. The train-test split is essential to validate that a model learns general rules rather than memorizing specific examples. Always start with clean data, simple models, and rigorous evaluation.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Introduction to Machine Learning – FAQs

Quick answers about learning Introduction to Machine Learning in Data Science.

This free note from Coding Now Tech Institute explains Introduction to Machine Learning in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Introduction to Machine Learning, is 100% free with no signup required.
With focused practice, most students grasp Introduction to Machine Learning in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now