Back to Python Notes
Topic #279

Getting Started with Machine Learning

By the end of this lesson, you will understand the standard workflow for building a machine learning model in Python using scikit-learn, specifically focusing on splitting data to prevent overfitting.

What it is

Machine Learning (ML) is a subset of artificial intelligence where algorithms learn patterns from data rather than following explicit instructions. The core workflow involves preparing data, training a model, and evaluating its performance on unseen data. A critical component of this process is data splitting, which divides your dataset into separate groups: one for teaching the model (training set) and one for testing its generalization ability (test set). Related terms include features (input variables), labels (output targets), and overfitting (when a model memorizes training data but fails on new data).

Why it matters

  • Prevents Overfitting: Testing on data the model has never seen ensures the results reflect real-world performance.
  • Objective Evaluation: Provides a reliable metric (like accuracy or error rate) to compare different models.
  • Data Integrity: Ensures that information from the test set does not leak into the training phase, which would invalidate results.
  • Standard Practice: This workflow is the industry standard for reproducible ML experiments.

Syntax or steps

The most common way to split data in Python is using the train_test_split function from the sklearn.model_selection module. The basic pattern requires importing the function, passing in your features (X) and labels (y), and specifying the proportion of data reserved for testing.

Example

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

# 1. Load a sample dataset
data = load_iris()
X = data.data   # Features (sepal/petal measurements)
y = data.target # Labels (species type)

# 2. Split the data
# test_size=0.2 means 20% goes to testing, 80% to training
# random_state ensures reproducibility
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

print(f"Training samples: {len(X_train)}")
print(f"Testing samples: {len(X_test)}")

Part-by-part explanation:

  • load_iris(): Fetches a built-in dataset so you can run the code immediately without external files.
  • X and y: Separate inputs from outputs. ML models require distinct feature matrices and target vectors.
  • train_test_split(...): Randomly shuffles and splits the arrays. It returns four objects: training features, testing features, training labels, and testing labels.
  • random_state=42: Sets a seed for the random number generator. This guarantees that every time you run the script, the same rows are selected for training and testing, making debugging easier.

Common mistakes

  • Forgetting to shuffle: If your data is ordered by class (e.g., all cats first, then dogs), a simple slice might result in a training set with only cats. train_test_split shuffles by default, but manual slicing often does not.
  • Using the test set for tuning: Never adjust hyperparameters based on test set performance. Use a validation set or cross-validation instead.
  • Ignoring class imbalance: In datasets where one class is rare, a random split might leave the minority class out of the test set entirely. Use stratify=y in train_test_split to maintain class proportions.
  • Not setting random_state: Without a fixed seed, results vary between runs, making it impossible to reproduce bugs or improvements.

When to use it

Use train_test_split for quick prototyping and when you have sufficient data. For smaller datasets or rigorous evaluation, consider Cross-Validation.

MethodBest ForProsCons
train_test_split Large datasets, fast iteration Simple, fast, easy to implement High variance if split is unlucky; wastes some data
Cross-Validation Small/Medium datasets, final reporting Uses all data for training/testing; more robust estimate Computationally expensive; slower to run

Practice

Guided Exercise: Modify the example above to use test_size=0.3. Print the shape of X_train and X_test using .shape. Verify that the total number of rows equals the original dataset size.

Challenge: Add the argument stratify=y to the train_test_split call. Explain why this is important for classification problems like Iris.

Hint: Stratification ensures that the percentage of each species in the training and test sets matches the original dataset.

Quick check

Question: Why do we pass both X and y to train_test_split?

Answer: We must split features and labels simultaneously and identically so that each row in X_train corresponds correctly to its label in y_train. If they were split separately, the alignment between input and output would be lost.

Summary

Splitting data into training and testing sets is the foundational step in any machine learning project, ensuring that models are evaluated fairly on unseen data. Using train_test_split with a fixed random_state provides a reproducible and standard way to initiate this workflow in Python.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Getting Started with Machine Learning – FAQs

Quick answers about learning Getting Started with Machine Learning in Python.

This free note from Coding Now Tech Institute explains Getting Started with Machine Learning in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Getting Started with Machine Learning, is 100% free with no signup required.
With focused practice, most students grasp Getting Started with Machine Learning in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now