Back to Python Notes
Topic #283

Train and Test Sets

Learn how to split your dataset into training and testing sets to accurately evaluate machine learning models without data leakage.

What it is

In supervised learning, a train/test split divides your dataset into two subsets: one used to teach the model (training set) and one held back to evaluate its performance (test set). The mental model is similar to studying for an exam using practice questions (training) and then taking the actual exam with unseen questions (testing). If you study the exact questions on the exam, your score is meaningless. Similarly, if you test on data the model has already seen, the evaluation is biased and overly optimistic. Related terms include validation set, cross-validation, and data leakage.

Why it matters

  • Prevents Overfitting: Ensures the model generalizes to new, unseen data rather than memorizing noise in the training set.
  • Realistic Evaluation: Provides an unbiased estimate of how the model will perform in production.
  • Model Selection: Allows comparison between different algorithms or hyperparameters based on fair metrics.
  • Data Integrity: Maintains the independence of training and testing samples, which is crucial for statistical validity.

Syntax or steps

The most common approach uses the train_test_split function from the sklearn.model_selection module. The standard workflow involves importing the function, defining your features (X) and target variable (y), and calling the split function. You can control the size of the test set using the test_size parameter and ensure reproducibility using random_state.

Example

from sklearn.model_selection import train_test_split
from sklearn.datasets import load_iris
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

# Load sample data
data = load_iris()
X = data.data
y = data.target

# Split the data: 80% training, 20% testing
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train the model
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)

# Evaluate the model
predictions = model.predict(X_test)
accuracy = accuracy_score(y_test, predictions)

print(f"Test Accuracy: {accuracy:.2f}")

Part-by-part explanation: First, we import necessary tools and load the Iris dataset. We define X as the feature matrix and y as the label vector. The train_test_split function shuffles the data and splits it; test_size=0.2 means 20% goes to testing, while random_state=42 ensures the same split every time this code runs. We then fit the logistic regression model only on X_train and y_train. Finally, we predict labels for X_test and calculate accuracy by comparing these predictions against the true y_test values.

Common mistakes

  • Testing on Training Data: Accidentally evaluating the model on X_train instead of X_test. Always verify which variables are passed to predict().
  • Ignoring Class Imbalance: A simple random split might result in a test set missing minority classes. Use the stratify=y parameter in train_test_split to maintain class proportions.
  • Scaling Before Splitting: Applying normalization or scaling to the entire dataset before splitting leaks information from the test set into the training process. Always split first, then scale X_train and transform X_test using the training parameters.
  • Forgetting Random State: Not setting random_state makes results non-reproducible, complicating debugging and comparisons.

When to use it

Use a simple train/test split when you have a large dataset and need a quick baseline evaluation. For smaller datasets or more robust validation, consider K-Fold Cross-Validation.

Method Best For Pros Cons
Train/Test Split Large datasets (>10k rows) Fast, simple implementation High variance in results depending on split
K-Fold CV Small/Medium datasets More reliable performance estimate Computationally expensive

Practice

Guided Exercise: Modify the example above to use stratify=y in the train_test_split call. Observe that the distribution of species in the test set matches the original dataset.

Challenge: Try changing test_size to 0.5 and rerun the code. Note how the accuracy changes. Why might a larger test set sometimes yield lower accuracy? (Hint: Consider the trade-off between training data volume and evaluation reliability.)

Quick check

Question: Why should you not apply StandardScaler to the entire dataset before calling train_test_split?
Answer: Because calculating mean and variance on the full dataset includes information from the test set, causing data leakage. You must fit the scaler on X_train only, then transform both sets.

Summary

Splitting data into training and testing sets is fundamental for honest model evaluation. It prevents overfitting by ensuring the model is judged on unseen data. Always remember to handle preprocessing after splitting to avoid data leakage.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Train and Test Sets – FAQs

Quick answers about learning Train and Test Sets in Python.

This free note from Coding Now Tech Institute explains Train and Test Sets in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Train and Test Sets, is 100% free with no signup required.
With focused practice, most students grasp Train and Test Sets in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now