Back to Python Notes
Topic #284

Categorical Data Preprocessing

Learn how to transform categorical variables into numerical formats suitable for machine learning models using one-hot encoding.

What it is

Categorical data consists of labels or names rather than numbers (e.g., "Red", "Blue", "Green"). Most machine learning algorithms require numerical input. One-hot encoding converts each category into a new binary column, where 1 indicates the presence of that category and 0 indicates its absence. This avoids implying an ordinal relationship between categories (e.g., assuming "Red" > "Blue"). Related terms include dummy variables, indicator features, and label encoding (which assigns integers but can introduce bias).

Why it matters

  • Algorithm Compatibility: Linear regression, neural networks, and support vector machines cannot process text strings directly.
  • Prevents False Ordering: Unlike label encoding, one-hot encoding treats all categories as distinct and equal, preventing models from interpreting arbitrary numerical order as significance.
  • Sparse Representation: It creates clear, independent features that help models isolate the impact of specific categories.
  • Standardization: It aligns with industry standards for feature engineering in tabular datasets.

Syntax or steps

The most common approach uses scikit-learn's OneHotEncoder. The basic workflow involves initializing the encoder, fitting it to training data to learn unique categories, and transforming both training and test sets.

  1. Import OneHotEncoder from sklearn.preprocessing.
  2. Create an instance, optionally setting sparse_output=False for dense arrays.
  3. Call fit_transform() on the training data.
  4. Call transform() on the test data to ensure consistency.

Example

from sklearn.preprocessing import OneHotEncoder
import numpy as np

# Sample categorical data: Color and Size
data = np.array([['Red', 'S'], ['Blue', 'M'], ['Red', 'L']])

# Initialize encoder
encoder = OneHotEncoder(sparse_output=False)

# Fit and transform
encoded_data = encoder.fit_transform(data)

print("Encoded Data:")
print(encoded_data)

print("\nFeature Names:")
print(encoder.get_feature_names_out())

Explanation: The input array contains two columns. The encoder identifies unique values ("Red", "Blue", "S", "M", "L"). It generates five new columns. For the first row ["Red", "S"], the output is [1, 0, 1, 0, 0], indicating "Red" is present (index 0), "Blue" is absent (index 1), "S" is present (index 2), etc. get_feature_names_out() provides readable column names like x0_Red.

Common mistakes

  • Fitting on Test Data: Never call fit_transform() on test data. Use transform() only. Fitting on test data leaks information about unseen categories, causing overfitting.
  • Ignoring High Cardinality: If a column has thousands of unique values (e.g., ZIP codes), one-hot encoding creates massive sparse matrices. Consider target encoding or dropping such features instead.
  • Forgetting Sparse Output: By default, newer versions may return sparse matrices. Set sparse_output=False if you need a standard NumPy array for visualization or certain libraries.
  • Dropping Categories Incorrectly: To avoid multicollinearity in linear models, use drop='first' in the encoder initialization. This removes one category per feature to prevent perfect correlation among dummy variables.

When to use it

MethodBest ForRisk
One-Hot EncodingNominal data (no order), low cardinality (<10 categories)Curse of dimensionality
Label EncodingOrdinal data (ordered categories), tree-based modelsImplies false numerical relationships
Target EncodingHigh cardinality nominal dataData leakage if not cross-validated

Use one-hot encoding when categories have no inherent order and the number of unique values is manageable. Use label encoding only for tree-based models (like Random Forests) which can handle integer splits without assuming linearity.

Practice

Guided Exercise: Encode a dataset with countries: ["USA", "Canada", "Mexico"]. Check the shape of the output.

Challenge: Modify the example to drop the first category for each feature to reduce multicollinearity. What is the new shape?

Hint: Add drop='first' to the OneHotEncoder constructor. The shape will decrease by the number of original features.

Quick check

Q: Why should you not fit the encoder on your entire dataset before splitting?

A: Because it introduces data leakage; the model learns about categories present in the test set during training, leading to overly optimistic performance estimates.

Summary

One-hot encoding transforms nominal categorical variables into binary vectors, enabling numerical computation while preserving category independence. Always fit the encoder on training data only and consider dropping categories or alternative methods for high-cardinality features to maintain model efficiency and accuracy.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Categorical Data Preprocessing – FAQs

Quick answers about learning Categorical Data Preprocessing in Python.

This free note from Coding Now Tech Institute explains Categorical Data Preprocessing in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Categorical Data Preprocessing, is 100% free with no signup required.
With focused practice, most students grasp Categorical Data Preprocessing in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now