Back to Data Science Notes
Topic #74

Principal Component Analysis (PCA)

By the end of this lesson, you will understand how Principal Component Analysis (PCA) reduces data dimensions while preserving maximum variance, and you will be able to implement it using Python.

What it is

Principal Component Analysis (PCA) is a statistical technique used for dimensionality reduction. It transforms a set of correlated variables into a smaller number of uncorrelated variables called principal components. The first principal component captures the largest possible variance in the data, the second captures the next largest variance orthogonal to the first, and so on.

The mental model is projecting high-dimensional data onto lower-dimensional axes that best represent the spread of the data. Related terms include eigenvalues (variance explained by each component), eigenvectors (directions of the components), and singular value decomposition (the underlying linear algebra method).

Why it matters

  • Visualization: Reduces complex datasets to 2D or 3D for plotting.
  • Noise Reduction: Discards low-variance components which often represent noise.
  • Computational Efficiency: Fewer features mean faster training times for machine learning models.
  • Multicollinearity Handling: Creates orthogonal features, solving issues where input variables are highly correlated.

Syntax or steps

  1. Standardize Data: PCA is sensitive to scale; always center and scale features to zero mean and unit variance.
  2. Compute Covariance Matrix: Determine how features vary together.
  3. Calculate Eigenvectors/Eigenvalues: Find directions of maximum variance.
  4. Select Components: Choose top k eigenvectors based on desired variance retention.
  5. Project Data: Transform original data onto the new subspace.

Example

import numpy as np
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler

# Generate synthetic data: 100 samples, 5 features
np.random.seed(42)
data = np.random.rand(100, 5) * 10

# Step 1: Standardize the data
scaler = StandardScaler()
data_scaled = scaler.fit_transform(data)

# Step 2: Apply PCA to reduce to 2 components
pca = PCA(n_components=2)
principal_components = pca.fit_transform(data_scaled)

# Output results
print("Original shape:", data.shape)
print("Reduced shape:", principal_components.shape)
print("Explained Variance Ratio:", pca.explained_variance_ratio_)

This code standardizes random data, then projects it from 5 dimensions down to 2. The explained_variance_ratio_ tells you how much information was retained by the two new components.

Common mistakes

  • Forgetting to Scale: If one feature ranges from 0-1 and another from 0-1000, PCA will bias towards the larger range. Always use StandardScaler.
  • Interpreting Components as Original Features: Principal components are linear combinations of original features. They do not have direct physical meaning unless carefully analyzed via loadings.
  • Using PCA on Non-Linear Data: PCA assumes linear relationships. For curved manifolds, consider t-SNE or UMAP instead.
  • Ignoring Explained Variance: Blindly reducing dimensions without checking if significant variance is lost can degrade model performance.

When to use it

MethodBest ForLimitation
PCALinear dimensionality reduction, preprocessing for ML, visualization of global structure.Fails with non-linear patterns; requires scaling.
t-SNEVisualizing clusters in high-dimensional data (local structure).Computationally expensive; distances between clusters are not meaningful.
LDASupervised dimensionality reduction (maximizes class separation).Requires labeled data; limited to C-1 components.

Practice

Guided Exercise: Modify the example above to retain 95% of the variance automatically. Use PCA(n_components=0.95) and print the resulting number of components.

Challenge: Plot the first two principal components using matplotlib. Color points by their distance from the origin to visualize density.

Quick check

Question: Why must data be standardized before applying PCA?

Answer: PCA maximizes variance. Without standardization, features with larger scales dominate the variance calculation, leading to biased principal components that reflect measurement units rather than intrinsic data structure.

Summary

PCA is a powerful unsupervised learning technique for reducing dimensionality by identifying orthogonal axes of maximum variance. It is essential for preprocessing, visualization, and improving computational efficiency, provided the data is scaled and relationships are approximately linear.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Principal Component Analysis (PCA) – FAQs

Quick answers about learning Principal Component Analysis (PCA) in Data Science.

This free note from Coding Now Tech Institute explains Principal Component Analysis (PCA) in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including Principal Component Analysis (PCA), is 100% free with no signup required.
With focused practice, most students grasp Principal Component Analysis (PCA) in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now