Back to Python Notes
Topic #262

Handling Empty Cells

Learn how to identify, inspect, and replace missing values in Pandas DataFrames using isna(), dropna(), and fillna().

What it is

In data analysis, "empty cells" usually refer to missing values. In Python's Pandas library, these are represented by the special floating-point value NaN (Not a Number) or the object None. These placeholders indicate that data was not recorded, was invalid, or is intentionally absent.

The mental model is that NaN is contagious: any arithmetic operation involving NaN results in NaN. Therefore, you must explicitly handle these values before performing calculations or visualizations. Related terms include nulls, missing data, and imputation (the process of filling gaps).

Why it matters

  • Prevents Errors: Many statistical functions return NaN if even one value is missing, breaking your analysis pipeline.
  • Data Integrity: Ignoring missing data can skew averages and correlations, leading to incorrect business or scientific conclusions.
  • Visualization Quality: Charts may break or display misleading gaps if missing values are not handled consistently.
  • Model Training: Most machine learning algorithms cannot process NaN values directly; they require complete datasets.

Syntax or steps

Pandas provides three primary strategies for handling missing data:

  1. Detect: Use df.isna() or df.isnull() to create a boolean mask identifying missing entries.
  2. Remove: Use df.dropna() to delete rows or columns containing missing values.
  3. Fill: Use df.fillna(value) to replace missing values with a specific constant, mean, median, or forward/backward filled value.

Example

import pandas as pd
import numpy as np

# Create a DataFrame with some missing values
data = {
    'Name': ['Alice', 'Bob', None, 'David'],
    'Age': [25, np.nan, 30, 28],
    'Score': [88, 92, np.nan, 76]
}
df = pd.DataFrame(data)

print("Original DataFrame:")
print(df)

# 1. Check for missing values
print("\nMissing Value Count per Column:")
print(df.isna().sum())

# 2. Fill numeric missing values with the column mean
df['Age'] = df['Age'].fillna(df['Age'].mean())
df['Score'] = df['Score'].fillna(df['Score'].median())

# 3. Fill text missing values with a placeholder string
df['Name'] = df['Name'].fillna('Unknown')

print("\nCleaned DataFrame:")
print(df)

Explanation: First, we import pandas and numpy because np.nan is the standard representation for missing floats. We create a sample dataset with mixed types. df.isna().sum() quickly reveals which columns have issues. We then use fillna() on specific columns. For numerical data like Age, using the mean() preserves the distribution better than zero. For Name, replacing None with 'Unknown' ensures the column remains string-typed without errors.

Common mistakes

  • Using inplace=True unnecessarily: Modern Pandas versions discourage inplace operations due to potential side effects and confusion. Prefer assignment: df = df.fillna(0).
  • Filling strings with numbers: Attempting to fill a text column with 0 will cause type coercion issues or errors. Always match the fill value to the column's data type.
  • Ignoring None vs NaN: While both represent missing data, None is often used in object columns and NaN in float columns. isna() handles both correctly, but manual checks might miss one.
  • Dropping too much data: Using dropna() on a large dataset with sparse missing values can result in losing significant information. Consider imputation instead.

When to use it

StrategyBest Used When...Risk
dropna()Missing data is random and minimal (<5% of rows).Bias if missingness correlates with target variable.
fillna(mean/median)Numerical data where central tendency is representative.Reduces variance; distorts distribution tails.
fillna(method='ffill')Time-series data where previous values persist.Carries forward errors or stale data.

Practice

Guided Exercise: Create a DataFrame with a column 'Price' containing [10, np.nan, 20, np.nan]. Calculate the mean of non-missing values and fill the NaNs with this mean.

Challenge: How would you fill missing values in a categorical column (e.g., 'Color') with the most frequent color? Hint: Look up mode().

Quick check

Q: Why does df['col'].fillna(0) return a new Series instead of modifying df?

A: By default, Pandas methods return a copy of the data to prevent unintended side effects. You must assign the result back (df['col'] = ...) to update the original DataFrame.

Summary

Handling empty cells requires choosing between removal and imputation based on data volume and context. Always verify your strategy with isna().sum() to ensure no unexpected missing values remain after processing.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Handling Empty Cells – FAQs

Quick answers about learning Handling Empty Cells in Python.

This free note from Coding Now Tech Institute explains Handling Empty Cells in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Handling Empty Cells, is 100% free with no signup required.
With focused practice, most students grasp Handling Empty Cells in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now