Learn how to identify, inspect, and replace missing values in Pandas DataFrames using isna(), dropna(), and fillna().
What it is
In data analysis, "empty cells" usually refer to missing values. In Python's Pandas library, these are represented by the special floating-point value NaN (Not a Number) or the object None. These placeholders indicate that data was not recorded, was invalid, or is intentionally absent.
The mental model is that NaN is contagious: any arithmetic operation involving NaN results in NaN. Therefore, you must explicitly handle these values before performing calculations or visualizations. Related terms include nulls, missing data, and imputation (the process of filling gaps).
Why it matters
- Prevents Errors: Many statistical functions return
NaNif even one value is missing, breaking your analysis pipeline. - Data Integrity: Ignoring missing data can skew averages and correlations, leading to incorrect business or scientific conclusions.
- Visualization Quality: Charts may break or display misleading gaps if missing values are not handled consistently.
- Model Training: Most machine learning algorithms cannot process
NaNvalues directly; they require complete datasets.
Syntax or steps
Pandas provides three primary strategies for handling missing data:
- Detect: Use
df.isna()ordf.isnull()to create a boolean mask identifying missing entries. - Remove: Use
df.dropna()to delete rows or columns containing missing values. - Fill: Use
df.fillna(value)to replace missing values with a specific constant, mean, median, or forward/backward filled value.
Example
import pandas as pd
import numpy as np
# Create a DataFrame with some missing values
data = {
'Name': ['Alice', 'Bob', None, 'David'],
'Age': [25, np.nan, 30, 28],
'Score': [88, 92, np.nan, 76]
}
df = pd.DataFrame(data)
print("Original DataFrame:")
print(df)
# 1. Check for missing values
print("\nMissing Value Count per Column:")
print(df.isna().sum())
# 2. Fill numeric missing values with the column mean
df['Age'] = df['Age'].fillna(df['Age'].mean())
df['Score'] = df['Score'].fillna(df['Score'].median())
# 3. Fill text missing values with a placeholder string
df['Name'] = df['Name'].fillna('Unknown')
print("\nCleaned DataFrame:")
print(df)
Explanation: First, we import pandas and numpy because np.nan is the standard representation for missing floats. We create a sample dataset with mixed types. df.isna().sum() quickly reveals which columns have issues. We then use fillna() on specific columns. For numerical data like Age, using the mean() preserves the distribution better than zero. For Name, replacing None with 'Unknown' ensures the column remains string-typed without errors.
Common mistakes
- Using
inplace=Trueunnecessarily: Modern Pandas versions discourageinplaceoperations due to potential side effects and confusion. Prefer assignment:df = df.fillna(0). - Filling strings with numbers: Attempting to fill a text column with
0will cause type coercion issues or errors. Always match the fill value to the column's data type. - Ignoring
NonevsNaN: While both represent missing data,Noneis often used in object columns andNaNin float columns.isna()handles both correctly, but manual checks might miss one. - Dropping too much data: Using
dropna()on a large dataset with sparse missing values can result in losing significant information. Consider imputation instead.
When to use it
| Strategy | Best Used When... | Risk |
|---|---|---|
dropna() | Missing data is random and minimal (<5% of rows). | Bias if missingness correlates with target variable. |
fillna(mean/median) | Numerical data where central tendency is representative. | Reduces variance; distorts distribution tails. |
fillna(method='ffill') | Time-series data where previous values persist. | Carries forward errors or stale data. |
Practice
Guided Exercise: Create a DataFrame with a column 'Price' containing [10, np.nan, 20, np.nan]. Calculate the mean of non-missing values and fill the NaNs with this mean.
Challenge: How would you fill missing values in a categorical column (e.g., 'Color') with the most frequent color? Hint: Look up mode().
Quick check
Q: Why does df['col'].fillna(0) return a new Series instead of modifying df?
A: By default, Pandas methods return a copy of the data to prevent unintended side effects. You must assign the result back (df['col'] = ...) to update the original DataFrame.
Summary
Handling empty cells requires choosing between removal and imputation based on data volume and context. Always verify your strategy with isna().sum() to ensure no unexpected missing values remain after processing.