By the end of this lesson, you will be able to quickly inspect and summarize Pandas DataFrames using built-in methods to understand data structure, types, and statistical distributions.
What it is
Analyzing a DataFrame involves examining its shape, column names, data types, and basic statistics without modifying the data. This process is often called Exploratory Data Analysis (EDA). Key concepts include:
head(): Displays the first few rows to verify data loading.info(): Provides a concise summary including index dtype, column dtypes, non-null values, and memory usage.describe(): Generates descriptive statistics for numerical columns (count, mean, std, min, max, quartiles).
Why it matters
- Data Integrity Check: Quickly identify if data loaded correctly or if there are unexpected null values.
- Type Verification: Ensure numeric columns are not stored as strings, which prevents calculation errors.
- Outlier Detection: Use
describe()to spot extreme values that might skew analysis. - Memory Optimization: Identify large datasets where downcasting types could save resources.
- Contextual Understanding: Gain immediate insight into the range and distribution of features before modeling.
Syntax or steps
The standard workflow for initial inspection follows these steps:
- Load the DataFrame into a variable (e.g.,
df). - Call
df.head(n)to view the first n rows (default 5). - Call
df.info()to check schema and missing values. - Call
df.describe()to get statistical summaries for numeric data.
Example
import pandas as pd
# Create a sample DataFrame
data = {
'Age': [25, 30, None, 45, 29],
'Salary': [50000, 60000, 75000, 80000, 55000],
'Department': ['HR', 'IT', 'IT', 'Finance', 'HR']
}
df = pd.DataFrame(data)
# 1. Inspect first few rows
print("Head:")
print(df.head())
# 2. Get structural info
print("\nInfo:")
df.info()
# 3. Get statistical summary
print("\nDescribe:")
print(df.describe())
Explanation:
df.head()prints the first five rows, allowing visual confirmation of the data structure.df.info()reveals that 'Age' has one missing value (None) and shows the total memory usage. It also confirms 'Department' is an object (string) type.df.describe()automatically calculates count, mean, standard deviation, min, max, and quartiles for numeric columns ('Age' and 'Salary'). Note that 'Department' is excluded because it is categorical.
Common mistakes
- Ignoring Nulls in Describe:
describe()excludes NaN values by default. If your count is lower than expected, checkdf.isnull().sum(). - Assuming All Columns Are Numeric: Calling
mean()on string columns raises an error. Always filter withselect_dtypes(include='number')if needed. - Overlooking Memory Usage: In large datasets,
info()shows memory consumption. Failing to optimize types (e.g., converting int64 to int32) can cause performance issues. - Misinterpreting Quartiles: The 25%, 50% (median), and 75% values in
describe()help detect skewness. A large gap between mean and median indicates outliers.
When to use it
Use these methods during the initial phase of any data project. Compare them with manual iteration:
| Method | Best For | Limitation |
|---|---|---|
df.head() | Quick visual verification of row content. | Does not show aggregate patterns. |
df.info() | Checking data types and missing values. | Does not provide statistical measures. |
df.describe() | Understanding distribution of numeric data. | Ignores categorical data by default. |
| Manual Looping | Custom complex checks. | Slow and verbose; avoid for basic EDA. |
Practice
Guided Exercise: Load a CSV file named sales.csv into a DataFrame. Print the number of rows and columns using shape, then display the first three rows.
Challenge: Using the same DataFrame, find out how many missing values exist in each column and print the mean salary only for employees in the 'IT' department.
Hint: Use df.isnull().sum() for missing values and boolean indexing like df[df['Dept'] == 'IT']['Salary'].mean() for conditional means.
Quick check
Question: Why does df.describe() not return statistics for string columns?
Answer: Because mathematical operations like mean and standard deviation are undefined for categorical text data. You must use include='all' to see counts and unique values for strings.
Summary
Inspecting DataFrames with head(), info(), and describe() provides a rapid, standardized way to understand data quality and distribution. These tools form the foundation of reliable data analysis by preventing assumptions about data types and completeness.