Back to Python Notes
Topic #260

Analyzing DataFrames

By the end of this lesson, you will be able to quickly inspect and summarize Pandas DataFrames using built-in methods to understand data structure, types, and statistical distributions.

What it is

Analyzing a DataFrame involves examining its shape, column names, data types, and basic statistics without modifying the data. This process is often called Exploratory Data Analysis (EDA). Key concepts include:

  • head(): Displays the first few rows to verify data loading.
  • info(): Provides a concise summary including index dtype, column dtypes, non-null values, and memory usage.
  • describe(): Generates descriptive statistics for numerical columns (count, mean, std, min, max, quartiles).

Why it matters

  • Data Integrity Check: Quickly identify if data loaded correctly or if there are unexpected null values.
  • Type Verification: Ensure numeric columns are not stored as strings, which prevents calculation errors.
  • Outlier Detection: Use describe() to spot extreme values that might skew analysis.
  • Memory Optimization: Identify large datasets where downcasting types could save resources.
  • Contextual Understanding: Gain immediate insight into the range and distribution of features before modeling.

Syntax or steps

The standard workflow for initial inspection follows these steps:

  1. Load the DataFrame into a variable (e.g., df).
  2. Call df.head(n) to view the first n rows (default 5).
  3. Call df.info() to check schema and missing values.
  4. Call df.describe() to get statistical summaries for numeric data.

Example

import pandas as pd

# Create a sample DataFrame
data = {
    'Age': [25, 30, None, 45, 29],
    'Salary': [50000, 60000, 75000, 80000, 55000],
    'Department': ['HR', 'IT', 'IT', 'Finance', 'HR']
}
df = pd.DataFrame(data)

# 1. Inspect first few rows
print("Head:")
print(df.head())

# 2. Get structural info
print("\nInfo:")
df.info()

# 3. Get statistical summary
print("\nDescribe:")
print(df.describe())

Explanation:

  • df.head() prints the first five rows, allowing visual confirmation of the data structure.
  • df.info() reveals that 'Age' has one missing value (None) and shows the total memory usage. It also confirms 'Department' is an object (string) type.
  • df.describe() automatically calculates count, mean, standard deviation, min, max, and quartiles for numeric columns ('Age' and 'Salary'). Note that 'Department' is excluded because it is categorical.

Common mistakes

  • Ignoring Nulls in Describe: describe() excludes NaN values by default. If your count is lower than expected, check df.isnull().sum().
  • Assuming All Columns Are Numeric: Calling mean() on string columns raises an error. Always filter with select_dtypes(include='number') if needed.
  • Overlooking Memory Usage: In large datasets, info() shows memory consumption. Failing to optimize types (e.g., converting int64 to int32) can cause performance issues.
  • Misinterpreting Quartiles: The 25%, 50% (median), and 75% values in describe() help detect skewness. A large gap between mean and median indicates outliers.

When to use it

Use these methods during the initial phase of any data project. Compare them with manual iteration:

MethodBest ForLimitation
df.head()Quick visual verification of row content.Does not show aggregate patterns.
df.info()Checking data types and missing values.Does not provide statistical measures.
df.describe()Understanding distribution of numeric data.Ignores categorical data by default.
Manual LoopingCustom complex checks.Slow and verbose; avoid for basic EDA.

Practice

Guided Exercise: Load a CSV file named sales.csv into a DataFrame. Print the number of rows and columns using shape, then display the first three rows.

Challenge: Using the same DataFrame, find out how many missing values exist in each column and print the mean salary only for employees in the 'IT' department.

Hint: Use df.isnull().sum() for missing values and boolean indexing like df[df['Dept'] == 'IT']['Salary'].mean() for conditional means.

Quick check

Question: Why does df.describe() not return statistics for string columns?

Answer: Because mathematical operations like mean and standard deviation are undefined for categorical text data. You must use include='all' to see counts and unique values for strings.

Summary

Inspecting DataFrames with head(), info(), and describe() provides a rapid, standardized way to understand data quality and distribution. These tools form the foundation of reliable data analysis by preventing assumptions about data types and completeness.

Want to go beyond the notes?

Join Coding Now Tech Institute's Python course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

Analyzing DataFrames – FAQs

Quick answers about learning Analyzing DataFrames in Python.

This free note from Coding Now Tech Institute explains Analyzing DataFrames in Python — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Python topic on Coding Now Tech Institute, including Analyzing DataFrames, is 100% free with no signup required.
With focused practice, most students grasp Analyzing DataFrames in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now