By the end of this lesson, you will be able to create, inspect, and manipulate pandas DataFrames to organize structured data in rows and columns.
What it is
A DataFrame is a two-dimensional labeled data structure with columns of potentially different types. Think of it as a spreadsheet or SQL table embedded directly in your Python code. It consists of three main components: index (row labels), columns (column labels), and data (the actual values). The most common library for handling DataFrames in Python is pandas.
Why it matters
- Structured Analysis: Enables efficient filtering, sorting, and aggregation of large datasets without writing complex loops.
- Interoperability: Seamlessly imports data from CSV, Excel, JSON, and databases, and exports back to these formats.
- Vectorized Operations: Mathematical operations apply to entire columns at once, which is significantly faster than iterating row-by-row.
- Missing Data Handling: Provides built-in tools to detect, fill, or drop missing values (
NaN) easily.
Syntax or steps
To create a DataFrame, import pandas and pass a dictionary where keys are column names and values are lists of data. You can also specify a custom index using the index parameter.
import pandas as pd
# Basic creation
df = pd.DataFrame({
"Name": ["Alice", "Bob"],
"Age": [25, 30]
})
Example
The following example creates a small sales dataset, filters it, and calculates a new column.
import pandas as pd
# Create a DataFrame
data = {
"Product": ["Laptop", "Mouse", "Keyboard"],
"Price": [1200, 25, 75],
"Quantity": [10, 50, 30]
}
df = pd.DataFrame(data)
# Add a calculated column
df["Total_Sales"] = df["Price"] * df["Quantity"]
# Filter rows where Price is greater than 50
expensive_items = df[df["Price"] > 50]
print(df)
print("\nExpensive Items:")
print(expensive_items)
Explanation: First, we define a dictionary mapping column headers to lists of values. We convert this into a DataFrame object df. Next, we perform vectorized multiplication on the Price and Quantity columns to create a new Total_Sales column. Finally, we use boolean indexing (df["Price"] > 50) to filter the DataFrame, returning only rows that meet the condition.
Common mistakes
- Mismatched List Lengths: If lists in the initial dictionary have different lengths, pandas raises a
ValueError. Ensure all columns have the same number of entries. - Chained Assignment: Avoid modifying a filtered subset directly like
df[df["A"] > 1]["B"] = 0. This often fails silently. Usedf.loc[df["A"] > 1, "B"] = 0instead. - Ignoring Indexes: When concatenating DataFrames, default indexes may duplicate. Use
ignore_index=Trueif unique sequential indexing is required. - Confusing Series vs. DataFrame: Selecting a single column returns a
Series, not a DataFrame. To keep it as a DataFrame, use double brackets:df[["Price"]].
When to use it
DataFrames are ideal for tabular data analysis. Compare them with NumPy arrays below.
| Feature | Pandas DataFrame | NumPy Array |
|---|---|---|
| Data Types | Heterogeneous (mixed types per column) | Homogeneous (single type) |
| Labels | Labeled axes (rows/columns) | Integer indices only |
| Best For | Exploratory data analysis, cleaning, reporting | High-performance numerical computation |
Practice
Guided Exercise: Create a DataFrame with columns "City" and "Temp". Add a new column "Is_Hot" that is True if Temp > 80, else False.
Challenge: Group the DataFrame by "City" and calculate the average temperature for each city.
Hint: Use df.groupby("City")["Temp"].mean().
Quick check
Question: How do you select the column named "Age" as a DataFrame rather than a Series?
Answer: Use double square brackets: df[["Age"]].
Summary
DataFrames provide a flexible, labeled structure for managing mixed-type tabular data in Python. By leveraging vectorized operations and intuitive indexing, they simplify data cleaning and analysis tasks that would otherwise require verbose loops.