By the end of this lesson, you will be able to create a Pandas DataFrame from raw data and perform basic filtering operations.
What it is
Pandas is an open-source Python library designed for data manipulation and analysis. Its core data structures are the Series (a one-dimensional labeled array) and the DataFrame (a two-dimensional table with rows and columns). Think of a DataFrame as a spreadsheet or SQL table that lives in your Python memory, allowing you to slice, filter, group, and aggregate data efficiently without writing complex loops.
Why it matters
- Efficiency: Operations are vectorized, meaning they run much faster than standard Python lists for large datasets.
- Readability: The syntax resembles natural language queries (e.g., "select where age > 30"), making code easier to maintain.
- Integration: It seamlessly connects with other libraries like NumPy, Matplotlib, and Scikit-learn.
- Handling Messy Data: Built-in methods easily handle missing values (
NaN) and inconsistent data types.
Syntax or steps
- Import the library using the conventional alias:
import pandas as pd. - Create a DataFrame by passing a dictionary of lists or loading from a file (CSV, Excel).
- Inspect the structure using
.head(),.info(), or.describe(). - Filter rows using boolean indexing inside square brackets.
Example
import pandas as pd
# Create sample data
data = {
'Name': ['Alice', 'Bob', 'Charlie', 'David'],
'Age': [25, 30, 35, 40],
'City': ['New York', 'London', 'New York', 'Paris']
}
# Convert to DataFrame
df = pd.DataFrame(data)
# Filter: Show only people older than 30 living in New York
filtered_df = df[(df['Age'] > 30) & (df['City'] == 'New York')]
print(filtered_df)
Explanation: First, we define a dictionary where keys become column headers and values become column data. We convert this into a DataFrame object named df. To filter, we create a boolean mask: df['Age'] > 30 returns a Series of True/False values. We combine conditions using bitwise operators (& for AND, | for OR) because standard Python logical operators (and, or) do not work element-wise on Series objects. Finally, we pass this combined mask back into df[] to retrieve matching rows.
Common mistakes
- Using
and/or: Always use&and|for combining filters. Usingandraises a ValueError. - Missing Parentheses: When combining multiple conditions, each condition must be wrapped in parentheses:
(cond1) & (cond2). - Chained Assignment: Avoid modifying data via chained indexing like
df[df['A']>0]['B'] = 1. Use.loc[]instead to prevent warnings and ensure changes persist. - Ignoring Indexes: Remember that filtering preserves the original index. If you need a clean reset, use
.reset_index(drop=True).
When to use it
| Scenario | Use Pandas | Use Alternatives |
|---|---|---|
| Small to medium tabular data (< 10M rows) | Yes - Easy API, rich features | No |
| Simple list operations | No - Overkill | Yes - Native Python Lists |
| Huge datasets exceeding RAM | No - Memory bound | Yes - Dask, Spark, or Polars |
Practice
Guided Exercise: Create a DataFrame with columns 'Product' and 'Price'. Add three items. Print the average price using .mean().
Challenge: Filter the DataFrame to show only products priced above the average. Hint: Calculate the mean first, store it in a variable, then use that variable in your boolean mask.
Quick check
Q: Why does df[df['Age'] > 30 and df['City'] == 'NY'] fail?
A: Python's and operator expects a single boolean value, but comparing a Series returns an array of booleans. You must use the bitwise & operator which works element-wise.
Summary
Pandas provides a powerful, intuitive interface for working with structured data in Python. By mastering DataFrame creation and boolean indexing, you unlock efficient data exploration and cleaning capabilities essential for any data science workflow.