By the end of this lesson, you will be able to select and manipulate Python’s core data structures—lists, tuples, dictionaries, and sets—to store and process data efficiently in data science workflows.
What it is
Python provides four primary built-in data structures that serve distinct purposes. A list is an ordered, mutable collection of items, ideal for sequences where order matters and elements may change. A tuple is an ordered, immutable collection, often used for fixed records or keys in dictionaries. A dictionary is an unordered (in older Python versions) or insertion-ordered (Python 3.7+) collection of key-value pairs, perfect for mapping relationships. A set is an unordered collection of unique elements, useful for membership testing and eliminating duplicates. Understanding these structures is fundamental because they form the backbone of data manipulation before converting them into more specialized formats like Pandas DataFrames or NumPy arrays.Why it matters
- Efficiency: Sets provide O(1) average time complexity for membership checks, far faster than lists for large datasets.
- Data Integrity: Tuples ensure that critical configuration parameters or coordinate points remain unchanged during processing.
- Flexibility: Dictionaries allow dynamic labeling of features or metadata, making code more readable than index-based access.
- Interoperability: Most data science libraries accept these native structures as input, serving as a bridge between raw data and analysis tools.
Syntax or steps
To create these structures, use specific literals: square brackets[] for lists, parentheses () for tuples, curly braces with colons {key: value} for dictionaries, and curly braces without colons {item1, item2} for sets. Accessing elements varies: lists and tuples use integer indices (e.g., my_list[0]), while dictionaries use keys (e.g., my_dict['name']). Modifying data involves assignment for lists and dictionaries, but not for tuples or sets directly via index/key; instead, use methods like .add() for sets or reassign variables for tuples.
Example
# Define different data structures
user_ids = [101, 102, 103] # List: ordered, mutable
coordinates = (45.5, -122.6) # Tuple: ordered, immutable
user_profile = { # Dictionary: key-value pairs
"id": 101,
"name": "Alice",
"active": True
}
unique_tags = {"python", "data", "ai"} # Set: unique, unordered
# Practical operations
user_ids.append(104) # Add to list
print(f"First ID: {user_ids[0]}") # Access by index
print(f"User Name: {user_profile['name']}") # Access by key
# Check membership efficiently
if "python" in unique_tags: # Fast lookup in set
print("Tag found")
# Convert list to set to remove duplicates
duplicate_list = [1, 2, 2, 3, 3, 3]
cleaned_set = set(duplicate_list)
print(f"Unique values: {cleaned_set}")
This example demonstrates creation, modification, and access patterns. The list allows appending new IDs. The tuple stores coordinates safely. The dictionary maps user attributes to values for easy retrieval. The set enables fast checking of tags and removing duplicates from a list.
Common mistakes
- Mutating Tuples: Attempting to assign a value to a tuple index (e.g.,
tup[0] = 5) raises a TypeError. Fix: Create a new tuple if changes are needed. - Unhashable Keys: Using a list or dictionary as a key in another dictionary causes a TypeError. Fix: Use tuples or strings as keys.
- Assuming Order in Sets: Iterating over a set does not guarantee the same order every time. Fix: Sort the set (
sorted(my_set)) if order is required. - Modifying While Iterating: Changing a list or dictionary size during iteration can cause unexpected behavior. Fix: Iterate over a copy or use list comprehensions.
When to use it
| Structure | Best For | Alternative |
|---|---|---|
| List | Ordered sequences needing frequent addition/removal. | NumPy Array (for numerical math) |
| Tuple | Fixed records, function return values, dict keys. | NamedTuple (for readability) |
| Dictionary | Mapping labels to values, JSON-like data. | Pandas Series/DataFrame (tabular) |
| Set | Unique items, membership tests, mathematical sets. | Bloom Filter (for massive scale) |
Practice
Guided Exercise: Create a dictionary representing a book with keys 'title', 'author', and 'year'. Print the author's name. Then, add a new key 'genre' with value 'Fiction'.Challenge: Given a list of email addresses
["a@x.com", "b@y.com", "a@x.com"], use a set to find how many unique domains exist. Hint: Extract domains using string splitting, then convert to a set.