By the end of this lesson, you will be able to create Python sets, understand their unique properties, and perform basic mathematical operations like union and intersection.
What it is
A set in Python is an unordered collection of unique elements. Unlike lists or tuples, sets do not allow duplicate values and are mutable (you can add or remove items). They are defined using curly braces {} or the set() constructor. Sets are ideal for membership testing and eliminating duplicate entries.
Key related terms include hashable (elements must be immutable types like numbers, strings, or tuples) and immutable set (frozenset), which cannot be changed after creation.
Why it matters
- Uniqueness: Automatically removes duplicates from data collections.
- Performance: Membership checks (
x in my_set) are significantly faster than in lists because sets use hash tables. - Mathematical Operations: Supports standard set theory operations like union, intersection, difference, and symmetric difference.
- Data Cleaning: Useful for finding common items between two datasets or identifying missing data.
Syntax or steps
To create a set, enclose comma-separated values in curly braces. Note that an empty pair of braces creates a dictionary, so use set() for an empty set.
# Creating a set with initial values
my_set = {1, 2, 3}
# Adding an element
my_set.add(4)
# Removing an element
my_set.remove(2)
Example
The following example demonstrates creating sets from lists to remove duplicates and performing union and intersection operations.
# Define two lists with some overlapping and unique items
list_a = [1, 2, 3, 4, 5]
list_b = [4, 5, 6, 7, 8]
# Convert lists to sets to ensure uniqueness
set_a = set(list_a)
set_b = set(list_b)
print(f"Set A: {set_a}")
print(f"Set B: {set_b}")
# Union: All unique elements from both sets
union_result = set_a | set_b
print(f"Union (A | B): {union_result}")
# Intersection: Elements present in both sets
intersection_result = set_a & set_b
print(f"Intersection (A & B): {intersection_result}")
# Difference: Elements in A but not in B
difference_result = set_a - set_b
print(f"Difference (A - B): {difference_result}")
Explanation: First, we convert lists to sets using set(). This automatically discards any duplicates within each list. We then use the pipe operator | for union, the ampersand & for intersection, and the minus sign - for difference. These operators return new sets containing the calculated results.
Common mistakes
- Trying to store unhashable types: You cannot put lists or dictionaries inside a set. Use tuples instead if you need nested structures.
- Assuming order: Sets are unordered. Do not rely on the sequence of elements when iterating; use a list if order matters.
- Using
{}for empty sets:empty = {}creates a dictionary. Useempty = set()instead. - Modifying during iteration: Changing a set while looping through it raises a
RuntimeError. Create a copy first if modification is needed.
When to use it
Compare sets with lists based on your primary need:
| Feature | List | Set |
|---|---|---|
| Duplicates | Allowed | Not Allowed |
| Ordering | Ordered | Unordered |
| Membership Test Speed | O(n) - Slow for large data | O(1) - Fast |
| Best For | Sequences, indexing | Unique items, math ops |
Use a list when you need to maintain insertion order or access items by index. Use a set when you need to check for existence quickly or ensure all items are unique.
Practice
Guided Exercise: Create two sets representing students in two different clubs. Find the students who are in both clubs.
Challenge: Given a list of words ['apple', 'banana', 'apple', 'orange'], write code to print only the unique words using a set.
Hint: Convert the list to a set, then back to a list if you need to iterate over them in a specific way, though printing the set directly works for uniqueness.
Quick check
Question: What happens if you try to add a duplicate item to an existing set?
Answer: Nothing changes. The set ignores the addition because it already contains that unique element.
Summary
Python sets are powerful tools for handling unique data and performing fast membership tests. By leveraging operations like union and intersection, you can efficiently analyze relationships between different collections of data without worrying about duplicates.