By the end of this lesson, you will be able to use Python's re module to search for patterns in text, extract specific data using groups, and replace substrings efficiently.
What it is
Regular expressions (regex) are sequences of characters that define a search pattern. In Python, the re module provides functions to work with these patterns. Think of regex as a powerful "find and replace" tool that understands structure rather than just literal text. Key concepts include patterns (the rule), targets (the text being searched), and matches (the results). Related terms include flags (modifiers like case-insensitivity) and groups (parenthesized sections used for extraction).
Why it matters
- Data Extraction: Pulling emails, phone numbers, or dates from unstructured text logs.
- Validation: Ensuring user input matches expected formats (e.g., passwords or IDs).
- Text Cleaning: Removing unwanted whitespace or special characters during preprocessing.
- Complex Search: Finding words based on prefixes, suffixes, or character classes without writing multiple loops.
Syntax or steps
The most common workflow involves three steps: compile the pattern (optional but efficient for reuse), search the text, and process the result. The primary functions are re.search() (finds the first match anywhere), re.findall() (returns all non-overlapping matches), and re.sub() (replaces matches).
- Import the module:
import re. - Define your pattern string. Use raw strings (
r'...') to avoid escaping backslashes. - Call a function like
re.search(pattern, text). - Check if a match object was returned before accessing data.
Example
import re
text = "Contact us at support@example.com or sales@company.org."
pattern = r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b'
# Find all email addresses
emails = re.findall(pattern, text)
print("Found emails:", emails)
# Replace emails with '[REDACTED]'
cleaned_text = re.sub(pattern, '[REDACTED]', text)
print("Cleaned text:", cleaned_text)
Explanation: The pattern uses \b for word boundaries to ensure we don't catch partial words. [A-Za-z0-9._%+-]+ matches the local part of the email, while @ is literal. The domain part follows similar logic, ending with a top-level domain check \.[A-Z|a-z]{2,}. re.findall returns a list of strings, while re.sub returns the modified original string.
Common mistakes
- Forgetting Raw Strings: Writing
'\d'instead ofr'\d'can cause issues because Python interprets\das an escape sequence before regex sees it. Always user''. - Ignoring None Returns:
re.search()returnsNoneif no match is found. Calling.group()onNonecrashes the program. Always checkif match:. - Greedy Matching: Patterns like
.*consume as much text as possible. If you want minimal matching, use.*?. - Catastrophic Backtracking: Complex nested quantifiers (like
(a+)+) can hang your program on long inputs. Keep patterns simple and specific.
When to use it
Regex is ideal for complex pattern matching but overkill for simple tasks. Compare it with standard string methods:
| Task | String Method | Regex |
|---|---|---|
| Check if substring exists | 'abc' in text | Overkill |
| Split by comma | text.split(',') | Unnecessary |
| Find all digits | Loop + isdigit() | re.findall(r'\d+', text) |
| Validate email format | Not feasible | Standard approach |
Practice
Guided Exercise: Write a regex to find all 4-digit years in the string "The events were in 1999, 2005, and 2023."
Challenge: Modify the pattern to only match years between 1900 and 1999.
Hint: For the challenge, use r'19\d{2}'.
Quick check
Question: What does re.match() do differently than re.search()?
Answer: re.match() checks for a match only at the beginning of the string, whereas re.search() scans the entire string for the first occurrence.
Summary
Python's re module enables precise text manipulation through pattern matching. Mastering raw strings, understanding greedy vs. lazy quantifiers, and knowing when to choose regex over simple string methods are key skills for robust text processing.