By the end of this lesson, you will be able to use Python's re module to search for and extract specific patterns within strings using regular expressions.
What it is
Regular expressions (regex) are a sequence of characters that define a search pattern. In Python, the re module provides functions to work with these patterns. Think of regex as a powerful text-processing tool that allows you to match complex structures like email addresses, phone numbers, or dates, rather than just simple substrings. Key concepts include patterns (the rule), strings (the data), and matches (the result).
Why it matters
- Data Extraction: Quickly pull specific information from unstructured text, such as finding all numbers in a log file.
- Validation: Ensure user input meets specific criteria, like checking if an email address has the correct format.
- Text Replacement: Modify parts of a string based on patterns, such as masking sensitive data.
- Flexibility: Handle variations in text formatting without writing extensive conditional logic.
Syntax or steps
The most common workflow involves importing the module and calling a function like findall(), search(), or sub(). Patterns are typically written as raw strings (prefixed with r) to prevent Python from interpreting backslashes as escape characters before the regex engine sees them.
- Import the module:
import re - Define the pattern using regex syntax.
- Call a method passing the pattern and the target string.
Example
import re
text = "Order 101 was placed by user 45678 on 2023-10-05."
# Find all sequences of digits
numbers = re.findall(r"\d+", text)
print(numbers)
Explanation:
import re: Loads the regular expression library.r"\d+": The pattern.\dmatches any digit (0-9). The+quantifier means "one or more" of the preceding character. Therprefix makes it a raw string.re.findall(...): Scans the entire string and returns a list of all non-overlapping matches.- Output:
['101', '45678']
Common mistakes
- Forgetting raw strings: Using
"\d"instead ofr"\d"can cause issues because Python interprets\das an escape sequence first. Always use raw strings for regex patterns. - Confusing
match()andsearch():match()only checks for a match at the beginning of the string, whilesearch()scans the whole string. Usesearch()unless you specifically need start-of-string anchoring. - Overly greedy quantifiers: By default,
*and+are greedy (they match as much as possible). If you need the shortest match, add a?after the quantifier (e.g.,.*?). - Ignoring case sensitivity: Regex is case-sensitive by default. Add the flag
re.IGNORECASEif you want to match both uppercase and lowercase letters.
When to use it
Compare regex with standard string methods (str.find(), in operator).
| Feature | Standard String Methods | Regular Expressions |
|---|---|---|
| Complexity | Low; easy to read for simple tasks. | High; requires learning syntax. |
| Performance | Faster for exact substring searches. | Slower due to pattern compilation/matching overhead. |
| Use Case | Checking if "apple" is in a sentence. | Extracting all dates in YYYY-MM-DD format. |
Use standard methods for simple containment checks. Use regex when you need to match patterns, positions, or variable formats.
Practice
Guided Exercise: Write a script that extracts all words starting with the letter 'P' from the string "Python is popular and powerful.".
Hint: Use the pattern r"\bP\w*". \b is a word boundary, P is the literal letter, and \w* matches zero or more word characters.
Challenge: Modify the code to validate if a string looks like a US phone number (e.g., 123-456-7890). Return True if it matches, False otherwise.
Quick check
Question: What does the regex pattern r"[a-z]+" match?
Answer: It matches one or more consecutive lowercase letters.
Summary
Python's re module enables powerful text processing through pattern matching. While standard string methods suffice for simple tasks, regex is essential for extracting or validating structured data within unstructured text. Mastering basic syntax like \d, \w, and quantifiers unlocks efficient data cleaning capabilities.