By the end of this lesson, you will be able to read and write text, CSV, and JSON files in Python using standard libraries, ensuring data integrity and proper resource management.
What it is
File handling in Python refers to the process of interacting with external storage systems to persist or retrieve data. The core mental model involves three steps: opening a file (creating a connection), performing operations (reading/writing), and closing the file (releasing resources). Python provides built-in modules for common formats:open() for raw text, csv for tabular data, and json for structured hierarchical data. Related terms include "file modes" (read/write/append) and "serialization" (converting objects to strings).
Why it matters
- Data Persistence: Allows analysis results and datasets to survive program termination.
- Interoperability: CSV and JSON are universal standards, enabling exchange between Python, Excel, databases, and web APIs.
- Memory Efficiency: Reading line-by-line or streaming large files prevents memory overflow compared to loading entire datasets into RAM.
- Automation: Scripts can batch-process thousands of files without manual intervention.
Syntax or steps
The safest way to handle files is using thewith statement (context manager), which automatically closes the file even if errors occur.
- Open: Use
with open('filename', 'mode') as f:. Common modes are'r'(read),'w'(write/overwrite), and'a'(append). - Process: For text, use
f.read()or iterate over lines. For CSV/JSON, pass the file objectfto the respective module's reader/writer. - Close: Handled automatically by exiting the
withblock.
Example
import csv
import json
# 1. Writing and Reading Text File
with open('data.txt', 'w') as f:
f.write("Hello Data Science\n")
with open('data.txt', 'r') as f:
content = f.read()
print(f"Text Content: {content.strip()}")
# 2. Writing and Reading CSV
rows = [["Name", "Age"], ["Alice", 30], ["Bob", 25]]
with open('people.csv', 'w', newline='') as f:
writer = csv.writer(f)
writer.writerows(rows)
with open('people.csv', 'r') as f:
reader = csv.reader(f)
for row in reader:
print(f"CSV Row: {row}")
# 3. Writing and Reading JSON
data = {"name": "Charlie", "skills": ["Python", "SQL"]}
with open('profile.json', 'w') as f:
json.dump(data, f)
with open('profile.json', 'r') as f:
loaded_data = json.load(f)
print(f"JSON Name: {loaded_data['name']}")
Explanation: The text section uses basic string writing. The CSV section uses newline='' on Windows to prevent extra blank lines; writerows writes multiple rows at once. The JSON section uses dump to serialize a dictionary to a file and load to deserialize it back into a Python dict.
Common mistakes
- Forgetting to close files: Leads to resource leaks. Always use
withblocks instead of manualclose(). - Overwriting data: Opening a file in
'w'mode deletes existing content immediately. Use'a'to append. - Incorrect CSV encoding: Failing to specify
encoding='utf-8'can cause errors when reading files with special characters. - Mixing binary and text modes: Do not open image files in text mode (
'r') or text files in binary mode ('rb') unless specifically required.
When to use it
| Format | Best For | Limitations |
|---|---|---|
| Text (.txt) | Logs, simple configurations, unstructured notes. | No structure; parsing requires custom logic. |
| CSV (.csv) | Tabular data, spreadsheets, database exports. | Struggles with nested structures or complex types. |
| JSON (.json) | API responses, configuration files, nested data. | Larger file size than CSV; slower to parse very large files. |
Practice
Guided Exercise: Create a script that readspeople.csv from the example above and prints only the names of people older than 26.
Challenge: Modify the script to write the filtered results to a new file called
young_adults.csv.
Hint: Convert the age string from CSV to an integer before comparing. Use
csv.DictReader for easier column access.
Quick check
Question: Why isnewline='' recommended when opening a file for writing with the csv module?
Answer: It prevents the operating system from adding extra carriage returns, which can result in blank lines between rows in the output file, especially on Windows.
Summary
Effective file handling relies on using context managers (with) to ensure safe resource management and selecting the appropriate format (Text, CSV, or JSON) based on data structure needs. Mastering these patterns allows for robust data ingestion and export pipelines in any data science workflow.