Back to Data Science Notes
Topic #50

File Handling in Python

By the end of this lesson, you will be able to read and write text, CSV, and JSON files in Python using standard libraries, ensuring data integrity and proper resource management.

What it is

File handling in Python refers to the process of interacting with external storage systems to persist or retrieve data. The core mental model involves three steps: opening a file (creating a connection), performing operations (reading/writing), and closing the file (releasing resources). Python provides built-in modules for common formats: open() for raw text, csv for tabular data, and json for structured hierarchical data. Related terms include "file modes" (read/write/append) and "serialization" (converting objects to strings).

Why it matters

  • Data Persistence: Allows analysis results and datasets to survive program termination.
  • Interoperability: CSV and JSON are universal standards, enabling exchange between Python, Excel, databases, and web APIs.
  • Memory Efficiency: Reading line-by-line or streaming large files prevents memory overflow compared to loading entire datasets into RAM.
  • Automation: Scripts can batch-process thousands of files without manual intervention.

Syntax or steps

The safest way to handle files is using the with statement (context manager), which automatically closes the file even if errors occur.
  1. Open: Use with open('filename', 'mode') as f:. Common modes are 'r' (read), 'w' (write/overwrite), and 'a' (append).
  2. Process: For text, use f.read() or iterate over lines. For CSV/JSON, pass the file object f to the respective module's reader/writer.
  3. Close: Handled automatically by exiting the with block.

Example

import csv
import json

# 1. Writing and Reading Text File
with open('data.txt', 'w') as f:
    f.write("Hello Data Science\n")

with open('data.txt', 'r') as f:
    content = f.read()
    print(f"Text Content: {content.strip()}")

# 2. Writing and Reading CSV
rows = [["Name", "Age"], ["Alice", 30], ["Bob", 25]]
with open('people.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerows(rows)

with open('people.csv', 'r') as f:
    reader = csv.reader(f)
    for row in reader:
        print(f"CSV Row: {row}")

# 3. Writing and Reading JSON
data = {"name": "Charlie", "skills": ["Python", "SQL"]}
with open('profile.json', 'w') as f:
    json.dump(data, f)

with open('profile.json', 'r') as f:
    loaded_data = json.load(f)
    print(f"JSON Name: {loaded_data['name']}")
Explanation: The text section uses basic string writing. The CSV section uses newline='' on Windows to prevent extra blank lines; writerows writes multiple rows at once. The JSON section uses dump to serialize a dictionary to a file and load to deserialize it back into a Python dict.

Common mistakes

  • Forgetting to close files: Leads to resource leaks. Always use with blocks instead of manual close().
  • Overwriting data: Opening a file in 'w' mode deletes existing content immediately. Use 'a' to append.
  • Incorrect CSV encoding: Failing to specify encoding='utf-8' can cause errors when reading files with special characters.
  • Mixing binary and text modes: Do not open image files in text mode ('r') or text files in binary mode ('rb') unless specifically required.

When to use it

FormatBest ForLimitations
Text (.txt)Logs, simple configurations, unstructured notes.No structure; parsing requires custom logic.
CSV (.csv)Tabular data, spreadsheets, database exports.Struggles with nested structures or complex types.
JSON (.json)API responses, configuration files, nested data.Larger file size than CSV; slower to parse very large files.

Practice

Guided Exercise: Create a script that reads people.csv from the example above and prints only the names of people older than 26.
Challenge: Modify the script to write the filtered results to a new file called young_adults.csv.
Hint: Convert the age string from CSV to an integer before comparing. Use csv.DictReader for easier column access.

Quick check

Question: Why is newline='' recommended when opening a file for writing with the csv module?
Answer: It prevents the operating system from adding extra carriage returns, which can result in blank lines between rows in the output file, especially on Windows.

Summary

Effective file handling relies on using context managers (with) to ensure safe resource management and selecting the appropriate format (Text, CSV, or JSON) based on data structure needs. Mastering these patterns allows for robust data ingestion and export pipelines in any data science workflow.

Want to go beyond the notes?

Join Coding Now Tech Institute's Data Science course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

File Handling in Python – FAQs

Quick answers about learning File Handling in Python in Data Science.

This free note from Coding Now Tech Institute explains File Handling in Python in Data Science — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Data Science topic on Coding Now Tech Institute, including File Handling in Python, is 100% free with no signup required.
With focused practice, most students grasp File Handling in Python in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now