Learn how to add single or multiple documents to a MongoDB collection using PyMongo, ensuring data persistence and efficient batch processing.
What it is
In the context of Python and MongoDB, "inserting documents" refers to adding new records to a collection. A document is essentially a JSON-like object (dictionary in Python) that stores data. Unlike relational databases where you define a strict schema upfront, MongoDB allows flexible structures, though consistency is recommended. The primary methods for this operation are insert_one() for single records and insert_many() for batches. Related terms include collections (tables), databases (schemas), and _id (the unique identifier automatically generated if not provided).
Why it matters
- Data Persistence: It is the fundamental step to saving application state or user input permanently.
- Performance Optimization: Using
insert_many()reduces network round-trips compared to looping through individual inserts. - Schema Flexibility: Allows rapid iteration during development without complex migration scripts.
- Atomicity Control: Provides options to handle errors gracefully when some documents fail to insert.
Syntax or steps
To insert data, you first need an active connection to a database and a reference to a specific collection. Once you have the collection object, you pass a dictionary representing the document as an argument. For multiple documents, you pass a list of dictionaries. If a document lacks an _id field, MongoDB generates a unique ObjectId automatically.
Example
from pymongo import MongoClient
# Connect to local MongoDB instance
client = MongoClient("mongodb://localhost:27017/")
db = client["my_database"]
collection = db["users"]
# Insert a single document
single_doc = {"name": "Ada", "age": 36}
result_one = collection.insert_one(single_doc)
print(f"Inserted ID: {result_one.inserted_id}")
# Insert multiple documents
many_docs = [
{"name": "Grace", "age": 45},
{"name": "Alan", "age": 41}
]
result_many = collection.insert_many(many_docs)
print(f"Inserted IDs: {result_many.inserted_ids}")
The code above establishes a connection and defines two operations. First, insert_one() adds Ada's record, returning an InsertOneResult object containing the generated _id. Second, insert_many() processes a list of dictionaries for Grace and Alan, returning an InsertManyResult with a list of all inserted IDs. This demonstrates both atomic single writes and efficient batch writes.
Common mistakes
- Forgetting to close connections: While PyMongo handles pooling well, explicitly closing clients in long-running scripts prevents resource leaks.
- Using deprecated methods: Avoid
save()orupdate()for inserting new records; useinsert_one()instead. - Ignoring duplicate key errors: If you manually specify
_id, ensure uniqueness to avoidDuplicateKeyError. - Large batch sizes: Sending extremely large lists in one
insert_many()call can exceed memory limits or timeout thresholds; chunk large datasets.
When to use it
Choose between single and bulk insertion based on volume and error handling needs.
| Method | Best Use Case | Error Handling |
|---|---|---|
insert_one() |
User form submissions, real-time events. | Fails immediately if the single doc is invalid. |
insert_many() |
Data migrations, log ingestion, seeding DB. | Can continue past errors if ordered=False. |
Practice
Guided Exercise: Create a script that connects to a test database and inserts three books into a "library" collection. Each book should have a title, author, and year. Print the inserted IDs.
Challenge: Modify the script to attempt inserting a document with a missing required field (simulate by checking logic before insert). How does insert_many() behave if one item in the list causes a validation error? Hint: Look at the ordered parameter.
Quick check
Question: What happens if you do not provide an _id field when calling insert_one()?
Answer: MongoDB automatically generates a unique ObjectId for the document.
Summary
Inserting documents in PyMongo is straightforward using insert_one() for individual records and insert_many() for batches. Understanding these methods ensures efficient data storage and proper handling of unique identifiers within your NoSQL applications.