MongoDB: When Documents Genuinely Fit
1 · The lesson
readThe honest 2026 take: Postgres JSONB has eaten roughly 70% of MongoDB's historical use cases. The "schemaless" pitch that won 2014 has been undermined by every team that learned, painfully, that data without a schema is data without invariants. Reach for Mongo when: your data really is hierarchical and read-mostly, when you're horizontally scaling collections beyond what a single Postgres box handles comfortably, or when you've inherited a Mongo deployment and rewriting isn't on the table.
Reach for Postgres + JSONB (db-postgres) for everything else. That said, when Mongo is the right tool, this is what knowing it well looks like.
Examples assume a running local MongoDB — Docker is the easiest way; see devops-docker. Pip-install the driver as shown. Expected output is in comments.
1. The Document Model
Mongo stores BSON documents — JSON plus a few extra types (binary ObjectId, native datetime, Decimal128, binary blobs). Collections are unordered bags of documents; databases are bags of collections.
{
"_id": ObjectId("6553f1..."),
"title": "Postgres vs Mongo in 2026",
"author_id": ObjectId("6553f0..."),
"tags": ["databases", "python"],
"comments": [
{"author": "Alice", "text": "Good post", "at": ISODate("2026-04-12T...")},
{"author": "Alex", "text": "Counterpoint", "at": ISODate("2026-04-13T...")}
],
"stats": {"views": 4210, "shares": 33}
}Two design decisions worth naming up front:
_idis mandatory and unique per collection. Mongo generates anObjectIdif you don't supply one. You can use any unique value — strings, integers, UUIDs — if you have a natural key.- Embedded vs referenced — comments are stored inside the post here. That's "embedding".
author_idpoints to a document in another collection. That's "referencing". Choosing between them is most of Mongo modelling (Section 9).
2. Connecting with PyMongo
pip install pymongo
from pymongo import MongoClient client = MongoClient("mongodb://localhost:27017/") db = client["myapp"] users = db["users"] users.insert_one({"name": "Alice", "email": "alice@example.com", "active": True}) print(users.find_one({"email": "alice@example.com"})) # {'_id': ObjectId('...'), 'name': 'Alice', 'email': 'alice@example.com', 'active': True}
MongoClient is already pooled — share one instance per process. Creating a new client per request gives up the whole point. The connection string honours every Mongo option; production deployments use a SRV record and TLS:
mongodb+srv://user:pass@cluster.example.net/myapp?retryWrites=true&w=majorityretryWrites=true and w=majority are the defaults you want for any production cluster — at-least-once writes acknowledged by a majority of replicas.
3. CRUD
The four-verb tour.
Insert:
users.insert_one({"name": "Alice", "email": "s@example.com"}) users.insert_many([ {"name": "Alex", "email": "alex@example.com"}, {"name": "Kira", "email": "kira@example.com"}, ])
setup added so this can run · defines users
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users')
Find:
users.find_one({"email": "s@example.com"}) list(users.find({"active": True}).limit(10)) list(users.find({"active": True}, {"name": 1, "email": 1, "_id": 0})) # projection
setup added so this can run · defines users
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users')
find() returns a cursor — iterate it, don't list() it on big result sets. Pass a projection (second arg) to limit which fields come back, exactly like SELECT col, col in SQL.
Update:
users.update_one( {"_id": user_id}, {"$set": {"name": "Alice T.", "active": True}}, ) users.update_many( {"active": False}, {"$set": {"archived": True}}, )
setup added so this can run · defines users, user_id
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users') user_id = _AutoMock('user_id')
Delete:
users.delete_one({"_id": user_id}) users.delete_many({"archived": True})
setup added so this can run · defines users, user_id
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users') user_id = _AutoMock('user_id')
update_one and delete_one modify at most one document. The plural variants act on every match — which is also "every document in the collection" if your filter is {}. The delete_everything footgun is real; review every multi-doc operation before it runs.
4. Query Operators
Where SQL has WHERE col > 10 AND col < 20, Mongo has nested dollar-prefixed operators:
users.find({"age": {"$gte": 18, "$lt": 65}}) users.find({"country": {"$in": ["UK", "IE"]}}) users.find({"email": {"$regex": "@example\\.com$"}}) users.find({"deleted_at": {"$exists": False}}) users.find({"$and": [{"active": True}, {"verified": True}]}) users.find({"$or": [{"role": "admin"}, {"role": "owner"}]})
setup added so this can run · defines users
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users')
The common ones to memorise:
| Operator | Meaning |
|---|---|
$eq, $ne | Equal / not equal |
$gt, $gte, $lt, $lte | Comparisons |
$in, $nin | In / not in a list |
$exists | Field is present (or not) |
$regex | Regex match |
$and, $or, $not, $nor | Logical combination |
5. Update Operators
$set is just the start. The atomic-update operators are why Mongo is good at counters and append-only structures:
| Operator | Meaning |
|---|---|
$set | Assign field |
$unset | Remove field |
$inc | Increment numerically |
$push | Append to an array |
$pull | Remove matching items from an array |
$addToSet | Append to an array if not already present |
$pop | Remove first or last element |
posts.update_one( {"_id": post_id}, { "$inc": {"stats.views": 1}, "$push": {"comments": {"author": "Alice", "text": "Hello", "at": datetime.utcnow()}}, }, )
setup added so this can run · defines posts, post_id, datetime
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) posts = _AutoMock('posts') post_id = _AutoMock('post_id') datetime = _AutoMock('datetime')
Both updates above are atomic at the document level — Mongo guarantees no other client sees a half-applied state of that document.
6. Aggregation Pipeline
The pipeline is Mongo's analytical engine — stages chained left to right, each transforming the document stream.
from pymongo import DESCENDING result = users.aggregate([ {"$match": {"active": True}}, {"$group": {"_id": "$country", "count": {"$sum": 1}, "avg_age": {"$avg": "$age"}}}, {"$sort": {"count": DESCENDING}}, {"$limit": 10}, ]) for row in result: print(row) # {'_id': 'UK', 'count': 1240, 'avg_age': 32.4} # {'_id': 'US', 'count': 980, 'avg_age': 35.1} # ...
setup added so this can run · defines users
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users')
Key stages:
| Stage | Equivalent |
|---|---|
$match | WHERE |
$group | GROUP BY |
$project | SELECT col1, col2 |
$sort | ORDER BY |
$limit / $skip | LIMIT / OFFSET |
$lookup | JOIN (slow vs SQL; see Section 9) |
$unwind | Explode an array field into one doc per element |
The pipeline is your reporting layer in pure-Mongo apps. For deep analytics, you'd usually ETL out to a real OLAP store anyway.
7. Indexes
The default _id index is the only one Mongo creates for you. Every other query that runs on a non-trivial collection needs one, or it's a full collection scan.
from pymongo import ASCENDING, DESCENDING users.create_index([("email", ASCENDING)], unique=True) posts.create_index([("author_id", ASCENDING), ("created", DESCENDING)]) posts.create_index([("title", "text")]) # text search index
setup added so this can run · defines users, posts
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) users = _AutoMock('users') posts = _AutoMock('posts')
Compound index order matters. A (author_id, created) index supports queries on author_id alone and on author_id + created, but not on created alone — same prefix rule as SQL composite indexes.
Detecting the missing index: run cursor.explain() and look for COLLSCAN in the plan. If you see one on a query that runs more than rarely, add the index.
unique=True enforces uniqueness at the database level — which is the only place uniqueness belongs.
8. Async with Motor
pymongo is sync. For asyncio apps, use Motor (the official async driver):
pip install motor
from motor.motor_asyncio import AsyncIOMotorClient import asyncio async def main(): client = AsyncIOMotorClient("mongodb://localhost:27017/") users = client.myapp.users await users.insert_one({"name": "Alice", "active": True}) async for user in users.find({"active": True}): print(user) client.close() asyncio.run(main())
setup added so this can run · defines user
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) user = _AutoMock('user')
Motor mirrors PyMongo's API almost call-for-call; every operation is awaitable and find() returns an async iterator.
If you'd rather use Pydantic models on top, Beanie (built on Motor) and odmantic are the modern ODM choices.
9. Embed or Reference
The one modelling question that defines whether your Mongo app stays fast.
Embed when:
- The child data only matters in the context of the parent (comments on a post, line items on an order).
- One-to-few cardinality — under a few hundred children.
- The parent is the natural read unit ("show me a post" naturally wants its comments).
- Total document stays well under the 16 MB BSON limit.
Reference when:
- The child is independently queryable (users → posts; you query "all posts by X").
- One-to-many or many-to-many at any scale.
- The child is shared across parents.
# Embedded — one read fetches everything to render the post page { "_id": post_id, "title": "...", "comments": [{"author": "Alice", "text": "..."}] } # Referenced — author lives in users; posts link via author_id {"_id": post_id, "title": "...", "author_id": author_obj_id} {"_id": author_obj_id, "name": "Alice", "email": "..."}
setup added so this can run · defines post_id, author_obj_id
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) post_id = _AutoMock('post_id') author_obj_id = _AutoMock('author_obj_id')
$lookup joins them in the aggregation pipeline — but it's substantially slower than a SQL JOIN because Mongo isn't built for it. If you find yourself reaching for $lookup constantly, that's a smell — your data is relational and Postgres is the right tool.
10. Transactions
Mongo transactions exist (since 4.0) and span multiple documents and collections. They require a replica set (a single standalone Mongo doesn't support them) and impose performance overhead.
with client.start_session() as session: with session.start_transaction(): accounts.update_one({"_id": from_id}, {"$inc": {"balance": -amount}}, session=session) accounts.update_one({"_id": to_id}, {"$inc": {"balance": amount}}, session=session) # commit on clean exit; abort on exception
setup added so this can run · defines client, accounts, from_id, to_id, amount
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) client = _AutoMock('client') accounts = _AutoMock('accounts') from_id = _AutoMock('from_id') to_id = _AutoMock('to_id') amount = 1
The advice: minimise transactions. Mongo's design assumption is one atomic document at a time. If you constantly need multi-document atomicity, the data is probably relational — and again, Postgres is the right tool.
11. When Mongo Fails You
Be honest about where the tool stops being a fit.
- Cross-collection relational queries —
$lookupworks, but slowly, and aggregation pipelines get gnarly fast. - Strong consistency across multiple documents — possible via transactions, but at a cost; Postgres does this trivially.
- Mature analytics tooling — every BI tool speaks SQL natively; many speak Mongo via translation layers that lose features.
- Constrained data shapes — Postgres
NOT NULL/CHECKconstraints and foreign keys catch entire classes of bugs at the database; Mongo schema validation (JSON Schema based) is improving but optional and lighter-touch. - Tight memory — Mongo wants RAM; the working set should fit in memory for acceptable performance.
If you're three months into a project and reaching for Mongo features that don't exist or perform poorly, the migration to Postgres is rarely fun but usually correct.
Common Mistakes
1. No indexes on common queries. Default behaviour is a collection scan. On a 10M-document collection that's hundreds of milliseconds per query, then a kernel page-fault storm, then your service falls over. explain() and create_index() on day one.
2. Storing dates as strings. "2026-05-14" sorts lexically, which works for ISO 8601 strings — until someone writes "14/05/2026". Use real datetime objects; Mongo stores them as BSON dates and indexes them properly.
3. Over-embedding. A document that grows unbounded (every event in a user's life embedded into the user doc) hits the 16 MB cap and slows every read to a crawl. Cap embedded arrays at "small and bounded" — comments-on-a-post yes; events-forever no.
4. Modelling like a relational DB with refs everywhere. Every document points to three others and every operation is three round trips. That's not Mongo's strength — that's an indication you should be in Postgres. Embed where Mongo wants you to embed.
5. Using Mongo when Postgres+JSONB would do. "Our data is hierarchical" is true of almost everyone's data. JSONB in Postgres gives you hierarchical storage plus joins, transactions, constraints, and mature tooling.
6. One MongoClient per request. MongoClient is a connection pool. Build one per process at startup; share it.
7. Not setting _id deliberately. Letting Mongo generate ObjectIds is fine. But if your natural key is the email, use the email as _id and skip the secondary unique index.
🎯 Your Turn — Blog Schema with Embedded Comments and Referenced Authors
Model a tiny blog: users and posts. Embed comments inside posts (small, read-with-post). Reference author_id to link posts to users. Implement four functions.
Skeleton:
from datetime import datetime from pymongo import MongoClient, ASCENDING from bson import ObjectId client = MongoClient("mongodb://localhost:27017/") db = client["blogdemo"] users = db["users"] posts = db["posts"] users.create_index([("email", ASCENDING)], unique=True) posts.create_index([("author_id", ASCENDING)]) def create_user(name: str, email: str) -> ObjectId: # TODO 1 ... def create_post(author_id: ObjectId, title: str, body: str) -> ObjectId: # TODO 2: insert with empty comments array; return new id ... def add_comment(post_id: ObjectId, author_name: str, text: str) -> None: # TODO 3: $push a comment onto the embedded comments array ... def posts_by_author_name(name: str) -> list[dict]: # TODO 4: aggregation pipeline — match user by name, $lookup posts, # return post docs with author name attached ...
Hint 1 — $push and the comments array
posts.update_one({"_id": post_id}, {"$push": {"comments": {...}}}). Store the comment as a small embedded doc with author, text, and an at datetime.
Hint 2 — $lookup as a JOIN
Start the pipeline on theusers collection ($match by name), then $lookup from posts using localField: "_id" and foreignField: "author_id". The matched posts come back nested under a new field; $unwind to flatten and $project to shape the output.
Show full solution
from datetime import datetime from pymongo import MongoClient, ASCENDING from bson import ObjectId client = MongoClient("mongodb://localhost:27017/") db = client["blogdemo"] users = db["users"] posts = db["posts"] users.create_index([("email", ASCENDING)], unique=True) posts.create_index([("author_id", ASCENDING)]) def create_user(name: str, email: str) -> ObjectId: res = users.insert_one({"name": name, "email": email, "created": datetime.utcnow()}) return res.inserted_id def create_post(author_id: ObjectId, title: str, body: str) -> ObjectId: res = posts.insert_one({ "author_id": author_id, "title": title, "body": body, "comments": [], "created": datetime.utcnow(), }) return res.inserted_id def add_comment(post_id: ObjectId, author_name: str, text: str) -> None: posts.update_one( {"_id": post_id}, {"$push": {"comments": {"author": author_name, "text": text, "at": datetime.utcnow()}}}, ) def posts_by_author_name(name: str) -> list[dict]: pipeline = [ {"$match": {"name": name}}, {"$lookup": { "from": "posts", "localField": "_id", "foreignField": "author_id", "as": "posts", }}, {"$unwind": "$posts"}, {"$project": { "_id": "$posts._id", "title": "$posts.title", "body": "$posts.body", "author": "$name", "comments": "$posts.comments", "created": "$posts.created", }}, ] return list(users.aggregate(pipeline)) # Demo surya_id = create_user("Alice", "alice@example.com") post_id = create_post(surya_id, "First post", "Hello, Mongo") add_comment(post_id, "Alex", "Nice one") add_comment(post_id, "Kira", "Welcome") for p in posts_by_author_name("Alice"): print(p["title"], "by", p["author"], "—", len(p["comments"]), "comments") # First post by Alice — 2 comments
What the solution gets right:
- Indexes up front —
emailunique on users,author_idon posts. Both queries we run hit indexes. - Embedded comments — small bounded array per post; one read fetches the post with its comments.
- Referenced author — authors are independent entities you might list, edit, or query on their own.
$pushfor comments — atomic at the document; concurrent comments don't race.- Aggregation
$lookupfor the join — works fine here because the lookup is one-to-many and the join key is indexed. - Datetimes as real
datetime— not strings.
For a high-traffic version, denormalise the author name onto the post itself ({"author_id": ..., "author_name": "Alice"}) and update both places when the user renames — trading write complexity for read speed. That trade-off is the whole Mongo modelling craft.
What You Learned
- Documents and BSON — JSON plus
ObjectId,datetime,Decimal128, binary. - One
MongoClientper process — it's already pooled. - CRUD with
_onevs_many— the_manyvariants are powerful and dangerous. - Query operators —
$gt,$in,$exists,$regex,$and,$or— nested-dict style. - Update operators —
$set,$inc,$push,$pull,$addToSet— atomic at the document. - Aggregation pipelines for grouping, sorting, projection, and (carefully) joins via
$lookup. - Indexes are mandatory; compound-index order follows the prefix rule.
- Async via Motor; ODMs Beanie and odmantic on top.
- Embed for read-with-parent, reference for independent or shared data.
- Transactions exist but signal a relational model. Use sparingly.
- Be honest about the fit — Postgres + JSONB has eaten most of Mongo's old territory.
Next: ORM Patterns That Don't Blow Up in Production — N+1, eager loading, bulk ops, migrations, and when to drop to raw SQL.