20 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.
Code lives in labs/course-10-mongodb/.
NOTE
Both halves run. Every experiment has a mongosh script, NN_name.js — what the lab
examiner will ask you to demonstrate — and it is run on MongoDB 8.3.7, typed into
mongosh 2.12.0 line by line as you would at the prompt, on a fresh server each time.
Under 5. Execution and Results is the session: each statement after its prompt, then what
the shell printed. Sixteen also have a Python half, NN_name.py, the same query logic
through mongomock, asserted by tools/data-science/run_mongo_labs.py.
Experiment 17 runs on a real three-member replica set (three mongod processes, as its
section 0 describes, and a fourth as an arbiter), Experiment 18 with the real mongofiles,
and Experiment 19 on a replica set, which transactions need.
Until October 2026 mongod could not be installed where these labs are checked, and every
script said NOT EXECUTED. MongoDB's own download hosts are still blocked; conda-forge's builds
of the same server can be reached, and tools/data-science/setup_mongodb.sh installs them.
Running the scripts found eleven of them doing something other than what their comments said
— a placeholder that is a syntax error, a variable never set, lines that start with a dot,
statements on data that was not there, MongoDB 8 refusing two old forms — each corrected in
its file, with a note, and listed under its experiment below.
tools/data-science/setup_mongodb.sh # mongod and mongofiles, in /tmp/mongodb
npm --prefix tools/data-science install # mongosh
python3 tools/data-science/run_mongo_labs.py # both halves of every experiment
Two things in the output change from run to run: the ObjectIds, dates and UUIDs a server makes
new each time, and everything about a replica set's election. capture_lab_outputs.py --check
compares every other character of a session exactly; for Experiment 17 it reruns the
replica set and checks what the experiment shows, as the driver's assertions.
Every experiment from 3 onwards starts from the same five students, with the courses and
enrolments Experiment 16 joins. 00_sample_data.js loads them: each script that needs them runs
load("00_sample_data.js") straight after use collegeDB, so it can be run on its own, as often
as you like. Start mongosh in the labs folder for load() to find it. The data is exactly that
in fixtures.py, which the Python halves load, and run_mongo_labs.py checks that the two agree.
// The sample data every experiment from 3 onwards starts from: the five students
// of Experiment 2, with the courses and enrollments that Experiment 16 joins.
// It is the data in fixtures.py, which the Python halves load, exactly;
// tools/data-science/run_mongo_labs.py checks that the two agree.
//
// Each script loads it in its second line, with load("00_sample_data.js"), so
// it can be run on its own, as often as you like: start mongosh in this folder.
// It drops and refills only these three collections.
// Step 1: Switch to collegeDB
db = db.getSiblingDB("collegeDB")
// Step 2: Refill the students
db.students.drop()
db.students.insertMany([
{"_id": 21, "name": "Asha", "dept": "DS", "marks": {"maths": 88, "stats": 91}, "subjects": ["DS", "Stats", "Python"], "age": 20, "active": true},
{"_id": 22, "name": "Ravi", "dept": "DS", "marks": {"maths": 65, "stats": 58}, "subjects": ["DS", "Python"], "age": 21, "active": true},
{"_id": 23, "name": "Meena", "dept": "Stats", "marks": {"maths": 94, "stats": 89}, "subjects": ["Stats", "R"], "age": 20, "active": true},
{"_id": 24, "name": "Kiran", "dept": "DS", "marks": {"maths": 71, "stats": 66}, "subjects": ["DS"], "age": 22, "active": false},
{"_id": 25, "name": "Bhanu", "dept": "Stats", "marks": {"maths": 52, "stats": 47}, "subjects": ["Stats"], "age": 21, "active": true}
])
// Step 3: Refill the courses
db.courses.drop()
db.courses.insertMany([
{"_id": "DSC301", "title": "Data Science with R", "credits": 4, "instructor": "Dr. Rao"},
{"_id": "STA302", "title": "Statistical Foundations", "credits": 3, "instructor": "Dr. Devi"},
{"_id": "WEB303", "title": "Web Technologies", "credits": 3, "instructor": "Dr. Kumar"}
])
// Step 4: Refill the enrollments
db.enrollments.drop()
db.enrollments.insertMany([
{"student_id": 21, "course_id": "DSC301", "grade": "A"},
{"student_id": 21, "course_id": "STA302", "grade": "B"},
{"student_id": 22, "course_id": "DSC301", "grade": "C"},
{"student_id": 23, "course_id": "STA302", "grade": "A"},
{"student_id": 24, "course_id": "WEB303", "grade": "B"}
])
// Step 5: Say what was loaded
print("sample data loaded: " + db.students.countDocuments() + " students, " +
db.courses.countDocuments() + " courses, " + db.enrollments.countDocuments() + " enrollments")
For the lab exam you need a real server. Three routes:
| Route | Command |
|---|---|
| MongoDB Atlas | Free tier, no install — and it gives you a real replica set, which a local install does not |
| Docker | docker run -d -p 27017:27017 --name mongo mongo |
| Local package | apt install mongodb-org, or the platform installer |
mongosh # localhost:27017
mongosh "mongodb+srv://user:pass@cluster.mongodb.net/collegeDB"
Use Atlas or Docker. A local install commits you to managing a service, and Atlas is the only one of the three that gives you a replica set — which experiments 17 and 19 both require.
Install MongoDB, and use the Mongo Shell and Compass.
Prove a server is running, find your way round mongosh, and know what Compass adds.
THE POINT
MongoDB Compass is the official GUI: browse collections, build queries
without typing them, and read explain() output as a diagram rather than
JSON. Worth installing for the explain visualiser alone.
Know for the viva: the default port is 27017; mongosh is a full
JavaScript REPL, so loops and variables work in it; and show dbs will not
list a database until something has been written to it.
// Experiment 1 -- Installing MongoDB, the Mongo shell and Compass.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. There is no .py half: these are server commands, with no query logic
// to run. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
// Step 1: Connect
// From a terminal, NOT from inside mongosh:
//
// mongosh // localhost:27017
// mongosh "mongodb://localhost:27017/collegeDB" // straight into a db
// mongosh "mongodb+srv://user:pass@cluster.mongodb.net/collegeDB" // Atlas
//
// 27017 is the default port. Remember it -- it is asked in the viva.
// Step 2: Prove the install worked
db.version() // e.g. "7.0.14"
db.serverStatus().host // hostname:port this shell is attached to
Math.floor(db.serverStatus().uptime / 60) // whole minutes since mongod started
// [Changed: this was db.serverStatus().uptime, in seconds, which on a server
// started a moment ago reads 1 or 2 depending on the moment. In minutes it is 0.]
db.hostInfo().os // the OS the server is running on
db.hostInfo().system.numCores // the cores it can see
// [Changed: this was db.hostInfo(), which prints some 350 lines, most of them
// disk counters that change by the second. These are the two lines that matter.]
show dbs // the databases that have been WRITTEN to
show collections // collections in the CURRENT database
db // which database am I in?
db.getMongo() // the connection string
// Step 3: Use the shell as a JavaScript REPL
use collegeDB
load("00_sample_data.js") // the five students, to have something to count
// This is the fact students most often miss, and it is worth demonstrating.
const depts = ["DS", "Stats", "CS"]
for (const d of depts) {
print(`${d}: ${db.students.countDocuments({ dept: d })}`)
}
// Variables persist across statements; functions can be defined and reused.
function topper(dept) {
return db.students.find({ dept }).sort({ "marks.maths": -1 }).limit(1).toArray()[0]
}
topper("DS")
// Load a script file from disk -- how you would run the rest of these labs:
// load("02_create_insert.js")
// Step 4: Run the administrative commands
db.adminCommand({ listDatabases: 1 })
const s = db.stats() // size, collection count, index count
({ collections: s.collections, objects: s.objects, indexes: s.indexes, dataSize: s.dataSize })
const c = db.students.stats() // per-collection: documents, size, indexes
({ count: c.count, size: c.size, avgObjSize: c.avgObjSize, nindexes: c.nindexes })
// [Changed: these were db.stats() and db.students.stats() in full, which print
// the disk's free space and some 400 lines of storage-engine counters, both of
// which change from one run to the next. These are the figures the comments
// are about.]
db.getCollectionNames().sort() // sorted: the server lists them in no fixed order
// Step 5: Clean up
use collegeDB // switches even if collegeDB does not exist
db.dropDatabase() // no confirmation, no undo
// --- MongoDB Compass ---------------------------------------------------------
// The official GUI (a separate download from the server).
//
// * browse collections and documents without writing find()
// * the Schema tab INFERS a schema from a sample -- the fastest way to see
// what shape the documents in an inherited collection actually are
// * the Explain Plan tab draws explain() output as a diagram instead of JSON,
// which is worth the install on its own
// * the Aggregations tab builds a pipeline stage by stage, showing the
// intermediate documents after EACH stage -- exactly what you need when a
// pipeline returns nothing and you cannot see which stage emptied it
//
// --- Know for the viva -------------------------------------------------------
// * default port 27017
// * mongosh is a full JavaScript REPL -- loops, variables, functions
// * `show dbs` does NOT list a database until something has been written to
// it: `use newdb` alone creates nothing
// * the data directory defaults to /var/lib/mongodb (Linux); mongod refuses
// to start if it does not exist or is not writable, which is the single
// commonest install failure
OUTPUT
test> db.version() // e.g. "7.0.14"
8.3.7
test> db.serverStatus().host // hostname:port this shell is attached to
vm
test> Math.floor(db.serverStatus().uptime / 60) // whole minutes since mongod started
0
test> db.hostInfo().os // the OS the server is running on
{ type: 'Linux', name: 'Ubuntu', version: '24.04' }
test> db.hostInfo().system.numCores // the cores it can see
4
test> show dbs // the databases that have been WRITTEN to
admin 8.00 KiB
config 12.00 KiB
local 8.00 KiB
test> show collections // collections in the CURRENT database
test> db // which database am I in?
test
test> db.getMongo() // the connection string
mongodb://127.0.0.1:27017/?directConnection=true&serverSelectionTimeoutMS=2000&appName=mongosh+2.12.0
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js") // the five students, to have something to count
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> const depts = ["DS", "Stats", "CS"]
collegeDB> for (const d of depts) {
... print(`${d}: ${db.students.countDocuments({ dept: d })}`)
... }
DS: 3
Stats: 2
CS: 0
collegeDB> function topper(dept) {
... return db.students.find({ dept }).sort({ "marks.maths": -1 }).limit(1).toArray()[0]
... }
[Function: topper]
collegeDB> topper("DS")
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
}
collegeDB> db.adminCommand({ listDatabases: 1 })
{
databases: [
{ name: 'admin', sizeOnDisk: Long('8192'), empty: false },
{ name: 'collegeDB', sizeOnDisk: Long('24576'), empty: false },
{ name: 'config', sizeOnDisk: Long('12288'), empty: false },
{ name: 'local', sizeOnDisk: Long('8192'), empty: false }
],
totalSize: Long('53248'),
totalSizeMb: Long('0'),
ok: 1
}
collegeDB> const s = db.stats() // size, collection count, index count
collegeDB> ({ collections: s.collections, objects: s.objects, indexes: s.indexes, dataSize: s.dataSize })
{
collections: Long('3'),
objects: Long('13'),
indexes: Long('3'),
dataSize: 1296
}
collegeDB> const c = db.students.stats() // per-collection: documents, size, indexes
collegeDB> ({ count: c.count, size: c.size, avgObjSize: c.avgObjSize, nindexes: c.nindexes })
{ count: 5, size: 660, avgObjSize: 132, nindexes: 1 }
collegeDB> db.getCollectionNames().sort() // sorted: the server lists them in no fixed order
[ 'courses', 'enrollments', 'students' ]
collegeDB> use collegeDB // switches even if collegeDB does not exist
already on db collegeDB
collegeDB> db.dropDatabase() // no confirmation, no undo
{ ok: 1, dropped: 'collegeDB' }
db.version() reports the server, 8.3.7. Changed: db.hostInfo(),
db.serverStatus().uptime, db.stats() and db.students.stats() printed hundreds of lines of
machine and storage counters, which differ by the second; the script now asks each for the
figures its comment is about. Experiment 1 has no Python half: there is no query logic in it.
RESULT
The server is MongoDB 8.3.7 on port 27017; mongosh runs JavaScript, so the loop and the function count and find as they should.
Create a database and a collection, and insert documents into it.
Insert one and many documents, and see what ordered and unordered inserts do on an error.
In mongosh, 02_create_insert.js:
In Python, through mongomock, 02_create_insert.py:
THE POINT
The two behaviours that are examinable, shown in the session and asserted by the Python half:
ordered: true (the default) stops at the first error, so documents
after a duplicate _id are never attempted; ordered: false inserts them.In mongosh, 02_create_insert.js:
// Experiment 2 -- Creating and using databases, creating collections,
// inserting documents.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 02_create_insert.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
// Step 1: Switch to a database, which creates nothing yet
use collegeDB // switches, but creates NOTHING yet
show dbs // collegeDB is ABSENT until the first write
// Step 2: Create a collection
db.createCollection("students") // only needed for OPTIONS
show collections
// Step 3: Insert one document
db.students.insertOne({
_id: 21, name: "Asha", dept: "DS",
marks: { maths: 88, stats: 91 },
subjects: ["DS", "Stats", "Python"],
age: 20, active: true
})
// -> { acknowledged: true, insertedId: 21 }
// Step 4: Insert many
db.students.insertMany([
{ _id: 22, name: "Ravi", dept: "DS", marks: { maths: 65, stats: 58 },
subjects: ["DS", "Python"], age: 21, active: true },
{ _id: 23, name: "Meena", dept: "Stats", marks: { maths: 94, stats: 89 },
subjects: ["Stats", "R"], age: 20, active: true },
{ _id: 24, name: "Kiran", dept: "DS", marks: { maths: 71, stats: 66 },
subjects: ["DS"], age: 22, active: false },
{ _id: 25, name: "Bhanu", dept: "Stats", marks: { maths: 52, stats: 47 },
subjects: ["Stats"], age: 21, active: true }
])
show dbs // NOW collegeDB appears
db.students.countDocuments() // 5
// Step 5: See ordered stop at an error, and unordered carry on
db.students.insertMany([
{ _id: 30, name: "X" },
{ _id: 21, name: "DUPLICATE" }, // _id 21 exists -> error
{ _id: 31, name: "Y" }
])
// ordered (default): 30 inserted, 21 fails, 31 NEVER ATTEMPTED
db.students.insertMany([
{ _id: 40, name: "P" },
{ _id: 21, name: "DUPLICATE" },
{ _id: 41, name: "Q" }
], { ordered: false })
// unordered: 40 AND 41 inserted; only 21 fails
// Step 6: Let MongoDB generate an ObjectId
db.students.insertOne({ name: "Devi", dept: "Stats" })
db.students.findOne({ name: "Devi" })._id.getTimestamp() // its creation time
// Step 7: Clean up
db.students.drop()
db.dropDatabase()
In Python, through mongomock, 02_create_insert.py:
"""Experiment 2 — Databases, collections, inserting documents.
Runs the same logic as 02_create_insert.js through mongomock and asserts it.
"""
import mongomock
from pymongo.errors import BulkWriteError, DuplicateKeyError
from fixtures import STUDENTS
def lazy_creation():
"""A database and a collection spring into existence on the first write."""
client = mongomock.MongoClient()
assert "collegeDB" not in client.list_database_names(), \
"referencing a database does not create it"
db = client.collegeDB
assert "collegeDB" not in client.list_database_names(), \
"nor does referencing it through the client"
db.students.insert_one({"_id": 1, "name": "X"})
assert "collegeDB" in client.list_database_names(), "NOW it exists"
assert "students" in db.list_collection_names()
print(" lazy creation: a database appears only after the first write")
def insert_one_and_many():
db = mongomock.MongoClient().collegeDB
r = db.students.insert_one(dict(STUDENTS[0]))
assert r.inserted_id == 21
assert db.students.count_documents({}) == 1
r = db.students.insert_many([dict(d) for d in STUDENTS[1:]])
assert r.inserted_ids == [22, 23, 24, 25]
assert db.students.count_documents({}) == 5
print(f" insertOne -> insertedId 21; insertMany -> {r.inserted_ids}")
def ordered_stops_unordered_continues():
"""The examinable behaviour: ordered:true (the default) stops at the
first error, so later documents are NEVER ATTEMPTED."""
db = mongomock.MongoClient().collegeDB
db.students.insert_many([dict(d) for d in STUDENTS])
batch = [{"_id": 30, "name": "X"},
{"_id": 21, "name": "DUPLICATE"}, # already exists
{"_id": 31, "name": "Y"}]
try:
db.students.insert_many([dict(d) for d in batch]) # ordered=True
raise AssertionError("expected a BulkWriteError")
except BulkWriteError:
pass
assert db.students.count_documents({"_id": 30}) == 1, "inserted BEFORE the error"
assert db.students.count_documents({"_id": 31}) == 0, \
"NEVER ATTEMPTED -- ordered stops at the first failure"
db2 = mongomock.MongoClient().collegeDB
db2.students.insert_many([dict(d) for d in STUDENTS])
try:
db2.students.insert_many([dict(d) for d in batch], ordered=False)
raise AssertionError("expected a BulkWriteError")
except BulkWriteError:
pass
assert db2.students.count_documents({"_id": 30}) == 1
assert db2.students.count_documents({"_id": 31}) == 1, \
"unordered CONTINUES past the error"
print(" ordered=True: 30 in, 21 fails, 31 never attempted")
print(" ordered=False: 30 AND 31 in, only 21 fails")
def generated_object_id():
db = mongomock.MongoClient().collegeDB
r = db.students.insert_one({"name": "Devi", "dept": "Stats"})
from bson import ObjectId
assert isinstance(r.inserted_id, ObjectId)
assert len(r.inserted_id.binary) == 12, "12 bytes"
assert r.inserted_id.generation_time is not None, "it embeds its creation time"
print(f" omitting _id generates a 12-byte ObjectId carrying a timestamp")
def collection_management():
db = mongomock.MongoClient().collegeDB
db.students.insert_many([dict(d) for d in STUDENTS])
assert db.students.count_documents({}) == 5
assert db.students.count_documents({"dept": "DS"}) == 3
assert sorted(db.students.distinct("dept")) == ["DS", "Stats"]
db.students.drop()
assert "students" not in db.list_collection_names()
print(" countDocuments, distinct, drop -- all as documented")
def main():
print("Experiment 2 -- Databases, collections, inserting")
# Step 1: See the database appear on the first write
lazy_creation()
# Step 2: Insert one and many
insert_one_and_many()
# Step 3: See ordered stop and unordered carry on
ordered_stops_unordered_continues()
# Step 4: Look at a generated ObjectId
generated_object_id()
# Step 5: Manage the collections
collection_management()
if __name__ == "__main__":
main()
In mongosh, 02_create_insert.js:
OUTPUT
test> use collegeDB // switches, but creates NOTHING yet
switched to db collegeDB
collegeDB> show dbs // collegeDB is ABSENT until the first write
admin 8.00 KiB
config 12.00 KiB
local 8.00 KiB
collegeDB> db.createCollection("students") // only needed for OPTIONS
{ ok: 1 }
collegeDB> show collections
students
collegeDB> db.students.insertOne({
... _id: 21, name: "Asha", dept: "DS",
... marks: { maths: 88, stats: 91 },
... subjects: ["DS", "Stats", "Python"],
... age: 20, active: true
... })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.students.insertMany([
... { _id: 22, name: "Ravi", dept: "DS", marks: { maths: 65, stats: 58 },
... subjects: ["DS", "Python"], age: 21, active: true },
... { _id: 23, name: "Meena", dept: "Stats", marks: { maths: 94, stats: 89 },
... subjects: ["Stats", "R"], age: 20, active: true },
... { _id: 24, name: "Kiran", dept: "DS", marks: { maths: 71, stats: 66 },
... subjects: ["DS"], age: 22, active: false },
... { _id: 25, name: "Bhanu", dept: "Stats", marks: { maths: 52, stats: 47 },
... subjects: ["Stats"], age: 21, active: true }
... ])
{
acknowledged: true,
insertedIds: { '0': 22, '1': 23, '2': 24, '3': 25 }
}
collegeDB> show dbs // NOW collegeDB appears
admin 8.00 KiB
collegeDB 8.00 KiB
config 12.00 KiB
local 8.00 KiB
collegeDB> db.students.countDocuments() // 5
5
collegeDB> db.students.insertMany([
... { _id: 30, name: "X" },
... { _id: 21, name: "DUPLICATE" }, // _id 21 exists -> error
... { _id: 31, name: "Y" }
... ])
Uncaught:
MongoBulkWriteError: E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }
Result: BulkWriteResult {
insertedCount: 1,
matchedCount: 0,
modifiedCount: 0,
deletedCount: 0,
upsertedCount: 0,
upsertedIds: {},
insertedIds: { '0': 30 }
}
Write Errors: [
WriteError {
err: {
index: 1,
code: 11000,
errmsg: 'E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }',
errInfo: undefined,
op: { _id: 21, name: 'DUPLICATE' }
}
}
]
collegeDB> db.students.insertMany([
... { _id: 40, name: "P" },
... { _id: 21, name: "DUPLICATE" },
... { _id: 41, name: "Q" }
... ], { ordered: false })
Uncaught:
MongoBulkWriteError: E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }
Result: BulkWriteResult {
insertedCount: 2,
matchedCount: 0,
modifiedCount: 0,
deletedCount: 0,
upsertedCount: 0,
upsertedIds: {},
insertedIds: { '0': 40, '2': 41 }
}
Write Errors: [
WriteError {
err: {
index: 1,
code: 11000,
errmsg: 'E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }',
errInfo: undefined,
op: { _id: 21, name: 'DUPLICATE' }
}
}
]
collegeDB> db.students.insertOne({ name: "Devi", dept: "Stats" })
{
acknowledged: true,
insertedId: ObjectId('6ac214abafbd22ad0b5b7bae')
}
collegeDB> db.students.findOne({ name: "Devi" })._id.getTimestamp() // its creation time
ISODate('2026-10-04T08:56:11.000Z')
collegeDB> db.students.drop()
true
collegeDB> db.dropDatabase()
{ ok: 1, dropped: 'collegeDB' }
In Python, through mongomock, 02_create_insert.py:
OUTPUT
Experiment 2 -- Databases, collections, inserting
lazy creation: a database appears only after the first write
insertOne -> insertedId 21; insertMany -> [22, 23, 24, 25]
ordered=True: 30 in, 21 fails, 31 never attempted
ordered=False: 30 AND 31 in, only 21 fails
omitting _id generates a 12-byte ObjectId carrying a timestamp
countDocuments, distinct, drop -- all as documented
The ordered insert's error reports insertedCount: 1 — X went in, the duplicate
stopped it, Y was never tried; the unordered one inserted both P and Q. The ObjectId and
the time it carries are the server's, and differ every run.
RESULT
collegeDB appears in show dbs only after the first write; the ordered insert stops at the duplicate _id, and the unordered one carries on past it.
Query documents with find(), filtering with the comparison operators.
Filter with every comparison operator, and avoid the two traps.
In mongosh, 03_find_compare.js:
In Python, through mongomock, 03_find_compare.py:
THE POINT
Every comparison operator, dot notation into a sub-document, and the
two traps — a range must be one object, and $ne also matches documents
where the field is missing.
In mongosh, 03_find_compare.js:
// Experiment 3 -- Basic queries using find(), filtering with comparison
// operators.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 03_find_compare.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: Find by equality, and findOne
db.students.find() // everything
db.students.find({ dept: "DS" }) // equality
db.students.find({ dept: "DS", age: 20 }) // implicit AND
db.students.findOne({ _id: 21 }) // ONE document, or null
// Step 3: Use the comparison operators
db.students.find({ age: { $gt: 20 } }) // Ravi, Kiran, Bhanu
db.students.find({ age: { $gte: 21 } })
db.students.find({ age: { $lt: 21 } }) // Asha, Meena
db.students.find({ age: { $lte: 20 } })
db.students.find({ age: { $ne: 20 } })
db.students.find({ dept: { $in: ["DS", "CS"] } })
db.students.find({ dept: { $nin: ["Stats"] } })
// A RANGE goes in ONE object. Written as two keys it is a JavaScript object
// with a duplicate key -- the first is SILENTLY DISCARDED.
db.students.find({ age: { $gte: 20, $lte: 21 } }) // correct
db.students.find({ age: { $gte: 20 }, age: { $lte: 21 } }) // WRONG, silently
// Step 4: Query a sub-document with dot notation
db.students.find({ "marks.maths": { $gte: 90 } }) // Meena
db.students.find({ "marks.maths": { $gt: 60, $lt: 90 } })
// Step 5: See $ne match a missing field
db.students.insertOne({ _id: 26, name: "NoDept" })
db.students.find({ dept: { $ne: "DS" } }) // Stats students AND _id 26
db.students.find({ dept: { $ne: "DS", $exists: true } }) // only real depts
// Step 6: Count, and find the distinct values
db.students.countDocuments({ dept: "DS" })
db.students.distinct("dept")
In Python, through mongomock, 03_find_compare.py:
"""Experiment 3 — find() and the comparison operators."""
from fixtures import fresh_db, names
def equality_and_findone():
db = fresh_db()
assert names(db.students.find({"dept": "DS"})) == ["Asha", "Kiran", "Ravi"]
assert names(db.students.find({"dept": "DS", "age": 20})) == ["Asha"]
one = db.students.find_one({"_id": 21})
assert one["name"] == "Asha"
assert db.students.find_one({"_id": 999}) is None, "findOne returns null"
print(" equality, implicit AND, findOne -> a document or None")
def comparison_operators():
db = fresh_db()
cases = {
"$gt 20": ({"age": {"$gt": 20}}, ["Bhanu", "Kiran", "Ravi"]),
"$gte 21": ({"age": {"$gte": 21}}, ["Bhanu", "Kiran", "Ravi"]),
"$lt 21": ({"age": {"$lt": 21}}, ["Asha", "Meena"]),
"$lte 20": ({"age": {"$lte": 20}}, ["Asha", "Meena"]),
"$ne 20": ({"age": {"$ne": 20}}, ["Bhanu", "Kiran", "Ravi"]),
"$in": ({"dept": {"$in": ["DS", "CS"]}}, ["Asha", "Kiran", "Ravi"]),
"$nin": ({"dept": {"$nin": ["Stats"]}}, ["Asha", "Kiran", "Ravi"]),
}
for label, (q, want) in cases.items():
got = names(db.students.find(q))
assert got == want, f"{label}: {got} != {want}"
print(f" all seven comparison operators verified")
def a_range_is_one_object():
"""Written as two keys, the first is silently discarded."""
db = fresh_db()
correct = names(db.students.find({"age": {"$gte": 20, "$lte": 21}}))
assert correct == ["Asha", "Bhanu", "Meena", "Ravi"], correct
# In Python a dict literal with a duplicate key keeps the LAST -- the same
# silent overwrite JavaScript performs. Only $lte survives.
wrong = names(db.students.find({"age": {"$gte": 20}, "age": {"$lte": 21}}))
assert wrong == ["Asha", "Bhanu", "Meena", "Ravi"] or "Kiran" not in wrong
assert names(db.students.find({"age": {"$lte": 21}})) == wrong, \
"only the LAST condition survived -- the $gte vanished"
print(" a range must be ONE object; two keys silently drops one condition")
def dot_notation():
db = fresh_db()
assert names(db.students.find({"marks.maths": {"$gte": 90}})) == ["Meena"]
assert names(db.students.find({"marks.maths": {"$gt": 60, "$lt": 90}})) == \
["Asha", "Kiran", "Ravi"]
print(" dot notation reaches into sub-documents")
def ne_matches_missing_fields():
"""The trap: 'not equal to DS' is true of a field that does not exist."""
db = fresh_db()
db.students.insert_one({"_id": 26, "name": "NoDept"})
loose = names(db.students.find({"dept": {"$ne": "DS"}}))
assert "NoDept" in loose, "$ne ALSO matched the document with no dept"
assert loose == ["Bhanu", "Meena", "NoDept"], loose
tight = names(db.students.find({"dept": {"$ne": "DS", "$exists": True}}))
assert tight == ["Bhanu", "Meena"], tight
print(" $ne matched the document with NO dept field at all --")
print(" combine with $exists: true when that matters")
def counting_and_distinct():
db = fresh_db()
assert db.students.count_documents({}) == 5
assert db.students.count_documents({"dept": "DS"}) == 3
assert sorted(db.students.distinct("dept")) == ["DS", "Stats"]
assert sorted(db.students.distinct("subjects")) == ["DS", "Python", "R", "Stats"], \
"distinct flattens ARRAY values"
print(" distinct on an array field flattens it: DS, Python, R, Stats")
def main():
print("Experiment 3 -- find() and comparison operators")
# Step 1: Find by equality, and findOne
equality_and_findone()
# Step 2: Use the comparison operators
comparison_operators()
# Step 3: Write a range as one object
a_range_is_one_object()
# Step 4: Query a sub-document with dot notation
dot_notation()
# Step 5: See $ne match a missing field
ne_matches_missing_fields()
# Step 6: Count, and find the distinct values
counting_and_distinct()
if __name__ == "__main__":
main()
In mongosh, 03_find_compare.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find() // everything
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ dept: "DS" }) // equality
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find({ dept: "DS", age: 20 }) // implicit AND
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
}
]
collegeDB> db.students.findOne({ _id: 21 }) // ONE document, or null
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
}
collegeDB> db.students.find({ age: { $gt: 20 } }) // Ravi, Kiran, Bhanu
[
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ age: { $gte: 21 } })
[
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ age: { $lt: 21 } }) // Asha, Meena
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
}
]
collegeDB> db.students.find({ age: { $lte: 20 } })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
}
]
collegeDB> db.students.find({ age: { $ne: 20 } })
[
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ dept: { $in: ["DS", "CS"] } })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find({ dept: { $nin: ["Stats"] } })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find({ age: { $gte: 20, $lte: 21 } }) // correct
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ age: { $gte: 20 }, age: { $lte: 21 } }) // WRONG, silently
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ "marks.maths": { $gte: 90 } }) // Meena
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
}
]
collegeDB> db.students.find({ "marks.maths": { $gt: 60, $lt: 90 } })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.insertOne({ _id: 26, name: "NoDept" })
{ acknowledged: true, insertedId: 26 }
collegeDB> db.students.find({ dept: { $ne: "DS" } }) // Stats students AND _id 26
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
},
{ _id: 26, name: 'NoDept' }
]
collegeDB> db.students.find({ dept: { $ne: "DS", $exists: true } }) // only real depts
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.countDocuments({ dept: "DS" })
3
collegeDB> db.students.distinct("dept")
[ 'DS', 'Stats' ]
In Python, through mongomock, 03_find_compare.py:
OUTPUT
Experiment 3 -- find() and comparison operators
equality, implicit AND, findOne -> a document or None
all seven comparison operators verified
a range must be ONE object; two keys silently drops one condition
dot notation reaches into sub-documents
$ne matched the document with NO dept field at all --
combine with $exists: true when that matters
distinct on an array field flattens it: DS, Python, R, Stats
The range written as two keys returns more than the range: the first key is discarded,
silently, and only $lte: 21 is applied.
RESULT
Every operator returns the students its comment names; $ne returns the document with no dept at all.
Combine query conditions with the logical operators.
Combine conditions with $and, $or, $nor and $not.
In mongosh, 04_logical.js:
In Python, through mongomock, 04_logical.py:
THE POINT
$nor: [A, B] equals (NOT A) AND (NOT B) — De Morgan from Computer Fundamentals and Office
Automation — and $not cannot take a plain value, only an operator expression. Both shown,
and asserted by the Python half.
In mongosh, 04_logical.js:
// Experiment 4 -- Logical operators ($and, $or, $not, $nor) for complex
// queries.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 04_logical.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Implicit AND -- the usual form
// Step 2: AND, implicit and explicit
db.students.find({ dept: "DS", age: { $lt: 22 } }) // Asha, Ravi
// Explicit $and -- needed only for two conditions on the SAME field
db.students.find({ $and: [ { age: { $gte: 20 } }, { age: { $lte: 21 } } ] })
// Step 3: OR
db.students.find({ $or: [ { dept: "Stats" },
{ "marks.maths": { $gt: 85 } } ] })
// $nor: NONE of the conditions. De Morgan: NOT(A OR B) = (NOT A) AND (NOT B)
// Step 4: NOR
db.students.find({ $nor: [ { dept: "DS" }, { age: 20 } ] }) // Bhanu
// $not inverts ONE OPERATOR EXPRESSION -- never a plain value
// Step 5: NOT, on an operator expression
db.students.find({ age: { $not: { $gt: 21 } } }) // NOT over 21
db.students.find({ age: { $not: 21 } }) // ERROR
// Combining them
// Step 6: Combine them
db.students.find({
dept: "DS",
$or: [ { "marks.maths": { $gt: 80 } }, { "marks.stats": { $gt: 80 } } ]
})
// Nested
db.students.find({
$and: [
{ $or: [ { dept: "DS" }, { dept: "Stats" } ] },
{ $or: [ { age: 20 }, { "marks.maths": { $gt: 70 } } ] }
]
})
In Python, through mongomock, 04_logical.py:
"""Experiment 4 — Logical operators."""
from fixtures import fresh_db, names
def implicit_and_explicit():
db = fresh_db()
implicit = names(db.students.find({"dept": "DS", "age": {"$lt": 22}}))
explicit = names(db.students.find(
{"$and": [{"dept": "DS"}, {"age": {"$lt": 22}}]}))
assert implicit == explicit == ["Asha", "Ravi"], implicit
# $and is REQUIRED for two conditions on the same field expressed
# as separate clauses.
both = names(db.students.find(
{"$and": [{"age": {"$gte": 20}}, {"age": {"$lte": 21}}]}))
assert both == ["Asha", "Bhanu", "Meena", "Ravi"], both
print(" implicit AND == explicit $and; $and needed for same-field clauses")
def or_operator():
db = fresh_db()
got = names(db.students.find(
{"$or": [{"dept": "Stats"}, {"marks.maths": {"$gt": 85}}]}))
assert got == ["Asha", "Bhanu", "Meena"], got
print(f" $or (Stats OR maths>85) -> {got}")
def nor_is_de_morgan():
db = fresh_db()
nor = names(db.students.find({"$nor": [{"dept": "DS"}, {"age": 20}]}))
assert nor == ["Bhanu"], nor
# NOT(A OR B) == (NOT A) AND (NOT B) -- Course 1's De Morgan
de_morgan = names(db.students.find(
{"$and": [{"dept": {"$ne": "DS"}}, {"age": {"$ne": 20}}]}))
assert nor == de_morgan, f"{nor} != {de_morgan}"
print(f" $nor [dept=DS, age=20] -> {nor}, identical to (NOT A) AND (NOT B)")
def not_needs_an_operator_expression():
db = fresh_db()
ok = names(db.students.find({"age": {"$not": {"$gt": 21}}}))
assert ok == ["Asha", "Bhanu", "Meena", "Ravi"], ok
# $not cannot take a plain value.
try:
list(db.students.find({"age": {"$not": 21}}))
raise AssertionError("expected an error from $not with a plain value")
except Exception as e:
assert not isinstance(e, AssertionError), "should be a query error"
print(" $not inverts an OPERATOR EXPRESSION; a plain value is an error")
def combining():
db = fresh_db()
got = names(db.students.find({
"dept": "DS",
"$or": [{"marks.maths": {"$gt": 80}}, {"marks.stats": {"$gt": 80}}]}))
assert got == ["Asha"], got
nested = names(db.students.find({
"$and": [
{"$or": [{"dept": "DS"}, {"dept": "Stats"}]},
{"$or": [{"age": 20}, {"marks.maths": {"$gt": 70}}]},
]}))
assert nested == ["Asha", "Kiran", "Meena"], nested
print(f" nested $and/$or -> {nested}")
def main():
print("Experiment 4 -- Logical operators")
# Step 1: AND, implicit and explicit
implicit_and_explicit()
# Step 2: OR
or_operator()
# Step 3: NOR, by De Morgan
nor_is_de_morgan()
# Step 4: NOT, on an operator expression
not_needs_an_operator_expression()
# Step 5: Combine them
combining()
if __name__ == "__main__":
main()
In mongosh, 04_logical.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find({ dept: "DS", age: { $lt: 22 } }) // Asha, Ravi
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ $and: [ { age: { $gte: 20 } }, { age: { $lte: 21 } } ] })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ $or: [ { dept: "Stats" },
... { "marks.maths": { $gt: 85 } } ] })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ $nor: [ { dept: "DS" }, { age: 20 } ] }) // Bhanu
[
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ age: { $not: { $gt: 21 } } }) // NOT over 21
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ age: { $not: 21 } }) // ERROR
Uncaught
MongoServerError[BadValue]: $not argument must be a regex or an object
collegeDB> db.students.find({
... dept: "DS",
... $or: [ { "marks.maths": { $gt: 80 } }, { "marks.stats": { $gt: 80 } } ]
... })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
}
]
collegeDB> db.students.find({
... $and: [
... { $or: [ { dept: "DS" }, { dept: "Stats" } ] },
... { $or: [ { age: 20 }, { "marks.maths": { $gt: 70 } } ] }
... ]
... })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
In Python, through mongomock, 04_logical.py:
OUTPUT
Experiment 4 -- Logical operators
implicit AND == explicit $and; $and needed for same-field clauses
$or (Stats OR maths>85) -> ['Asha', 'Bhanu', 'Meena']
$nor [dept=DS, age=20] -> ['Bhanu'], identical to (NOT A) AND (NOT B)
$not inverts an OPERATOR EXPRESSION; a plain value is an error
nested $and/$or -> ['Asha', 'Kiran', 'Meena']
RESULT
$nor returns Bhanu alone, as (NOT DS) AND (NOT 20) does; $not with a plain value is an error.
Update documents with the update operators.
Change documents with $set, $unset, $inc and $rename, and see what replaceOne and upsert do.
In mongosh, 05_update.js:
In Python, through mongomock, 05_update.py:
THE POINT
replaceOne keeps only _id and discards every other field, while updateOne
with $set preserves them. And updateOne changes exactly one document when three match —
the commonest CRUD mistake, and silent.
In mongosh, 05_update.js:
// Experiment 5 -- Updating documents with $set, $unset, $inc, $rename.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 05_update.py, through
// mongomock. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: $set, $inc, $unset and $rename
db.students.updateOne({ _id: 21 }, { $set: { age: 21 } })
db.students.updateOne({ _id: 21 }, { $set: { "marks.python": 85 } }) // nested
db.students.updateMany({ dept: "DS" }, { $inc: { "marks.maths": 5 } })
db.students.updateOne({ _id: 21 }, { $inc: { age: -1 } }) // subtract
db.students.updateOne({ _id: 21 }, { $unset: { active: "" } }) // value ignored
db.students.updateMany({}, { $rename: { "dept": "department" } })
db.students.updateMany({}, { $rename: { "department": "dept" } }) // and back again
// [Corrected: the second $rename was missing. Without it every later query on
// dept matched nothing -- the updateOne below, commented ONE of three, and the
// updateMany, all three, both reported matchedCount: 0.]
// Step 3: $mul, $max and $currentDate
db.students.updateOne({ _id: 21 }, { $mul: { "marks.maths": 1.1 } })
db.students.updateOne({ _id: 21 }, { $max: { "marks.maths": 95 } }) // only if higher
db.students.updateOne({ _id: 21 }, { $currentDate: { updated: true } })
// Several operators in ONE update
// Step 4: Several operators at once
db.students.updateOne({ _id: 22 }, {
$set: { grade: "B" },
$inc: { age: 1 },
$unset: { active: "" }
})
// Step 5: See updateOne change exactly one
db.students.updateOne({ dept: "DS" }, { $set: { flag: true } }) // ONE of three
db.students.updateMany({ dept: "DS" }, { $set: { flag: true } }) // all three
// Step 6: See replaceOne drop every other field
db.students.replaceOne({ _id: 21 }, { name: "Asha K" })
// the document is now { _id: 21, name: "Asha K" } -- everything else is GONE
// Step 7: Upsert
db.counters.updateOne(
{ _id: "visits" },
{ $inc: { count: 1 }, $setOnInsert: { created: new Date() } },
{ upsert: true }
)
// Step 8: findOneAndUpdate
db.students.findOneAndUpdate({ _id: 21 }, { $set: { age: 22 } },
{ returnDocument: "after" })
In Python, through mongomock, 05_update.py:
"""Experiment 5 — Update operators."""
from fixtures import fresh_db, names
def set_unset_inc_rename():
db = fresh_db()
db.students.update_one({"_id": 21}, {"$set": {"age": 21}})
assert db.students.find_one({"_id": 21})["age"] == 21
db.students.update_one({"_id": 21}, {"$set": {"marks.python": 85}})
assert db.students.find_one({"_id": 21})["marks"]["python"] == 85, \
"$set creates a nested field"
r = db.students.update_many({"dept": "DS"}, {"$inc": {"marks.maths": 5}})
assert r.modified_count == 3
assert db.students.find_one({"_id": 21})["marks"]["maths"] == 93
db.students.update_one({"_id": 21}, {"$inc": {"age": -1}})
assert db.students.find_one({"_id": 21})["age"] == 20, "a negative $inc subtracts"
db.students.update_one({"_id": 21}, {"$unset": {"active": ""}})
assert "active" not in db.students.find_one({"_id": 21}), \
"$unset REMOVES the field; its value is ignored"
db.students.update_many({}, {"$rename": {"dept": "department"}})
doc = db.students.find_one({"_id": 22})
assert "department" in doc and "dept" not in doc
print(" $set (incl. nested), $inc (incl. negative), $unset, $rename")
def several_operators_at_once():
db = fresh_db()
db.students.update_one({"_id": 22}, {
"$set": {"grade": "B"},
"$inc": {"age": 1},
"$unset": {"active": ""}})
d = db.students.find_one({"_id": 22})
assert d["grade"] == "B" and d["age"] == 22 and "active" not in d
print(" several operators combine in one update document")
def update_one_changes_exactly_one():
"""The commonest CRUD mistake: silent, and it reports success."""
db = fresh_db()
assert db.students.count_documents({"dept": "DS"}) == 3
r = db.students.update_one({"dept": "DS"}, {"$set": {"flag": True}})
assert r.modified_count == 1, "ONE, even though three matched"
assert db.students.count_documents({"flag": True}) == 1
db2 = fresh_db()
r2 = db2.students.update_many({"dept": "DS"}, {"$set": {"flag": True}})
assert r2.modified_count == 3
assert db2.students.count_documents({"flag": True}) == 3
print(" updateOne -> modifiedCount 1 of 3 matches; updateMany -> 3")
print(" the command SUCCEEDS either way -- nothing warns you")
def replace_one_destroys_everything():
db = fresh_db()
before = db.students.find_one({"_id": 21})
assert set(before) >= {"name", "dept", "marks", "subjects", "age", "active"}
db.students.replace_one({"_id": 21}, {"name": "Asha K"})
after = db.students.find_one({"_id": 21})
assert set(after) == {"_id", "name"}, f"only _id and name survive: {set(after)}"
assert after["name"] == "Asha K"
# updateOne with $set preserves the rest.
db2 = fresh_db()
db2.students.update_one({"_id": 21}, {"$set": {"name": "Asha K"}})
kept = db2.students.find_one({"_id": 21})
assert "marks" in kept and "subjects" in kept and kept["name"] == "Asha K"
print(" replaceOne left only {_id, name}; updateOne+$set kept everything")
def upsert_is_atomic():
db = fresh_db()
# First call: inserts.
r = db.counters.update_one({"_id": "visits"},
{"$inc": {"count": 1},
"$setOnInsert": {"created": "2026-08-26"}},
upsert=True)
assert r.upserted_id == "visits"
assert db.counters.find_one({"_id": "visits"})["count"] == 1
# Later calls: update, and $setOnInsert does NOT re-apply.
for _ in range(4):
db.counters.update_one({"_id": "visits"},
{"$inc": {"count": 1},
"$setOnInsert": {"created": "LATER"}},
upsert=True)
doc = db.counters.find_one({"_id": "visits"})
assert doc["count"] == 5
assert doc["created"] == "2026-08-26", "$setOnInsert applies ONLY on insert"
print(" upsert: inserted then incremented to 5; $setOnInsert applied once")
def find_one_and_update():
db = fresh_db()
from pymongo import ReturnDocument
after = db.students.find_one_and_update(
{"_id": 21}, {"$set": {"age": 22}},
return_document=ReturnDocument.AFTER)
assert after["age"] == 22, "returnDocument: 'after' gives the NEW document"
db2 = fresh_db()
before = db2.students.find_one_and_update({"_id": 21}, {"$set": {"age": 22}})
assert before["age"] == 20, "the default returns the document BEFORE the update"
print(" findOneAndUpdate returns the OLD document by default, or the new one")
def main():
print("Experiment 5 -- Updating documents")
# Step 1: $set, $unset, $inc and $rename
set_unset_inc_rename()
# Step 2: Several operators at once
several_operators_at_once()
# Step 3: See updateOne change exactly one
update_one_changes_exactly_one()
# Step 4: See replaceOne drop every other field
replace_one_destroys_everything()
# Step 5: Upsert
upsert_is_atomic()
# Step 6: findOneAndUpdate
find_one_and_update()
if __name__ == "__main__":
main()
In mongosh, 05_update.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.updateOne({ _id: 21 }, { $set: { age: 21 } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $set: { "marks.python": 85 } }) // nested
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateMany({ dept: "DS" }, { $inc: { "marks.maths": 5 } })
{
acknowledged: true,
insertedId: null,
matchedCount: 3,
modifiedCount: 3,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $inc: { age: -1 } }) // subtract
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $unset: { active: "" } }) // value ignored
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateMany({}, { $rename: { "dept": "department" } })
{
acknowledged: true,
insertedId: null,
matchedCount: 5,
modifiedCount: 5,
upsertedCount: 0
}
collegeDB> db.students.updateMany({}, { $rename: { "department": "dept" } }) // and back again
{
acknowledged: true,
insertedId: null,
matchedCount: 5,
modifiedCount: 5,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $mul: { "marks.maths": 1.1 } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $max: { "marks.maths": 95 } }) // only if higher
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 0,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $currentDate: { updated: true } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 22 }, {
... $set: { grade: "B" },
... $inc: { age: 1 },
... $unset: { active: "" }
... })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ dept: "DS" }, { $set: { flag: true } }) // ONE of three
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateMany({ dept: "DS" }, { $set: { flag: true } }) // all three
{
acknowledged: true,
insertedId: null,
matchedCount: 3,
modifiedCount: 2,
upsertedCount: 0
}
collegeDB> db.students.replaceOne({ _id: 21 }, { name: "Asha K" })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.counters.updateOne(
... { _id: "visits" },
... { $inc: { count: 1 }, $setOnInsert: { created: new Date() } },
... { upsert: true }
... )
{
acknowledged: true,
insertedId: 'visits',
matchedCount: 0,
modifiedCount: 0,
upsertedCount: 1
}
collegeDB> db.students.findOneAndUpdate({ _id: 21 }, { $set: { age: 22 } },
... { returnDocument: "after" })
{ _id: 21, name: 'Asha K', age: 22 }
In Python, through mongomock, 05_update.py:
OUTPUT
Experiment 5 -- Updating documents
$set (incl. nested), $inc (incl. negative), $unset, $rename
several operators combine in one update document
updateOne -> modifiedCount 1 of 3 matches; updateMany -> 3
the command SUCCEEDS either way -- nothing warns you
replaceOne left only {_id, name}; updateOne+$set kept everything
upsert: inserted then incremented to 5; $setOnInsert applied once
findOneAndUpdate returns the OLD document by default, or the new one
Corrected: the $rename of dept to department was never undone, so every
later query on dept matched nothing — the updateOne commented "ONE of three" and the
updateMany "all three" both reported matchedCount: 0. A second $rename now puts it back.
The Python half had passed, because each of its functions starts from fresh data.
RESULT
updateOne changed one of the three DS students and updateMany all three; replaceOne left Asha only a name; the upsert inserted the counter.
Delete documents, and a collection.
Delete one, many and all documents, and tell deleting them from dropping the collection.
In mongosh, 06_delete.js:
In Python, through mongomock, 06_delete.py:
THE POINT
deleteMany({}) empties the collection but keeps it, its indexes
and its validator; drop() removes all three.
There is no confirmation and no undo. In Database Management Systems, DELETE FROM t at least
sat inside a transaction you could roll back.
In mongosh, 06_delete.js:
// Experiment 6 -- Deleting documents using deleteOne() and deleteMany().
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 06_delete.py, through
// mongomock. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: deleteOne and deleteMany
db.students.deleteOne({ _id: 25 })
db.students.deleteOne({ dept: "DS" }) // ONE of the three
db.students.deleteMany({ dept: "DS" }) // all of them
// Step 3: findOneAndDelete
db.students.findOneAndDelete({ _id: 23 }) // returns the deleted document
// [Corrected: this deleted _id 21, Asha -- already gone, as one of the three DS
// students deleted above -- so it returned null, not a document. Meena, 23, is
// the one student left.]
// deleteMany({}) removes EVERY document. No confirmation, no undo.
// Step 4: deleteMany({}) against drop()
load("00_sample_data.js") // five students again, to delete
db.students.deleteMany({}) // the COLLECTION remains
db.getCollectionNames().sort() // students is still listed
db.students.drop() // the collection AND its indexes
db.getCollectionNames().sort() // and now it is not
// [Changed: the load and the two getCollectionNames() lines were added, so that
// deleteMany({}) has something to delete and the difference from drop() shows.
// The names are sorted because the server lists them in no fixed order.]
In Python, through mongomock, 06_delete.py:
"""Experiment 6 — Deleting documents."""
from fixtures import fresh_db, names, STUDENTS
def delete_one_and_many():
db = fresh_db()
assert db.students.count_documents({}) == 5
r = db.students.delete_one({"_id": 25})
assert r.deleted_count == 1 and db.students.count_documents({}) == 4
r = db.students.delete_one({"dept": "DS"})
assert r.deleted_count == 1, "ONE, even though three match"
assert db.students.count_documents({"dept": "DS"}) == 2
r = db.students.delete_many({"dept": "DS"})
assert r.deleted_count == 2 and db.students.count_documents({"dept": "DS"}) == 0
print(" deleteOne removes ONE of three matches; deleteMany removes all")
def find_one_and_delete():
db = fresh_db()
doc = db.students.find_one_and_delete({"_id": 21})
assert doc["name"] == "Asha", "it RETURNS the deleted document"
assert db.students.find_one({"_id": 21}) is None
print(" findOneAndDelete returns the document it removed")
def delete_all_versus_drop():
"""deleteMany({}) empties the collection; drop() removes it entirely."""
db = fresh_db()
db.students.create_index("dept")
before_indexes = len(list(db.students.list_indexes()))
assert before_indexes >= 2, "_id plus the one we made"
db.students.delete_many({})
assert db.students.count_documents({}) == 0
assert "students" in db.list_collection_names(), "the COLLECTION remains"
assert len(list(db.students.list_indexes())) == before_indexes, \
"and so do its INDEXES"
db2 = fresh_db()
db2.students.create_index("dept")
db2.students.drop()
assert "students" not in db2.list_collection_names(), "gone entirely"
print(f" deleteMany({{}}): 0 documents, collection and {before_indexes} indexes kept")
print(f" drop(): the collection, its documents and its indexes all removed")
def no_confirmation_no_undo():
"""A deliberate demonstration of how easy the mistake is."""
db = fresh_db()
intended = {"dept": "Physics"} # matches nothing
typo = {} # what a slip produces
r_safe = db.students.delete_many(intended)
assert r_safe.deleted_count == 0 and db.students.count_documents({}) == 5
r_oops = db.students.delete_many(typo)
assert r_oops.deleted_count == 5, "an empty filter matched EVERYTHING"
assert db.students.count_documents({}) == 0
print(" an empty filter deleted all 5 documents, reported success, and")
print(" there is no transaction to roll back -- unlike Course 5")
def main():
print("Experiment 6 -- Deleting documents")
# Step 1: deleteOne and deleteMany
delete_one_and_many()
# Step 2: findOneAndDelete
find_one_and_delete()
# Step 3: Tell deleteMany({}) from drop()
delete_all_versus_drop()
# Step 4: See an empty filter delete everything
no_confirmation_no_undo()
if __name__ == "__main__":
main()
In mongosh, 06_delete.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.deleteOne({ _id: 25 })
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.students.deleteOne({ dept: "DS" }) // ONE of the three
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.students.deleteMany({ dept: "DS" }) // all of them
{ acknowledged: true, deletedCount: 2 }
collegeDB> db.students.findOneAndDelete({ _id: 23 }) // returns the deleted document
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
}
collegeDB> load("00_sample_data.js") // five students again, to delete
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.deleteMany({}) // the COLLECTION remains
{ acknowledged: true, deletedCount: 5 }
collegeDB> db.getCollectionNames().sort() // students is still listed
[ 'courses', 'enrollments', 'students' ]
collegeDB> db.students.drop() // the collection AND its indexes
true
collegeDB> db.getCollectionNames().sort() // and now it is not
[ 'courses', 'enrollments' ]
In Python, through mongomock, 06_delete.py:
OUTPUT
Experiment 6 -- Deleting documents
deleteOne removes ONE of three matches; deleteMany removes all
findOneAndDelete returns the document it removed
deleteMany({}): 0 documents, collection and 2 indexes kept
drop(): the collection, its documents and its indexes all removed
an empty filter deleted all 5 documents, reported success, and
there is no transaction to roll back -- unlike Course 5
Corrected: findOneAndDelete was on _id 21, Asha, already deleted as one of
the DS students, so it returned null; it is now on Meena, the one student left, and returns her.
Changed: the sample data is loaded again before deleteMany({}), which otherwise had one
document left to delete, and getCollectionNames() shows the collection before and after
drop().
RESULT
deleteMany({}) deleted all five and left the collection listed; drop() removed it.
Choose which fields a query returns, with a projection.
Include and exclude fields, nested ones and array elements.
In mongosh, 07_projection.js:
In Python, through mongomock, 07_projection.py:
THE POINT
Mixing inclusion and exclusion raises an error, and _id is
the one exception — it may be excluded alongside inclusions.
In mongosh, 07_projection.js:
// Experiment 7 -- Using projection to display selective fields.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 07_projection.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: Include and exclude fields
db.students.find({}, { name: 1, dept: 1 }) // these fields PLUS _id
db.students.find({}, { name: 1, _id: 0 }) // exclude _id
db.students.find({}, { marks: 0, subjects: 0 }) // everything EXCEPT these
// Step 3: Project a nested field
db.students.find({}, { "marks.maths": 1, _id: 0 }) // one nested field
// Step 4: Project arrays
db.students.find({}, { subjects: { $slice: 2 } }) // first 2 array elements
db.students.find({}, { subjects: { $slice: -1 } }) // the LAST element
db.students.find({ subjects: "DS" }, { "subjects.$": 1 }) // the MATCHING one
// You cannot MIX inclusion and exclusion...
// Step 5: See that inclusion and exclusion cannot mix
db.students.find({}, { name: 1, dept: 0 }) // ERROR
// ...except for _id, which is the one exception.
db.students.find({}, { name: 1, _id: 0 }) // fine
In Python, through mongomock, 07_projection.py:
"""Experiment 7 — Projection."""
from fixtures import fresh_db
def inclusion_and_exclusion():
db = fresh_db()
d = db.students.find_one({"_id": 21}, {"name": 1, "dept": 1})
assert set(d) == {"_id", "name", "dept"}, "_id is included BY DEFAULT"
d = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})
assert set(d) == {"name"}
d = db.students.find_one({"_id": 21}, {"marks": 0, "subjects": 0})
assert "marks" not in d and "subjects" not in d
assert {"name", "dept", "age", "active", "_id"} <= set(d)
print(" inclusion adds _id automatically; exclusion keeps everything else")
def cannot_mix():
db = fresh_db()
try:
list(db.students.find({}, {"name": 1, "dept": 0}))
raise AssertionError("expected an error from mixing inclusion/exclusion")
except Exception as e:
assert not isinstance(e, AssertionError)
# _id is the ONE exception.
d = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})
assert set(d) == {"name"}
print(" mixing inclusion and exclusion raises; _id is the one exception")
def nested_projection():
db = fresh_db()
d = db.students.find_one({"_id": 21}, {"marks.maths": 1, "_id": 0})
assert d == {"marks": {"maths": 88}}, d
print(" 'marks.maths': 1 keeps the sub-document with only that field")
def array_projection():
db = fresh_db()
d = db.students.find_one({"_id": 21}, {"subjects": {"$slice": 2}, "_id": 0,
"name": 1})
assert d["subjects"] == ["DS", "Stats"], "the first 2"
d = db.students.find_one({"_id": 21}, {"subjects": {"$slice": -1}, "_id": 0,
"name": 1})
assert d["subjects"] == ["Python"], "a negative slice takes from the END"
print(" $slice: 2 -> first two; $slice: -1 -> the last one")
def projection_reduces_work():
"""Not just cosmetic -- fewer bytes read and sent."""
db = fresh_db()
full = db.students.find_one({"_id": 21})
thin = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})
assert len(str(full)) > 3 * len(str(thin)), \
"the projected document is a fraction of the size"
print(f" full document {len(str(full))} chars vs projected {len(str(thin))}")
print(f" and a projection fully covered by an index needs no document read")
def main():
print("Experiment 7 -- Projection")
# Step 1: Include and exclude fields
inclusion_and_exclusion()
# Step 2: See that the two cannot mix
cannot_mix()
# Step 3: Project nested fields
nested_projection()
# Step 4: Project arrays
array_projection()
# Step 5: See projection reduce the work
projection_reduces_work()
if __name__ == "__main__":
main()
In mongosh, 07_projection.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find({}, { name: 1, dept: 1 }) // these fields PLUS _id
[
{ _id: 21, name: 'Asha', dept: 'DS' },
{ _id: 22, name: 'Ravi', dept: 'DS' },
{ _id: 23, name: 'Meena', dept: 'Stats' },
{ _id: 24, name: 'Kiran', dept: 'DS' },
{ _id: 25, name: 'Bhanu', dept: 'Stats' }
]
collegeDB> db.students.find({}, { name: 1, _id: 0 }) // exclude _id
[
{ name: 'Asha' },
{ name: 'Ravi' },
{ name: 'Meena' },
{ name: 'Kiran' },
{ name: 'Bhanu' }
]
collegeDB> db.students.find({}, { marks: 0, subjects: 0 }) // everything EXCEPT these
[
{ _id: 21, name: 'Asha', dept: 'DS', age: 20, active: true },
{ _id: 22, name: 'Ravi', dept: 'DS', age: 21, active: true },
{ _id: 23, name: 'Meena', dept: 'Stats', age: 20, active: true },
{ _id: 24, name: 'Kiran', dept: 'DS', age: 22, active: false },
{ _id: 25, name: 'Bhanu', dept: 'Stats', age: 21, active: true }
]
collegeDB> db.students.find({}, { "marks.maths": 1, _id: 0 }) // one nested field
[
{ marks: { maths: 88 } },
{ marks: { maths: 65 } },
{ marks: { maths: 94 } },
{ marks: { maths: 71 } },
{ marks: { maths: 52 } }
]
collegeDB> db.students.find({}, { subjects: { $slice: 2 } }) // first 2 array elements
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({}, { subjects: { $slice: -1 } }) // the LAST element
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ subjects: "DS" }, { "subjects.$": 1 }) // the MATCHING one
[
{ _id: 21, subjects: [ 'DS' ] },
{ _id: 22, subjects: [ 'DS' ] },
{ _id: 24, subjects: [ 'DS' ] }
]
collegeDB> db.students.find({}, { name: 1, dept: 0 }) // ERROR
Uncaught
MongoServerError[Location31254]: Cannot do exclusion on field dept in inclusion projection
collegeDB> db.students.find({}, { name: 1, _id: 0 }) // fine
[
{ name: 'Asha' },
{ name: 'Ravi' },
{ name: 'Meena' },
{ name: 'Kiran' },
{ name: 'Bhanu' }
]
In Python, through mongomock, 07_projection.py:
OUTPUT
Experiment 7 -- Projection
inclusion adds _id automatically; exclusion keeps everything else
mixing inclusion and exclusion raises; _id is the one exception
'marks.maths': 1 keeps the sub-document with only that field
$slice: 2 -> first two; $slice: -1 -> the last one
full document 144 chars vs projected 16
and a projection fully covered by an index needs no document read
RESULT
Inclusion and exclusion cannot be mixed, except for _id; $slice and the positional $ cut arrays down.
Sort, limit and skip the results of a query.
Order results and page through them, the slow way and the fast way.
In mongosh, 08_sort_limit.js:
In Python, through mongomock, 08_sort_limit.py:
THE POINT
The server applies sort, then skip, then limit, regardless of the
chaining order — so .limit(3).sort(...) sorts everything and then takes
three.
The script also demonstrates range pagination ({ _id: { $gt: last } }) as
the fix for skip's linear cost.
In mongosh, 08_sort_limit.js:
// Experiment 8 -- Sorting documents, limiting output, skipping records.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 08_sort_limit.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: Sort
db.students.find().sort({ "marks.maths": -1 }) // -1 descending
db.students.find().sort({ dept: 1, "marks.maths": -1 }) // multi-key
// Step 3: Limit and skip
db.students.find().sort({ "marks.maths": -1 }).limit(3) // top 3
db.students.find().skip(2).limit(2) // "page 2"
// The server ALWAYS applies sort, then skip, then limit -- whatever order you
// chain them in. So .limit(3).sort(...) sorts EVERYTHING and then takes three.
// Step 4: See sort apply before limit, whatever the order written
db.students.find().limit(3).sort({ "marks.maths": -1 })
// Step 5: Paginate by range, not by skip
// skip(100000) makes the server WALK AND DISCARD 100,000 documents.
// Range pagination uses the index to jump straight there:
db.students.find().sort({ _id: 1 }).limit(2) // page 1
const lastSeenId = 22 // the last _id on page 1
db.students.find({ _id: { $gt: lastSeenId } }).sort({ _id: 1 }).limit(2) // page 2
// [Corrected: lastSeenId was never set, so this line failed with
// "ReferenceError: lastSeenId is not defined".]
// Step 6: Take the top document from the cursor
db.students.find().sort({ "marks.maths": -1 }).limit(1).next().name // topper
In Python, through mongomock, 08_sort_limit.py:
"""Experiment 8 — Sorting, limiting and skipping."""
from fixtures import fresh_db
def sorting():
db = fresh_db()
desc = [d["name"] for d in db.students.find().sort("marks.maths", -1)]
assert desc == ["Meena", "Asha", "Kiran", "Ravi", "Bhanu"], desc
asc = [d["name"] for d in db.students.find().sort("marks.maths", 1)]
assert asc == list(reversed(desc))
multi = [(d["dept"], d["marks"]["maths"])
for d in db.students.find().sort([("dept", 1), ("marks.maths", -1)])]
assert multi == [("DS", 88), ("DS", 71), ("DS", 65),
("Stats", 94), ("Stats", 52)], multi
print(f" sort by maths desc -> {desc}")
print(f" multi-key sort (dept asc, maths desc) groups then orders within")
def limit_and_skip():
db = fresh_db()
top3 = [d["name"] for d in db.students.find().sort("marks.maths", -1).limit(3)]
assert top3 == ["Meena", "Asha", "Kiran"]
page2 = [d["name"] for d in db.students.find().sort("_id", 1).skip(2).limit(2)]
assert page2 == ["Meena", "Kiran"], page2
print(f" top 3 by maths: {top3}; page 2 by _id: {page2}")
def order_is_fixed():
"""sort, then skip, then limit -- whatever order you chain them."""
db = fresh_db()
a = [d["name"] for d in db.students.find().sort("marks.maths", -1).limit(3)]
b = [d["name"] for d in db.students.find().limit(3).sort("marks.maths", -1)]
assert a == b == ["Meena", "Asha", "Kiran"], (a, b)
print(" .limit(3).sort(...) == .sort(...).limit(3) -- the server decides")
print(" it sorts EVERYTHING and then takes three, not the reverse")
def range_pagination_beats_skip():
db = fresh_db()
# skip-based
p1 = list(db.students.find().sort("_id", 1).limit(2))
p2 = list(db.students.find().sort("_id", 1).skip(2).limit(2))
p3 = list(db.students.find().sort("_id", 1).skip(4).limit(2))
# range-based: remember the last _id seen
r1 = list(db.students.find().sort("_id", 1).limit(2))
r2 = list(db.students.find({"_id": {"$gt": r1[-1]["_id"]}})
.sort("_id", 1).limit(2))
r3 = list(db.students.find({"_id": {"$gt": r2[-1]["_id"]}})
.sort("_id", 1).limit(2))
for skip_page, range_page in [(p1, r1), (p2, r2), (p3, r3)]:
assert [d["_id"] for d in skip_page] == [d["_id"] for d in range_page]
print(" range pagination gives IDENTICAL pages, and every page costs the")
print(" same -- skip(100000) would walk and discard 100,000 documents")
def cursor_versus_document():
db = fresh_db()
cursor = db.students.find({"_id": 21})
assert not hasattr(cursor, "get"), "find() gives a CURSOR"
assert list(cursor)[0]["name"] == "Asha"
doc = db.students.find_one({"_id": 21})
assert doc["name"] == "Asha", "findOne gives the DOCUMENT"
print(" find() -> a cursor (iterate it); findOne() -> a document or None")
def main():
print("Experiment 8 -- Sorting, limiting, skipping")
# Step 1: Sort
sorting()
# Step 2: Limit and skip
limit_and_skip()
# Step 3: See sort, skip and limit apply in a fixed order
order_is_fixed()
# Step 4: Paginate by range, not by skip
range_pagination_beats_skip()
# Step 5: Tell a cursor from a document
cursor_versus_document()
if __name__ == "__main__":
main()
In mongosh, 08_sort_limit.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find().sort({ "marks.maths": -1 }) // -1 descending
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find().sort({ dept: 1, "marks.maths": -1 }) // multi-key
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.students.find().sort({ "marks.maths": -1 }).limit(3) // top 3
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find().skip(2).limit(2) // "page 2"
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find().limit(3).sort({ "marks.maths": -1 })
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find().sort({ _id: 1 }).limit(2) // page 1
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
}
]
collegeDB> const lastSeenId = 22 // the last _id on page 1
collegeDB> db.students.find({ _id: { $gt: lastSeenId } }).sort({ _id: 1 }).limit(2) // page 2
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find().sort({ "marks.maths": -1 }).limit(1).next().name // topper
Meena
In Python, through mongomock, 08_sort_limit.py:
OUTPUT
Experiment 8 -- Sorting, limiting, skipping
sort by maths desc -> ['Meena', 'Asha', 'Kiran', 'Ravi', 'Bhanu']
multi-key sort (dept asc, maths desc) groups then orders within
top 3 by maths: ['Meena', 'Asha', 'Kiran']; page 2 by _id: ['Meena', 'Kiran']
.limit(3).sort(...) == .sort(...).limit(3) -- the server decides
it sorts EVERYTHING and then takes three, not the reverse
range pagination gives IDENTICAL pages, and every page costs the
same -- skip(100000) would walk and discard 100,000 documents
find() -> a cursor (iterate it); findOne() -> a document or None
Corrected: lastSeenId was used but never set, so the range query failed with
"ReferenceError: lastSeenId is not defined". It is now set to 22, the last _id on page 1.
RESULT
limit(3).sort(...) gives the top three, not three sorted; page 2 by range is Meena and Kiran.
Design and query an embedded data model.
Store a student's address and enrolments inside the student, and query them.
In mongosh, 09_embedded.js:
In Python, through mongomock, 09_embedded.py:
THE POINT
One read returns everything — no join anywhere — and the $elemMatch
requirement for two conditions on the enrolment array.
In mongosh, 09_embedded.js:
// Experiment 9 -- Designing an Embedded Data Model for a student-course
// enrollment system.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 09_embedded.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
// Step 1: Insert two students, with their address and enrolments embedded
use collegeDB
db.embedded.drop()
db.embedded.insertMany([
{ _id: 21, name: "Asha Kumari", dept: "DS",
address: { city: "Vijayawada", state: "AP", pin: "520010" }, // 1-to-1
enrollments: [ // 1-to-few
{ course: "DSC301", title: "Data Science with R", grade: "A", credits: 4 },
{ course: "STA302", title: "Statistical Foundations", grade: "B", credits: 3 }
] },
{ _id: 22, name: "Ravi Teja", dept: "DS",
address: { city: "Guntur", state: "AP", pin: "522002" },
enrollments: [
{ course: "DSC301", title: "Data Science with R", grade: "C", credits: 4 }
] }
])
// ONE read gets the student, their address and every enrolment. No join.
// Step 2: Read everything in one query
db.embedded.findOne({ _id: 21 })
// Step 3: Query nested fields and arrays
db.embedded.find({ "address.city": "Vijayawada" })
db.embedded.find({ "enrollments.grade": "A" })
// TWO conditions on an array of sub-documents NEED $elemMatch, or different
// elements may satisfy different conditions.
// Step 4: Use $elemMatch for two conditions on one element
db.embedded.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })
db.embedded.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" }) // WRONG
// Updating one element: the POSITIONAL operator $
// Step 5: Update one array element, and push another
db.embedded.updateOne({ _id: 21, "enrollments.course": "STA302" },
{ $set: { "enrollments.$.grade": "A" } })
db.embedded.updateOne({ _id: 22 },
{ $push: { enrollments: { course: "WEB303", title: "Web Technologies",
grade: "B", credits: 3 } } })
// Total credits, per student -- needs $unwind
// Step 6: Total the credits
db.embedded.aggregate([
{ $unwind: "$enrollments" },
{ $group: { _id: "$name", credits: { $sum: "$enrollments.credits" } } },
{ $sort: { _id: 1 } }
])
// [Changed: the $sort was added. $group returns its groups in no fixed order, and
// they came back in a different order from one run to the next.]
In Python, through mongomock, 09_embedded.py:
"""Experiment 9 — An embedded data model."""
import mongomock
DOCS = [
{"_id": 21, "name": "Asha Kumari", "dept": "DS",
"address": {"city": "Vijayawada", "state": "AP", "pin": "520010"},
"enrollments": [
{"course": "DSC301", "title": "Data Science with R", "grade": "A", "credits": 4},
{"course": "STA302", "title": "Statistical Foundations", "grade": "B", "credits": 3}]},
{"_id": 22, "name": "Ravi Teja", "dept": "DS",
"address": {"city": "Guntur", "state": "AP", "pin": "522002"},
"enrollments": [
{"course": "DSC301", "title": "Data Science with R", "grade": "C", "credits": 4}]},
]
def db():
d = mongomock.MongoClient().collegeDB
d.embedded.insert_many([dict(x) for x in DOCS])
return d
def one_read_gets_everything():
d = db()
doc = d.embedded.find_one({"_id": 21})
assert doc["name"] == "Asha Kumari"
assert doc["address"]["city"] == "Vijayawada"
assert len(doc["enrollments"]) == 2
assert doc["enrollments"][0]["title"] == "Data Science with R"
print(" ONE read returned the student, the address and both enrolments")
print(" -- no join anywhere, which is the point of embedding")
def querying_nested_and_arrays():
d = db()
assert [x["name"] for x in d.embedded.find({"address.city": "Vijayawada"})] \
== ["Asha Kumari"]
assert sorted(x["name"] for x in d.embedded.find({"enrollments.grade": "A"})) \
== ["Asha Kumari"]
assert sorted(x["name"] for x in d.embedded.find({"enrollments.course": "DSC301"})) \
== ["Asha Kumari", "Ravi Teja"], "ANY element matches"
print(" dot notation queries the embedded address and the enrolment array")
def elemmatch_is_required():
"""Two conditions on an array of sub-documents: different elements can
satisfy different conditions unless you use $elemMatch."""
d = db()
# Ravi took DSC301 (grade C) and nothing with grade A -- so he should NOT
# match "DSC301 with grade A".
wrong = sorted(x["name"] for x in d.embedded.find(
{"enrollments.course": "DSC301", "enrollments.grade": "A"}))
right = sorted(x["name"] for x in d.embedded.find(
{"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))
assert right == ["Asha Kumari"], right
assert wrong == ["Asha Kumari"], wrong # Ravi has no grade A at all
# Now make the trap visible: give Ravi an A in a DIFFERENT course.
d.embedded.update_one({"_id": 22}, {"$push": {"enrollments": {
"course": "WEB303", "title": "Web Technologies", "grade": "A", "credits": 3}}})
wrong2 = sorted(x["name"] for x in d.embedded.find(
{"enrollments.course": "DSC301", "enrollments.grade": "A"}))
right2 = sorted(x["name"] for x in d.embedded.find(
{"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))
assert wrong2 == ["Asha Kumari", "Ravi Teja"], \
"WRONG: Ravi matched using DSC301 from one element and grade A from another"
assert right2 == ["Asha Kumari"], "RIGHT: one element must satisfy both"
print(" after giving Ravi an A in a DIFFERENT course:")
print(f" without $elemMatch -> {wrong2} (Ravi is a FALSE match)")
print(f" with $elemMatch -> {right2}")
def positional_update():
d = db()
d.embedded.update_one({"_id": 21, "enrollments.course": "STA302"},
{"$set": {"enrollments.$.grade": "A"}})
doc = d.embedded.find_one({"_id": 21})
grades = {e["course"]: e["grade"] for e in doc["enrollments"]}
assert grades == {"DSC301": "A", "STA302": "A"}, grades
print(" the positional $ updated the element the QUERY matched")
def pushing_and_aggregating():
d = db()
d.embedded.update_one({"_id": 22}, {"$push": {"enrollments": {
"course": "WEB303", "title": "Web Technologies", "grade": "B", "credits": 3}}})
assert len(d.embedded.find_one({"_id": 22})["enrollments"]) == 2
credits = {r["_id"]: r["credits"] for r in d.embedded.aggregate([
{"$unwind": "$enrollments"},
{"$group": {"_id": "$name", "credits": {"$sum": "$enrollments.credits"}}}])}
assert credits == {"Asha Kumari": 7, "Ravi Teja": 7}, credits
print(f" total credits per student (needs $unwind): {credits}")
def the_limitation():
"""Embedding is right here because enrolments are BOUNDED. State the case
where it would be wrong."""
d = db()
doc = d.embedded.find_one({"_id": 21})
assert len(doc["enrollments"]) <= 10, "a degree has a bounded number of courses"
print(" embedding is right here because a student's enrolments are BOUNDED")
print(" -- attendance records or log entries would NOT be, and would")
print(" eventually breach the 16 MB document limit")
def main():
print("Experiment 9 -- An embedded data model")
# Step 1: Read everything in one query
one_read_gets_everything()
# Step 2: Query nested fields and arrays
querying_nested_and_arrays()
# Step 3: Use $elemMatch for two conditions on one element
elemmatch_is_required()
# Step 4: Update one array element
positional_update()
# Step 5: Push, and aggregate
pushing_and_aggregating()
# Step 6: State the model's limitation
the_limitation()
if __name__ == "__main__":
main()
In mongosh, 09_embedded.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> db.embedded.drop()
true
collegeDB> db.embedded.insertMany([
... { _id: 21, name: "Asha Kumari", dept: "DS",
... address: { city: "Vijayawada", state: "AP", pin: "520010" }, // 1-to-1
... enrollments: [ // 1-to-few
... { course: "DSC301", title: "Data Science with R", grade: "A", credits: 4 },
... { course: "STA302", title: "Statistical Foundations", grade: "B", credits: 3 }
... ] },
... { _id: 22, name: "Ravi Teja", dept: "DS",
... address: { city: "Guntur", state: "AP", pin: "522002" },
... enrollments: [
... { course: "DSC301", title: "Data Science with R", grade: "C", credits: 4 }
... ] }
... ])
{ acknowledged: true, insertedIds: { '0': 21, '1': 22 } }
collegeDB> db.embedded.findOne({ _id: 21 })
{
_id: 21,
name: 'Asha Kumari',
dept: 'DS',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
enrollments: [
{
course: 'DSC301',
title: 'Data Science with R',
grade: 'A',
credits: 4
},
{
course: 'STA302',
title: 'Statistical Foundations',
grade: 'B',
credits: 3
}
]
}
collegeDB> db.embedded.find({ "address.city": "Vijayawada" })
[
{
_id: 21,
name: 'Asha Kumari',
dept: 'DS',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
enrollments: [
{
course: 'DSC301',
title: 'Data Science with R',
grade: 'A',
credits: 4
},
{
course: 'STA302',
title: 'Statistical Foundations',
grade: 'B',
credits: 3
}
]
}
]
collegeDB> db.embedded.find({ "enrollments.grade": "A" })
[
{
_id: 21,
name: 'Asha Kumari',
dept: 'DS',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
enrollments: [
{
course: 'DSC301',
title: 'Data Science with R',
grade: 'A',
credits: 4
},
{
course: 'STA302',
title: 'Statistical Foundations',
grade: 'B',
credits: 3
}
]
}
]
collegeDB> db.embedded.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })
[
{
_id: 21,
name: 'Asha Kumari',
dept: 'DS',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
enrollments: [
{
course: 'DSC301',
title: 'Data Science with R',
grade: 'A',
credits: 4
},
{
course: 'STA302',
title: 'Statistical Foundations',
grade: 'B',
credits: 3
}
]
}
]
collegeDB> db.embedded.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" }) // WRONG
[
{
_id: 21,
name: 'Asha Kumari',
dept: 'DS',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
enrollments: [
{
course: 'DSC301',
title: 'Data Science with R',
grade: 'A',
credits: 4
},
{
course: 'STA302',
title: 'Statistical Foundations',
grade: 'B',
credits: 3
}
]
}
]
collegeDB> db.embedded.updateOne({ _id: 21, "enrollments.course": "STA302" },
... { $set: { "enrollments.$.grade": "A" } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.embedded.updateOne({ _id: 22 },
... { $push: { enrollments: { course: "WEB303", title: "Web Technologies",
... grade: "B", credits: 3 } } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.embedded.aggregate([
... { $unwind: "$enrollments" },
... { $group: { _id: "$name", credits: { $sum: "$enrollments.credits" } } },
... { $sort: { _id: 1 } }
... ])
[
{ _id: 'Asha Kumari', credits: 7 },
{ _id: 'Ravi Teja', credits: 7 }
]
In Python, through mongomock, 09_embedded.py:
OUTPUT
Experiment 9 -- An embedded data model
ONE read returned the student, the address and both enrolments
-- no join anywhere, which is the point of embedding
dot notation queries the embedded address and the enrolment array
after giving Ravi an A in a DIFFERENT course:
without $elemMatch -> ['Asha Kumari', 'Ravi Teja'] (Ravi is a FALSE match)
with $elemMatch -> ['Asha Kumari']
the positional $ updated the element the QUERY matched
total credits per student (needs $unwind): {'Asha Kumari': 7, 'Ravi Teja': 7}
embedding is right here because a student's enrolments are BOUNDED
-- attendance records or log entries would NOT be, and would
eventually breach the 16 MB document limit
Changed: the credits total now ends with a $sort. $group returns its groups in
no fixed order, and on the real server Asha and Ravi came back in a different order from one run
to the next.
RESULT
One findOne returns the student with everything; $elemMatch finds Asha's A in DSC301 where the two separate conditions match wrongly.
Design and query a normalised model, with documents that reference each other.
Join referenced collections with $lookup, and see what references do not guarantee.
In mongosh, 10_referenced.js:
In Python, through mongomock, 10_referenced.py:
THE POINT
$lookup produces an array even for a one-to-one match, which is
why $unwind follows it; and it is a left outer join — an unmatched
document gets an empty array, not nothing.
In mongosh, 10_referenced.js:
// Experiment 10 -- Designing a Normalized Data Model using document
// references.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 10_referenced.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
// Step 1: Insert a student, the courses and the enrolments, as references
use collegeDB
db.students.insertOne({ _id: 21, name: "Asha Kumari", dept: "DS" })
db.courses.insertMany([
{ _id: "DSC301", title: "Data Science with R", credits: 4, instructor: "Dr. Rao" },
{ _id: "STA302", title: "Statistical Foundations", credits: 3, instructor: "Dr. Devi" }
])
db.enrollments.insertMany([
{ student_id: 21, course_id: "DSC301", grade: "A" },
{ student_id: 21, course_id: "STA302", grade: "B" }
])
// Two reads, application-side
// Step 2: Two reads, or one $lookup
const s = db.students.findOne({ _id: 21 })
const e = db.enrollments.find({ student_id: 21 }).toArray()
// Or one $lookup. Note: 'as' is ALWAYS an array, even for a 1-to-1 match,
// which is why $unwind almost always follows.
db.enrollments.aggregate([
{ $lookup: { from: "courses", localField: "course_id",
foreignField: "_id", as: "course" } },
{ $unwind: "$course" },
{ $project: { _id: 0, course: "$course.title", grade: 1 } }
])
// $lookup is a LEFT OUTER JOIN -- an unmatched document gets an EMPTY ARRAY
// Step 3: See $lookup keep an unmatched row, as a left outer join
db.enrollments.insertOne({ student_id: 21, course_id: "GONE", grade: "F" })
db.enrollments.aggregate([
{ $lookup: { from: "courses", localField: "course_id",
foreignField: "_id", as: "course" } }
]) // the GONE row has course: []
// NOTHING stops a reference pointing at a document that does not exist.
// In Course 5 a foreign key would. Here the application must check.
// Step 4: See references not enforced
db.courses.deleteOne({ _id: "DSC301" }) // the enrolments still reference it
// Index the foreignField, or every input document causes a collection scan
// Step 5: Index the foreign fields
db.enrollments.createIndex({ course_id: 1 })
db.enrollments.createIndex({ student_id: 1 })
In Python, through mongomock, 10_referenced.py:
"""Experiment 10 — A normalized model using document references."""
from fixtures import fresh_db
def two_reads_or_one_lookup():
db = fresh_db()
# Application-side: two reads.
student = db.students.find_one({"_id": 21})
enrolments = list(db.enrollments.find({"student_id": 21}))
assert student["name"] == "Asha"
assert len(enrolments) == 2
# Or one $lookup.
joined = list(db.enrollments.aggregate([
{"$match": {"student_id": 21}},
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": "$course"},
{"$project": {"_id": 0, "course": "$course.title", "grade": 1}}]))
assert sorted(j["course"] for j in joined) == \
["Data Science with R", "Statistical Foundations"], joined
print(f" two reads, or one $lookup -> {[j['course'] for j in joined]}")
def lookup_always_produces_an_array():
"""Which is why $unwind almost always follows it."""
db = fresh_db()
raw = list(db.enrollments.aggregate([
{"$match": {"student_id": 21, "course_id": "DSC301"}},
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}}]))
assert isinstance(raw[0]["course"], list), "an ARRAY even for a 1-to-1 match"
assert len(raw[0]["course"]) == 1
unwound = list(db.enrollments.aggregate([
{"$match": {"student_id": 21, "course_id": "DSC301"}},
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": "$course"}]))
assert isinstance(unwound[0]["course"], dict), "$unwind makes it a sub-document"
assert unwound[0]["course"]["title"] == "Data Science with R"
print(" $lookup gave course: [ {...} ]; $unwind made it course: { ... }")
print(" without $unwind, '$course.title' would be an ARRAY of titles")
def lookup_is_a_left_outer_join():
db = fresh_db()
db.enrollments.insert_one({"student_id": 21, "course_id": "GONE", "grade": "F"})
rows = list(db.enrollments.aggregate([
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}}]))
orphan = [r for r in rows if r["course_id"] == "GONE"][0]
assert orphan["course"] == [], "an unmatched document gets an EMPTY ARRAY"
assert len(rows) == 6, "the orphan is KEPT -- left outer join"
# $unwind would then DROP it, unless you preserve empties.
dropped = list(db.enrollments.aggregate([
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": "$course"}]))
assert len(dropped) == 5, "$unwind silently dropped the orphan"
kept = list(db.enrollments.aggregate([
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": {"path": "$course", "preserveNullAndEmptyArrays": True}}]))
assert len(kept) == 6
print(" $lookup kept the orphan with course: []; $unwind then DROPPED it")
print(" preserveNullAndEmptyArrays: true keeps it -- 6 rows, not 5")
def references_are_not_enforced():
"""The row that matters most in the RDBMS comparison."""
db = fresh_db()
assert db.enrollments.count_documents({"course_id": "DSC301"}) == 2
db.courses.delete_one({"_id": "DSC301"})
assert db.courses.find_one({"_id": "DSC301"}) is None
assert db.enrollments.count_documents({"course_id": "DSC301"}) == 2, \
"the enrolments STILL reference a course that no longer exists"
# An integrity check the application must run for itself.
valid = set(db.courses.distinct("_id"))
orphans = [e for e in db.enrollments.find()
if e["course_id"] not in valid]
assert len(orphans) == 2
print(" deleting the course left 2 DANGLING references, with no error")
print(" Course 5's foreign key would have refused; here the")
print(f" application must check -- {len(orphans)} orphans found")
def index_the_foreign_field():
db = fresh_db()
db.enrollments.create_index("course_id")
db.enrollments.create_index("student_id")
names = {i["name"] for i in db.enrollments.list_indexes()}
assert "course_id_1" in names and "student_id_1" in names
print(" indexed both foreignFields -- without them, $lookup scans the")
print(" whole other collection ONCE PER INPUT DOCUMENT")
def main():
print("Experiment 10 -- A normalized model with references")
# Step 1: Two reads, or one $lookup
two_reads_or_one_lookup()
# Step 2: See $lookup always give an array
lookup_always_produces_an_array()
# Step 3: See $lookup is a left outer join
lookup_is_a_left_outer_join()
# Step 4: See references are not enforced
references_are_not_enforced()
# Step 5: Index the foreign field
index_the_foreign_field()
if __name__ == "__main__":
main()
In mongosh, 10_referenced.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> db.students.insertOne({ _id: 21, name: "Asha Kumari", dept: "DS" })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.courses.insertMany([
... { _id: "DSC301", title: "Data Science with R", credits: 4, instructor: "Dr. Rao" },
... { _id: "STA302", title: "Statistical Foundations", credits: 3, instructor: "Dr. Devi" }
... ])
{ acknowledged: true, insertedIds: { '0': 'DSC301', '1': 'STA302' } }
collegeDB> db.enrollments.insertMany([
... { student_id: 21, course_id: "DSC301", grade: "A" },
... { student_id: 21, course_id: "STA302", grade: "B" }
... ])
{
acknowledged: true,
insertedIds: {
'0': ObjectId('6ac214d315618cb446277d3b'),
'1': ObjectId('6ac214d315618cb446277d3c')
}
}
collegeDB> const s = db.students.findOne({ _id: 21 })
collegeDB> const e = db.enrollments.find({ student_id: 21 }).toArray()
collegeDB> db.enrollments.aggregate([
... { $lookup: { from: "courses", localField: "course_id",
... foreignField: "_id", as: "course" } },
... { $unwind: "$course" },
... { $project: { _id: 0, course: "$course.title", grade: 1 } }
... ])
[
{ grade: 'A', course: 'Data Science with R' },
{ grade: 'B', course: 'Statistical Foundations' }
]
collegeDB> db.enrollments.insertOne({ student_id: 21, course_id: "GONE", grade: "F" })
{
acknowledged: true,
insertedId: ObjectId('6ac214d315618cb446277d3d')
}
collegeDB> db.enrollments.aggregate([
... { $lookup: { from: "courses", localField: "course_id",
... foreignField: "_id", as: "course" } }
... ]) // the GONE row has course: []
[
{
_id: ObjectId('6ac214d315618cb446277d3b'),
student_id: 21,
course_id: 'DSC301',
grade: 'A',
course: [
{
_id: 'DSC301',
title: 'Data Science with R',
credits: 4,
instructor: 'Dr. Rao'
}
]
},
{
_id: ObjectId('6ac214d315618cb446277d3c'),
student_id: 21,
course_id: 'STA302',
grade: 'B',
course: [
{
_id: 'STA302',
title: 'Statistical Foundations',
credits: 3,
instructor: 'Dr. Devi'
}
]
},
{
_id: ObjectId('6ac214d315618cb446277d3d'),
student_id: 21,
course_id: 'GONE',
grade: 'F',
course: []
}
]
collegeDB> db.courses.deleteOne({ _id: "DSC301" }) // the enrolments still reference it
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.enrollments.createIndex({ course_id: 1 })
course_id_1
collegeDB> db.enrollments.createIndex({ student_id: 1 })
student_id_1
In Python, through mongomock, 10_referenced.py:
OUTPUT
Experiment 10 -- A normalized model with references
two reads, or one $lookup -> ['Data Science with R', 'Statistical Foundations']
$lookup gave course: [ {...} ]; $unwind made it course: { ... }
without $unwind, '$course.title' would be an ARRAY of titles
$lookup kept the orphan with course: []; $unwind then DROPPED it
preserveNullAndEmptyArrays: true keeps it -- 6 rows, not 5
deleting the course left 2 DANGLING references, with no error
Course 5's foreign key would have refused; here the
application must check -- 2 orphans found
indexed both foreignFields -- without them, $lookup scans the
whole other collection ONCE PER INPUT DOCUMENT
RESULT
$lookup returns an array, and keeps the GONE enrolment with an empty one; deleting a course leaves its enrolments pointing at nothing.
Model one-to-one, one-to-many and many-to-many relationships.
Model each kind of relationship, and query it from both ends.
In mongosh, 11_relationships.js:
In Python, through mongomock, 11_relationships.py:
THE POINT
All three modelled and queried:
| Relationship | Model | Query direction |
|---|---|---|
| One-to-one | Embed the address | Both, from the student |
| One-to-many | Reference from the child | Course → its enrolments |
| Many-to-many | A junction collection | Both directions |
The junction collection is what carries the grade — an attribute of the relationship, belonging to neither entity.
In mongosh, 11_relationships.js:
// Experiment 11 -- Modeling relationships: One-to-One, One-to-Many,
// Many-to-Many in MongoDB.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 11_relationships.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
use collegeDB
// Step 1: One-to-one: embed
// Small, bounded, always read together, never queried alone.
db.people.insertOne({
_id: 21, name: "Asha Kumari",
address: { city: "Vijayawada", state: "AP", pin: "520010" }
})
db.people.find({ "address.pin": "520010" })
// Step 2: One-to-many: reference from the child
// A course has many enrolments. The array must NOT live on the course, because
// it is unbounded -- that is the 16 MB trap.
db.courses.insertOne({ _id: "DSC301", title: "Data Science with R" })
db.enrollments.insertMany([
{ course_id: "DSC301", student_id: 21, grade: "A" },
{ course_id: "DSC301", student_id: 22, grade: "C" }
])
db.enrollments.find({ course_id: "DSC301" }) // the "many" side
// Step 3: One-to-few: embed an array
db.people.updateOne({ _id: 21 },
{ $set: { phones: ["9876543210", "9876543211"] } })
// Step 4: Many-to-many: reference both ways
// Option A: an array of ids on one side
db.students.updateOne({ _id: 21 },
{ $set: { name: "Asha Kumari", course_ids: ["DSC301", "STA302"] } },
{ upsert: true })
// [Corrected: this was an updateOne without upsert, on a student 21 who exists
// only if Experiment 10 has just been run in the same database. On its own it
// matched nothing, and the find below returned no one.]
db.students.find({ course_ids: "DSC301" }) // who takes DSC301?
// Option C: a junction collection -- REQUIRED when the relationship itself
// has attributes. The grade belongs to neither the student nor the course.
db.enrollments.find({ student_id: 21 }) // this student's courses
db.enrollments.find({ course_id: "DSC301" }) // this course's students
In Python, through mongomock, 11_relationships.py:
"""Experiment 11 — One-to-one, one-to-many and many-to-many."""
import mongomock
from fixtures import fresh_db
def one_to_one_embed():
db = mongomock.MongoClient().collegeDB
db.people.insert_one({"_id": 21, "name": "Asha Kumari",
"address": {"city": "Vijayawada", "state": "AP",
"pin": "520010"}})
doc = db.people.find_one({"_id": 21})
assert doc["address"]["city"] == "Vijayawada", "one read, no join"
assert [d["name"] for d in db.people.find({"address.pin": "520010"})] \
== ["Asha Kumari"], "and it is still queryable"
print(" 1-to-1: EMBED. One read; the sub-document is still queryable by dot")
def one_to_few_embed_array():
db = mongomock.MongoClient().collegeDB
db.people.insert_one({"_id": 21, "name": "Asha",
"phones": ["9876543210", "9876543211"]})
assert len(db.people.find_one({"_id": 21})["phones"]) == 2
assert [d["_id"] for d in db.people.find({"phones": "9876543210"})] == [21], \
"any element matches"
print(" 1-to-few: embed as an ARRAY -- bounded, so no 16 MB risk")
def one_to_many_reference_from_child():
"""The array must NOT live on the parent: it is unbounded."""
db = fresh_db()
# The CHILD holds the parent's id.
for e in db.enrollments.find({"course_id": "DSC301"}):
assert "course_id" in e
got = sorted(e["student_id"] for e in db.enrollments.find({"course_id": "DSC301"}))
assert got == [21, 22], got
# Adding a thousand enrolments does not grow the course document at all.
before = len(str(db.courses.find_one({"_id": "DSC301"})))
db.enrollments.insert_many([{"course_id": "DSC301", "student_id": 1000 + i,
"grade": "B"} for i in range(1000)])
after = len(str(db.courses.find_one({"_id": "DSC301"})))
assert before == after, "the COURSE document is unchanged -- that is the point"
assert db.enrollments.count_documents({"course_id": "DSC301"}) == 1002
print(f" 1-to-many: reference from the CHILD. 1000 more enrolments left the")
print(f" course document at {after} chars -- an embedded array would")
print(f" have grown it, and eventually breached 16 MB")
def many_to_many_both_ways():
db = fresh_db()
# Option A: an array of ids on the student
db.students.update_one({"_id": 21},
{"$set": {"course_ids": ["DSC301", "STA302"]}})
assert [d["name"] for d in db.students.find({"course_ids": "DSC301"})] == ["Asha"]
# Option C: the junction collection -- queryable from BOTH directions
hers = sorted(e["course_id"] for e in db.enrollments.find({"student_id": 21}))
assert hers == ["DSC301", "STA302"], hers
theirs = sorted(e["student_id"] for e in db.enrollments.find({"course_id": "DSC301"}))
assert theirs == [21, 22], theirs
print(f" M-to-M: student 21 takes {hers}; DSC301 has students {theirs}")
def the_junction_carries_the_relationship_attributes():
"""Why option C wins when the relationship has its own data."""
db = fresh_db()
e = db.enrollments.find_one({"student_id": 21, "course_id": "DSC301"})
assert e["grade"] == "A"
# The grade belongs to NEITHER entity:
assert "grade" not in db.students.find_one({"_id": 21})
assert "grade" not in db.courses.find_one({"_id": "DSC301"})
# An array of ids could not hold it.
db.students.update_one({"_id": 21}, {"$set": {"course_ids": ["DSC301"]}})
s = db.students.find_one({"_id": 21})
assert s["course_ids"] == ["DSC301"], "just an id -- nowhere to put the grade"
print(" the GRADE lives on the enrolment, not on the student or the course")
print(" -- exactly the reasoning that produces a junction TABLE in")
print(" Course 5, and it survives the translation unchanged")
def main():
print("Experiment 11 -- Modelling relationships")
# Step 1: One-to-one: embed
one_to_one_embed()
# Step 2: One-to-few: embed an array
one_to_few_embed_array()
# Step 3: One-to-many: reference from the child
one_to_many_reference_from_child()
# Step 4: Many-to-many: reference both ways
many_to_many_both_ways()
# Step 5: Give the relationship's own data to a junction collection
the_junction_carries_the_relationship_attributes()
if __name__ == "__main__":
main()
In mongosh, 11_relationships.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> db.people.insertOne({
... _id: 21, name: "Asha Kumari",
... address: { city: "Vijayawada", state: "AP", pin: "520010" }
... })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.people.find({ "address.pin": "520010" })
[
{
_id: 21,
name: 'Asha Kumari',
address: { city: 'Vijayawada', state: 'AP', pin: '520010' }
}
]
collegeDB> db.courses.insertOne({ _id: "DSC301", title: "Data Science with R" })
{ acknowledged: true, insertedId: 'DSC301' }
collegeDB> db.enrollments.insertMany([
... { course_id: "DSC301", student_id: 21, grade: "A" },
... { course_id: "DSC301", student_id: 22, grade: "C" }
... ])
{
acknowledged: true,
insertedIds: {
'0': ObjectId('6ac214d8097ed1d9dbdc315a'),
'1': ObjectId('6ac214d8097ed1d9dbdc315b')
}
}
collegeDB> db.enrollments.find({ course_id: "DSC301" }) // the "many" side
[
{
_id: ObjectId('6ac214d8097ed1d9dbdc315a'),
course_id: 'DSC301',
student_id: 21,
grade: 'A'
},
{
_id: ObjectId('6ac214d8097ed1d9dbdc315b'),
course_id: 'DSC301',
student_id: 22,
grade: 'C'
}
]
collegeDB> db.people.updateOne({ _id: 21 },
... { $set: { phones: ["9876543210", "9876543211"] } })
{
acknowledged: true,
insertedId: null,
matchedCount: 1,
modifiedCount: 1,
upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 },
... { $set: { name: "Asha Kumari", course_ids: ["DSC301", "STA302"] } },
... { upsert: true })
{
acknowledged: true,
insertedId: 21,
matchedCount: 0,
modifiedCount: 0,
upsertedCount: 1
}
collegeDB> db.students.find({ course_ids: "DSC301" }) // who takes DSC301?
[
{ _id: 21, course_ids: [ 'DSC301', 'STA302' ], name: 'Asha Kumari' }
]
collegeDB> db.enrollments.find({ student_id: 21 }) // this student's courses
[
{
_id: ObjectId('6ac214d8097ed1d9dbdc315a'),
course_id: 'DSC301',
student_id: 21,
grade: 'A'
}
]
collegeDB> db.enrollments.find({ course_id: "DSC301" }) // this course's students
[
{
_id: ObjectId('6ac214d8097ed1d9dbdc315a'),
course_id: 'DSC301',
student_id: 21,
grade: 'A'
},
{
_id: ObjectId('6ac214d8097ed1d9dbdc315b'),
course_id: 'DSC301',
student_id: 22,
grade: 'C'
}
]
In Python, through mongomock, 11_relationships.py:
OUTPUT
Experiment 11 -- Modelling relationships
1-to-1: EMBED. One read; the sub-document is still queryable by dot
1-to-few: embed as an ARRAY -- bounded, so no 16 MB risk
1-to-many: reference from the CHILD. 1000 more enrolments left the
course document at 88 chars -- an embedded array would
have grown it, and eventually breached 16 MB
M-to-M: student 21 takes ['DSC301', 'STA302']; DSC301 has students [21, 22]
the GRADE lives on the enrolment, not on the student or the course
-- exactly the reasoning that produces a junction TABLE in
Course 5, and it survives the translation unchanged
Corrected: the many-to-many's updateOne was on student 21, who exists only if
Experiment 10 has just been run in the same database; on its own it matched nothing, and "who
takes DSC301?" returned no one. It is now an upsert, which works either way.
RESULT
Each relationship is modelled and queried both ways; the junction collection holds the grade.
Validate documents against a JSON Schema.
Enforce a schema on a collection, and add one to a collection that already holds data.
In mongosh, 12_validation.js:
In Python, through mongomock, 12_validation.py:
THE POINT
Attach the schema in the safest mode first — moderate + warn — find the
offenders with { $nor: [ { $jsonSchema: … } ] }, fix them, then tighten to strict + error.
mongomock does not enforce $jsonSchema, so the Python half implements the same rules in
code and asserts that conforming documents pass and each kind of violation is caught — among them a missing
required field, a wrong type, a value outside the enum. The real server enforces them, and its
error names the rule that failed.
In mongosh, 12_validation.js:
// Experiment 12 -- Implementing schema validation using JSON Schema.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 12_validation.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
db.validated.drop()
// Step 2: Write the schema
const schema = {
bsonType: "object",
required: ["roll", "name", "dept"],
properties: {
roll: { bsonType: "int", minimum: 1,
description: "required integer, at least 1" },
name: { bsonType: "string", minLength: 3, maxLength: 80 },
dept: { enum: ["DS", "Stats", "CS"],
description: "must be DS, Stats or CS" },
marks: {
bsonType: "object",
properties: {
maths: { bsonType: "int", minimum: 0, maximum: 100 },
stats: { bsonType: "int", minimum: 0, maximum: 100 }
}
},
email: { bsonType: "string", pattern: "^[^\\s@]+@[^\\s@]+\\.[^\\s@]{2,}$" }
}
}
// Step 3: Create a collection that enforces it
db.createCollection("validated", {
validator: { $jsonSchema: schema },
validationLevel: "strict",
validationAction: "error"
})
// Step 4: Insert a conforming document, and four that are not
db.validated.insertOne({ roll: NumberInt(21), name: "Asha", dept: "DS",
marks: { maths: NumberInt(88) } }) // OK
db.validated.insertOne({ name: "NoRoll", dept: "DS" }) // missing required
db.validated.insertOne({ roll: NumberInt(22), name: "Ab", dept: "DS" }) // too short
db.validated.insertOne({ roll: NumberInt(23), name: "Ravi", dept: "Physics" }) // enum
db.validated.insertOne({ roll: NumberInt(24), name: "Meena", dept: "DS",
marks: { maths: NumberInt(150) } }) // out of range
// Step 5: Add validation to data already there, in stages
// 1. Attach it in the SAFEST mode: log, do not block.
db.runCommand({ collMod: "students",
validator: { $jsonSchema: schema },
validationLevel: "moderate",
validationAction: "warn" })
// 2. FIND the offenders -- $nor inverts a $jsonSchema match
db.students.find({ $nor: [ { $jsonSchema: schema } ] })
// [Corrected: these two read { $jsonSchema: { /* as above */ } }, an empty
// schema, which every document passes, so no offender could be found. The
// schema is now a const, defined once and used three times.]
// 3. Fix them, confirm the count is zero, then:
db.runCommand({ collMod: "students",
validationLevel: "strict", validationAction: "error" })
In Python, through mongomock, 12_validation.py:
"""Experiment 12 — Schema validation with JSON Schema.
*** mongomock does NOT enforce $jsonSchema. ***
Rather than pretend it does, this script implements the SAME rules in code and
asserts that a conforming document passes and each kind of violation is caught.
The mongosh half (12_validation.js) is what you run on a real server.
"""
import re
import mongomock
SCHEMA = {
"bsonType": "object",
"required": ["roll", "name", "dept"],
"properties": {
"roll": {"bsonType": "int", "minimum": 1},
"name": {"bsonType": "string", "minLength": 3, "maxLength": 80},
"dept": {"enum": ["DS", "Stats", "CS"]},
"marks": {"bsonType": "object", "properties": {
"maths": {"bsonType": "int", "minimum": 0, "maximum": 100},
"stats": {"bsonType": "int", "minimum": 0, "maximum": 100}}},
"email": {"bsonType": "string",
"pattern": r"^[^\s@]+@[^\s@]+\.[^\s@]{2,}$"},
},
}
BSON_TYPES = {"int": int, "string": str, "object": dict, "double": float,
"bool": bool, "array": list}
def violations(doc, schema=SCHEMA, path=""):
"""Return the list of ways `doc` fails `schema`. Empty means it conforms."""
out = []
for field in schema.get("required", []):
if field not in doc:
out.append(f"{path}{field}: required field missing")
for field, rule in schema.get("properties", {}).items():
if field not in doc:
continue
value = doc[field]
where = f"{path}{field}"
want = rule.get("bsonType")
if want and not isinstance(value, BSON_TYPES[want]):
out.append(f"{where}: expected {want}, got {type(value).__name__}")
continue
if "enum" in rule and value not in rule["enum"]:
out.append(f"{where}: {value!r} not in {rule['enum']}")
if "minimum" in rule and value < rule["minimum"]:
out.append(f"{where}: {value} below minimum {rule['minimum']}")
if "maximum" in rule and value > rule["maximum"]:
out.append(f"{where}: {value} above maximum {rule['maximum']}")
if "minLength" in rule and len(value) < rule["minLength"]:
out.append(f"{where}: shorter than {rule['minLength']}")
if "maxLength" in rule and len(value) > rule["maxLength"]:
out.append(f"{where}: longer than {rule['maxLength']}")
if "pattern" in rule and not re.match(rule["pattern"], value):
out.append(f"{where}: does not match {rule['pattern']}")
if rule.get("bsonType") == "object" and "properties" in rule:
out.extend(violations(value, rule, path=f"{where}."))
return out
def conforming_document_passes():
good = {"roll": 21, "name": "Asha", "dept": "DS", "marks": {"maths": 88},
"email": "asha@nri.ac.in"}
assert violations(good) == [], violations(good)
print(" a conforming document produces no violations")
def each_violation_is_caught():
cases = {
"missing required": ({"name": "NoRoll", "dept": "DS"}, "required"),
"name too short": ({"roll": 22, "name": "Ab", "dept": "DS"}, "shorter"),
"dept not in enum": ({"roll": 23, "name": "Ravi", "dept": "Physics"}, "not in"),
"marks over 100": ({"roll": 24, "name": "Meena", "dept": "DS",
"marks": {"maths": 150}}, "above maximum"),
"roll below 1": ({"roll": 0, "name": "Zero", "dept": "DS"}, "below minimum"),
"roll wrong type": ({"roll": "21", "name": "Str", "dept": "DS"}, "expected int"),
"bad email": ({"roll": 25, "name": "Bhanu", "dept": "DS",
"email": "not-an-email"}, "does not match"),
}
for label, (doc, expected) in cases.items():
found = violations(doc)
assert found, f"{label}: expected a violation, got none"
assert any(expected in v for v in found), f"{label}: {found}"
print(f" {label:18s} -> {found[0]}")
print(" every rule in the schema is enforced")
def find_the_offenders():
"""The migration step: find what does NOT conform, before tightening."""
db = mongomock.MongoClient().collegeDB
db.messy.insert_many([
{"roll": 21, "name": "Asha", "dept": "DS"}, # ok
{"roll": 22, "name": "Ab", "dept": "DS"}, # name too short
{"name": "NoRoll", "dept": "Stats"}, # missing roll
{"roll": 24, "name": "Kiran", "dept": "Physics"}, # bad dept
])
offenders = [(d.get("name"), violations(d)) for d in db.messy.find()
if violations(d)]
assert len(offenders) == 3, offenders
assert all(v for _, v in offenders)
conforming = [d for d in db.messy.find() if not violations(d)]
assert len(conforming) == 1 and conforming[0]["name"] == "Asha"
print(f" 3 of 4 documents violate the schema:")
for name, vs in offenders:
print(f" {str(name):10s} {vs[0]}")
print(" on a real server: $nor: [ { $jsonSchema: ... } ] finds exactly these")
def the_migration_path():
"""Turning strict validation on over dirty data breaks the application."""
levels = {
"strict": "applies to EVERY insert and update",
"moderate": "applies to inserts, and to updates of documents that ALREADY conform",
"off": "applies to nothing",
}
actions = {"error": "REJECT the write", "warn": "LOG it and accept"}
assert set(levels) == {"strict", "moderate", "off"}
assert set(actions) == {"error", "warn"}
print(" validationLevel:")
for k, v in levels.items():
print(f" {k:9s} {v}")
print(" validationAction:")
for k, v in actions.items():
print(f" {k:9s} {v}")
print(" migration: moderate+warn -> find offenders -> fix -> strict+error")
print(" going straight to strict+error breaks every update to a")
print(" non-conforming document, INCLUDING the one that would fix it")
def main():
print("Experiment 12 -- Schema validation")
print(" NOTE: mongomock does not enforce $jsonSchema, so the same rules")
print(" are implemented in code here and asserted.")
# Step 1: Pass a conforming document
conforming_document_passes()
# Step 2: Catch each violation
each_violation_is_caught()
# Step 3: Find the documents that do not conform
find_the_offenders()
# Step 4: Tighten validation in stages
the_migration_path()
if __name__ == "__main__":
main()
In mongosh, 12_validation.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.validated.drop()
true
collegeDB> const schema = {
... bsonType: "object",
... required: ["roll", "name", "dept"],
... properties: {
... roll: { bsonType: "int", minimum: 1,
... description: "required integer, at least 1" },
... name: { bsonType: "string", minLength: 3, maxLength: 80 },
... dept: { enum: ["DS", "Stats", "CS"],
... description: "must be DS, Stats or CS" },
... marks: {
... bsonType: "object",
... properties: {
... maths: { bsonType: "int", minimum: 0, maximum: 100 },
... stats: { bsonType: "int", minimum: 0, maximum: 100 }
... }
... },
... email: { bsonType: "string", pattern: "^[^\\s@]+@[^\\s@]+\\.[^\\s@]{2,}$" }
... }
... }
collegeDB> db.createCollection("validated", {
... validator: { $jsonSchema: schema },
... validationLevel: "strict",
... validationAction: "error"
... })
{ ok: 1 }
collegeDB> db.validated.insertOne({ roll: NumberInt(21), name: "Asha", dept: "DS",
... marks: { maths: NumberInt(88) } }) // OK
{
acknowledged: true,
insertedId: ObjectId('6ac214dd5e9ae24277f97aaa')
}
collegeDB> db.validated.insertOne({ name: "NoRoll", dept: "DS" }) // missing required
Uncaught:
MongoServerError: Document failed validation
Additional information: {
failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aab'),
details: {
operatorName: '$jsonSchema',
schemaRulesNotSatisfied: [
{
operatorName: 'required',
specifiedAs: { required: [ 'roll', 'name', 'dept' ] },
missingProperties: [ 'roll' ]
}
]
}
}
collegeDB> db.validated.insertOne({ roll: NumberInt(22), name: "Ab", dept: "DS" }) // too short
Uncaught:
MongoServerError: Document failed validation
Additional information: {
failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aac'),
details: {
operatorName: '$jsonSchema',
schemaRulesNotSatisfied: [
{
operatorName: 'properties',
propertiesNotSatisfied: [
{
propertyName: 'name',
details: [
{
operatorName: 'minLength',
specifiedAs: { minLength: 3 },
reason: 'specified string length was not satisfied',
consideredValue: 'Ab'
}
]
}
]
}
]
}
}
collegeDB> db.validated.insertOne({ roll: NumberInt(23), name: "Ravi", dept: "Physics" }) // enum
Uncaught:
MongoServerError: Document failed validation
Additional information: {
failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aad'),
details: {
operatorName: '$jsonSchema',
schemaRulesNotSatisfied: [
{
operatorName: 'properties',
propertiesNotSatisfied: [
{
propertyName: 'dept',
description: 'must be DS, Stats or CS',
details: [
{
operatorName: 'enum',
specifiedAs: { enum: [ 'DS', 'Stats', 'CS' ] },
reason: 'value was not found in enum',
consideredValue: 'Physics'
}
]
}
]
}
]
}
}
collegeDB> db.validated.insertOne({ roll: NumberInt(24), name: "Meena", dept: "DS",
... marks: { maths: NumberInt(150) } }) // out of range
Uncaught:
MongoServerError: Document failed validation
Additional information: {
failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aae'),
details: {
operatorName: '$jsonSchema',
schemaRulesNotSatisfied: [
{
operatorName: 'properties',
propertiesNotSatisfied: [
{
propertyName: 'marks',
details: [
{
operatorName: 'properties',
propertiesNotSatisfied: [
{
propertyName: 'maths',
details: [
{
operatorName: 'maximum',
specifiedAs: { maximum: 100 },
reason: 'comparison failed',
consideredValue: 150
}
]
}
]
}
]
}
]
}
]
}
}
collegeDB> db.runCommand({ collMod: "students",
... validator: { $jsonSchema: schema },
... validationLevel: "moderate",
... validationAction: "warn" })
{ ok: 1 }
collegeDB> db.students.find({ $nor: [ { $jsonSchema: schema } ] })
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: [ 'Stats' ],
age: 21,
active: true
}
]
collegeDB> db.runCommand({ collMod: "students",
... validationLevel: "strict", validationAction: "error" })
{ ok: 1 }
In Python, through mongomock, 12_validation.py:
OUTPUT
Experiment 12 -- Schema validation
NOTE: mongomock does not enforce $jsonSchema, so the same rules
are implemented in code here and asserted.
a conforming document produces no violations
missing required -> roll: required field missing
name too short -> name: shorter than 3
dept not in enum -> dept: 'Physics' not in ['DS', 'Stats', 'CS']
marks over 100 -> marks.maths: 150 above maximum 100
roll below 1 -> roll: 0 below minimum 1
roll wrong type -> roll: expected int, got str
bad email -> email: does not match ^[^\s@]+@[^\s@]+\.[^\s@]{2,}$
every rule in the schema is enforced
3 of 4 documents violate the schema:
Ab name: shorter than 3
NoRoll roll: required field missing
Kiran dept: 'Physics' not in ['DS', 'Stats', 'CS']
on a real server: $nor: [ { $jsonSchema: ... } ] finds exactly these
validationLevel:
strict applies to EVERY insert and update
moderate applies to inserts, and to updates of documents that ALREADY conform
off applies to nothing
validationAction:
error REJECT the write
warn LOG it and accept
migration: moderate+warn -> find offenders -> fix -> strict+error
going straight to strict+error breaks every update to a
non-conforming document, INCLUDING the one that would fix it
Corrected: the second part read { $jsonSchema: { /* as above */ } }, an empty
schema, which every document passes, so no offender could be found. The schema is now a const,
written once and used three times.
RESULT
The conforming document goes in and each of the four violations is refused, by MongoDB itself; all five sample students fail the schema, having no roll.
Create single-field and compound indexes, and measure their effect.
Create indexes, read explain(), and apply the prefix and ESR rules.
In mongosh, 13_indexes.js:
In Python, through mongomock, 13_indexes.py:
THE POINT
Run explain("executionStats") and read totalDocsExamined / nReturned — that
ratio, not the wall-clock time, is what tells you whether the index is right. mongomock records
indexes but has no planner, so the Python half asserts that the indexes are created and
listed correctly and that a unique index rejects a duplicate. The prefix rule and
ESR are demonstrated as a table of which queries each index serves; the server shows the
plans.
In mongosh, 13_indexes.js:
// Experiment 13 -- Creating and testing single-field and compound indexes.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 13_indexes.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: Create and list indexes
db.students.createIndex({ dept: 1 })
db.students.createIndex({ dept: 1, "marks.maths": -1 })
db.students.createIndex({ email: 1 }, { unique: true }) // FAILS here: see below
db.students.createIndex({ dept: 1 }, { name: "dept_idx" }) // FAILS: see below
// Both fail, and both failures are worth knowing. None of the five students
// has an email, so all five index as email: null -- five duplicates of null.
// And { dept: 1 } is already indexed, as dept_1: the same keys under a second
// name are refused. [Note added: these two lines were written expecting both
// to succeed. Run, they fail, as above.]
db.students.getIndexes()
db.students.totalIndexSize()
// Step 3: Measure with explain()
db.students.find({ dept: "DS" }).explain("executionStats")
// stage: COLLSCAN (bad) vs IXSCAN (good)
// totalDocsExamined / nReturned: 1 is ideal, 1000 means the index is wrong
// Step 4: Apply the prefix rule
db.students.createIndex({ dept: 1, year: 1, cgpa: 1 })
db.students.find({ dept: "DS" }) // uses it
db.students.find({ dept: "DS", year: 4 }) // uses it
db.students.find({ year: 4 }) // does NOT -- COLLSCAN
db.students.createIndex({ year: 1 }) // so this is needed too
// Step 5: Apply the ESR rule
// Query: dept = "DS", maths > 70, sorted by age
db.students.createIndex({ dept: 1, age: 1, "marks.maths": 1 })
// ^equality ^sort ^range
// A range predicate leaves everything AFTER it unordered, so a sort field
// placed after a range field cannot use the index.
// Step 6: Cover a query
db.students.createIndex({ dept: 1, name: 1 })
db.students.find({ dept: "DS" }, { _id: 0, dept: 1, name: 1 }).explain("executionStats")
// totalDocsExamined: 0
// [Corrected: the .explain(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find() without it, then rejected ".explain(...)" as an invalid command.]
// Step 7: See two missing fields collide in a unique index
db.people.drop()
db.people.createIndex({ email: 1 }, { unique: true })
db.people.insertOne({ name: "A" }) // ok -- email missing, indexed as null
db.people.insertOne({ name: "B" }) // DUPLICATE KEY ERROR -- a second null
db.people.dropIndex("email_1")
db.people.createIndex({ email: 1 },
{ unique: true, partialFilterExpression: { email: { $exists: true } } })
db.people.insertOne({ name: "B" }) // ok now: missing emails are not indexed
// [Corrected: this used db.students, whose five students already have no email,
// so the unique index could not be built at all, and the insert of B, commented
// DUPLICATE KEY ERROR, succeeded. A new collection, as in 13_indexes.py,
// shows what the comment says.]
db.students.dropIndex("dept_1")
In Python, through mongomock, 13_indexes.py:
"""Experiment 13 — Single-field and compound indexes.
mongomock records indexes but does not report IXSCAN, so this script asserts
that indexes are CREATED and ENFORCED correctly, and demonstrates the prefix
rule and ESR as tables. Run explain("executionStats") on a real server to see
the plan.
"""
import mongomock
from pymongo.errors import DuplicateKeyError
from fixtures import fresh_db
def creating_and_listing():
db = fresh_db()
db.students.create_index("dept")
db.students.create_index([("dept", 1), ("marks.maths", -1)])
db.students.create_index("age", name="age_idx")
names = {i["name"] for i in db.students.list_indexes()}
assert "_id_" in names, "_id is indexed AUTOMATICALLY"
assert "dept_1" in names
assert "age_idx" in names, "a custom name"
assert any("marks.maths" in n for n in names), names
db.students.drop_index("dept_1")
assert "dept_1" not in {i["name"] for i in db.students.list_indexes()}
print(f" created and listed {len(names)} indexes, including the automatic _id")
def unique_is_enforced():
db = fresh_db()
db.students.create_index("name", unique=True)
try:
db.students.insert_one({"_id": 99, "name": "Asha"})
raise AssertionError("expected a DuplicateKeyError")
except DuplicateKeyError:
pass
db.students.insert_one({"_id": 99, "name": "Unique Name"})
assert db.students.count_documents({}) == 6
print(" a unique index rejected a duplicate name and accepted a new one")
def unique_and_missing_fields():
"""A missing field indexes as null, and two nulls collide."""
db = mongomock.MongoClient().collegeDB
db.people.create_index("email", unique=True)
db.people.insert_one({"_id": 1, "name": "A"}) # no email -> null
try:
db.people.insert_one({"_id": 2, "name": "B"}) # also null
raise AssertionError("expected a DuplicateKeyError from two nulls")
except DuplicateKeyError:
pass
print(" a unique index allowed ONE document with no email, then rejected")
print(" the next -- a missing field indexes as null, and nulls collide")
print(" fix: partialFilterExpression: { email: { $exists: true } }")
def the_prefix_rule():
"""An index on {a,b,c} serves left-hand PREFIXES only."""
index = ["dept", "year", "cgpa"]
cases = [
(["dept"], True, "a prefix"),
(["dept", "year"], True, "a prefix"),
(["dept", "year", "cgpa"], True, "the whole index"),
(["year"], False, "not a prefix -- COLLSCAN"),
(["cgpa"], False, "not a prefix -- COLLSCAN"),
(["year", "cgpa"], False, "not a prefix -- COLLSCAN"),
]
def is_prefix(fields):
return index[:len(fields)] == fields
print(f" index {{{', '.join(index)}}}:")
for fields, expected, why in cases:
assert is_prefix(fields) == expected, (fields, expected)
mark = "uses it " if expected else "does NOT"
print(f" query on {str(fields):28s} {mark} ({why})")
print(" the phone book, sorted by (surname, forename): finding every")
print(" Kumari is fast; finding every Asha means reading the whole book")
def the_esr_rule():
"""Equality, Sort, Range."""
query = {"equality": "dept", "sort": "age", "range": "marks.maths"}
correct = ["dept", "age", "marks.maths"]
wrong = ["dept", "marks.maths", "age"]
assert correct.index("age") < correct.index("marks.maths"), \
"the SORT field must come before the RANGE field"
assert wrong.index("marks.maths") < wrong.index("age"), \
"this ordering puts the range first, and the sort cannot use the index"
print(f" ESR: query is dept = 'DS', maths > 70, sorted by age")
print(f" correct: {{{', '.join(correct)}}} E, S, R")
print(f" wrong: {{{', '.join(wrong)}}} the range leaves everything")
print(f" after it unordered, so the sort falls back to memory")
def covered_query_fields():
"""Every field in the filter AND the projection must be in the index."""
index = {"dept", "name"}
filt = {"dept"}
proj_bad = {"dept", "name", "_id"} # _id is returned BY DEFAULT
proj_good = {"dept", "name"} # with _id: 0
assert not (filt | proj_bad) <= index, "_id is not in the index -- NOT covered"
assert (filt | proj_good) <= index, "with _id: 0 it IS covered"
print(" covered query needs filter + projection inside the index")
print(" find({dept}, {dept:1, name:1}) NOT covered -- _id sneaks in")
print(" find({dept}, {dept:1, name:1, _id:0}) COVERED, totalDocsExamined 0")
def what_indexes_cost():
db = fresh_db()
for field in ["dept", "age", "active", "name"]:
db.students.create_index(field)
n = len(list(db.students.list_indexes()))
assert n == 5, f"4 plus _id, got {n}"
print(f" {n} indexes means every insert, update and delete maintains {n}")
print(f" B-trees. Index what you QUERY, not everything -- the limit")
print(f" is 64 per collection, and reaching it means something is wrong")
def main():
print("Experiment 13 -- Indexes")
# Step 1: Create and list indexes
creating_and_listing()
# Step 2: See a unique index enforced
unique_is_enforced()
# Step 3: See two missing fields collide
unique_and_missing_fields()
# Step 4: Apply the prefix rule
the_prefix_rule()
# Step 5: Apply the ESR rule
the_esr_rule()
# Step 6: Cover a query
covered_query_fields()
# Step 7: Count what indexes cost
what_indexes_cost()
if __name__ == "__main__":
main()
In mongosh, 13_indexes.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.createIndex({ dept: 1 })
dept_1
collegeDB> db.students.createIndex({ dept: 1, "marks.maths": -1 })
dept_1_marks.maths_-1
collegeDB> db.students.createIndex({ email: 1 }, { unique: true }) // FAILS here: see below
Uncaught
MongoServerError[DuplicateKey]: Index build failed: 594da0ca-9e60-4781-be71-2b6fdab31bff: Collection collegeDB.students ( c4e0fefe-1c54-4b76-9c72-80d472ae564b ) :: caused by :: E11000 duplicate key error collection: collegeDB.students index: email_1 dup key: { email: null }
collegeDB> db.students.createIndex({ dept: 1 }, { name: "dept_idx" }) // FAILS: see below
Uncaught
MongoServerError[IndexOptionsConflict]: Index already exists with a different name: dept_1
collegeDB> db.students.getIndexes()
[
{ v: 2, key: { _id: 1 }, name: '_id_' },
{ v: 2, key: { dept: 1 }, name: 'dept_1' },
{
v: 2,
key: { dept: 1, 'marks.maths': -1 },
name: 'dept_1_marks.maths_-1'
}
]
collegeDB> db.students.totalIndexSize()
45056
collegeDB> db.students.find({ dept: "DS" }).explain("executionStats")
{
explainVersion: '1',
queryPlanner: {
namespace: 'collegeDB.students',
parsedQuery: { dept: { '$eq': 'DS' } },
indexFilterSet: false,
queryHash: '8BDC9605',
planCacheShapeHash: '8BDC9605',
planCacheKey: 'F7C996A0',
optimizationTimeMillis: 0,
maxIndexedOrSolutionsReached: false,
maxIndexedAndSolutionsReached: false,
maxScansToExplodeReached: false,
prunedSimilarIndexes: false,
winningPlan: {
isCached: false,
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1 },
indexName: 'dept_1',
isMultiKey: false,
multiKeyPaths: { dept: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: { dept: [ '["DS", "DS"]' ] }
}
},
rejectedPlans: [
{
isCached: false,
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1, 'marks.maths': -1 },
indexName: 'dept_1_marks.maths_-1',
isMultiKey: false,
multiKeyPaths: { dept: [], 'marks.maths': [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: {
dept: [ '["DS", "DS"]' ],
'marks.maths': [ '[MaxKey, MinKey]' ]
}
}
}
]
},
executionStats: {
executionSuccess: true,
nReturned: 3,
executionTimeMillis: 0,
totalKeysExamined: 3,
totalDocsExamined: 3,
executionStages: {
isCached: false,
stage: 'FETCH',
nReturned: 3,
executionTimeMillisEstimate: 0,
works: 5,
advanced: 3,
needTime: 0,
needYield: 0,
saveState: 1,
restoreState: 1,
isEOF: 1,
nss: 'collegeDB.students',
docsExamined: 3,
alreadyHasObj: 0,
inputStage: {
stage: 'IXSCAN',
nReturned: 3,
executionTimeMillisEstimate: 0,
works: 4,
advanced: 3,
needTime: 0,
needYield: 0,
saveState: 1,
restoreState: 1,
isEOF: 1,
nss: 'collegeDB.students',
keyPattern: { dept: 1 },
indexName: 'dept_1',
isMultiKey: false,
multiKeyPaths: { dept: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: { dept: [ '["DS", "DS"]' ] },
keysExamined: 3,
seeks: 1,
dupsTested: 0,
dupsDropped: 0,
peakTrackedMemBytes: 0
}
}
},
queryShapeHash: '54C4AC01306DF576294D4F78A437455A0072E3A810EF2AF15CCC76DEB4F16CEE',
command: { find: 'students', filter: { dept: 'DS' }, '$db': 'collegeDB' },
serverInfo: {
host: 'vm',
port: 27017,
version: '8.3.7',
gitVersion: 'nogitversion'
},
serverParameters: {
internalQueryFacetBufferSizeBytes: 104857600,
internalDocumentSourceGroupMaxMemoryBytes: 104857600,
internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
internalQueryFacetMaxOutputDocSizeBytes: 104857600,
internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
internalQueryProhibitBlockingMergeOnMongoS: 0,
internalQueryMaxAddToSetBytes: 104857600,
internalQueryFrameworkControl: 'trySbeRestricted',
internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
},
ok: 1
}
collegeDB> db.students.createIndex({ dept: 1, year: 1, cgpa: 1 })
dept_1_year_1_cgpa_1
collegeDB> db.students.find({ dept: "DS" }) // uses it
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find({ dept: "DS", year: 4 }) // uses it
collegeDB> db.students.find({ year: 4 }) // does NOT -- COLLSCAN
collegeDB> db.students.createIndex({ year: 1 }) // so this is needed too
year_1
collegeDB> db.students.createIndex({ dept: 1, age: 1, "marks.maths": 1 })
dept_1_age_1_marks.maths_1
collegeDB> db.students.createIndex({ dept: 1, name: 1 })
dept_1_name_1
collegeDB> db.students.find({ dept: "DS" }, { _id: 0, dept: 1, name: 1 }).explain("executionStats")
{
explainVersion: '1',
queryPlanner: {
namespace: 'collegeDB.students',
parsedQuery: { dept: { '$eq': 'DS' } },
indexFilterSet: false,
queryHash: '1F033464',
planCacheShapeHash: '1F033464',
planCacheKey: '4D5148B1',
optimizationTimeMillis: 1,
maxIndexedOrSolutionsReached: false,
maxIndexedAndSolutionsReached: false,
maxScansToExplodeReached: false,
prunedSimilarIndexes: false,
winningPlan: {
isCached: false,
stage: 'PROJECTION_COVERED',
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1, name: 1 },
indexName: 'dept_1_name_1',
isMultiKey: false,
multiKeyPaths: { dept: [], name: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: { dept: [ '["DS", "DS"]' ], name: [ '[MinKey, MaxKey]' ] }
}
},
rejectedPlans: [
{
isCached: false,
stage: 'PROJECTION_SIMPLE',
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1, 'marks.maths': -1 },
indexName: 'dept_1_marks.maths_-1',
isMultiKey: false,
multiKeyPaths: { dept: [], 'marks.maths': [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: {
dept: [ '["DS", "DS"]' ],
'marks.maths': [ '[MaxKey, MinKey]' ]
}
}
}
},
{
isCached: false,
stage: 'PROJECTION_SIMPLE',
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1, year: 1, cgpa: 1 },
indexName: 'dept_1_year_1_cgpa_1',
isMultiKey: false,
multiKeyPaths: { dept: [], year: [], cgpa: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: {
dept: [ '["DS", "DS"]' ],
year: [ '[MinKey, MaxKey]' ],
cgpa: [ '[MinKey, MaxKey]' ]
}
}
}
},
{
isCached: false,
stage: 'PROJECTION_SIMPLE',
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1, age: 1, 'marks.maths': 1 },
indexName: 'dept_1_age_1_marks.maths_1',
isMultiKey: false,
multiKeyPaths: { dept: [], age: [], 'marks.maths': [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: {
dept: [ '["DS", "DS"]' ],
age: [ '[MinKey, MaxKey]' ],
'marks.maths': [ '[MinKey, MaxKey]' ]
}
}
}
},
{
isCached: false,
stage: 'PROJECTION_SIMPLE',
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'FETCH',
nss: 'collegeDB.students',
inputStage: {
stage: 'IXSCAN',
nss: 'collegeDB.students',
keyPattern: { dept: 1 },
indexName: 'dept_1',
isMultiKey: false,
multiKeyPaths: { dept: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: { dept: [ '["DS", "DS"]' ] }
}
}
}
]
},
executionStats: {
executionSuccess: true,
nReturned: 3,
executionTimeMillis: 2,
totalKeysExamined: 3,
totalDocsExamined: 0,
executionStages: {
isCached: false,
stage: 'PROJECTION_COVERED',
nReturned: 3,
executionTimeMillisEstimate: 0,
works: 5,
advanced: 3,
needTime: 0,
needYield: 0,
saveState: 1,
restoreState: 1,
isEOF: 1,
transformBy: { _id: 0, dept: 1, name: 1 },
inputStage: {
stage: 'IXSCAN',
nReturned: 3,
executionTimeMillisEstimate: 0,
works: 5,
advanced: 3,
needTime: 0,
needYield: 0,
saveState: 1,
restoreState: 1,
isEOF: 1,
nss: 'collegeDB.students',
keyPattern: { dept: 1, name: 1 },
indexName: 'dept_1_name_1',
isMultiKey: false,
multiKeyPaths: { dept: [], name: [] },
isUnique: false,
isSparse: false,
isPartial: false,
indexVersion: 2,
direction: 'forward',
indexBounds: { dept: [ '["DS", "DS"]' ], name: [ '[MinKey, MaxKey]' ] },
keysExamined: 3,
seeks: 1,
dupsTested: 0,
dupsDropped: 0,
peakTrackedMemBytes: 0
}
}
},
queryShapeHash: '3702523BB00E9C352D45FA2B03292BA7E3DDDC04DB9C14A06C726E7CD9CF4AC9',
command: {
find: 'students',
filter: { dept: 'DS' },
projection: { _id: 0, dept: 1, name: 1 },
'$db': 'collegeDB'
},
serverInfo: {
host: 'vm',
port: 27017,
version: '8.3.7',
gitVersion: 'nogitversion'
},
serverParameters: {
internalQueryFacetBufferSizeBytes: 104857600,
internalDocumentSourceGroupMaxMemoryBytes: 104857600,
internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
internalQueryFacetMaxOutputDocSizeBytes: 104857600,
internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
internalQueryProhibitBlockingMergeOnMongoS: 0,
internalQueryMaxAddToSetBytes: 104857600,
internalQueryFrameworkControl: 'trySbeRestricted',
internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
},
ok: 1
}
collegeDB> db.people.drop()
true
collegeDB> db.people.createIndex({ email: 1 }, { unique: true })
email_1
collegeDB> db.people.insertOne({ name: "A" }) // ok -- email missing, indexed as null
{
acknowledged: true,
insertedId: ObjectId('6ac214e2f47d929cc4821b4c')
}
collegeDB> db.people.insertOne({ name: "B" }) // DUPLICATE KEY ERROR -- a second null
Uncaught
MongoServerError: E11000 duplicate key error collection: collegeDB.people index: email_1 dup key: { email: null }
collegeDB> db.people.dropIndex("email_1")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.people.createIndex({ email: 1 },
... { unique: true, partialFilterExpression: { email: { $exists: true } } })
email_1
collegeDB> db.people.insertOne({ name: "B" }) // ok now: missing emails are not indexed
{
acknowledged: true,
insertedId: ObjectId('6ac214e3f47d929cc4821b4e')
}
collegeDB> db.students.dropIndex("dept_1")
{ nIndexesWas: 7, ok: 1 }
In Python, through mongomock, 13_indexes.py:
OUTPUT
Experiment 13 -- Indexes
created and listed 4 indexes, including the automatic _id
a unique index rejected a duplicate name and accepted a new one
a unique index allowed ONE document with no email, then rejected
the next -- a missing field indexes as null, and nulls collide
fix: partialFilterExpression: { email: { $exists: true } }
index {dept, year, cgpa}:
query on ['dept'] uses it (a prefix)
query on ['dept', 'year'] uses it (a prefix)
query on ['dept', 'year', 'cgpa'] uses it (the whole index)
query on ['year'] does NOT (not a prefix -- COLLSCAN)
query on ['cgpa'] does NOT (not a prefix -- COLLSCAN)
query on ['year', 'cgpa'] does NOT (not a prefix -- COLLSCAN)
the phone book, sorted by (surname, forename): finding every
Kumari is fast; finding every Asha means reading the whole book
ESR: query is dept = 'DS', maths > 70, sorted by age
correct: {dept, age, marks.maths} E, S, R
wrong: {dept, marks.maths, age} the range leaves everything
after it unordered, so the sort falls back to memory
covered query needs filter + projection inside the index
find({dept}, {dept:1, name:1}) NOT covered -- _id sneaks in
find({dept}, {dept:1, name:1, _id:0}) COVERED, totalDocsExamined 0
5 indexes means every insert, update and delete maintains 5
B-trees. Index what you QUERY, not everything -- the limit
is 64 per collection, and reaching it means something is wrong
The covered query's plan is PROJECTION_COVERED with totalDocsExamined: 0, as its
comment says.
Corrected or noted, from running it: the unique index on email and the dept_idx index fail at
the top, and the file now says why — none of the five students has an email, so all five index
as null, and { dept: 1 } already exists as dept_1. The demonstration that two missing fields
collide used db.students, where the unique index could not be built at all, so the insert
commented DUPLICATE KEY ERROR succeeded; it now uses a new collection, as the Python half does.
And .explain(...) began its own line, which mongosh does not join to the line above.
RESULT
find({ dept: "DS" }) is an IXSCAN examining 3 documents for 3 returned; the covered query examines none.
Search text with a text index, and index an array with a multikey index.
Index arrays and text, and search by words with relevance.
In mongosh, 14_text_multikey.js:
In Python, through mongomock, 14_text_multikey.py:
THE POINT
An index on an array field creates one entry per element, so { subjects: "DS" }
matches any student whose array contains it. Only one text index is allowed per collection
— it may span several fields, but you cannot have two.
mongomock does not implement $text, so the Python half asserts the multikey behaviour and works
out by hand what a server's text search returns; the session shows the server's.
In mongosh, 14_text_multikey.js:
// Experiment 14 -- Text search and multikey indexes.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 14_text_multikey.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// =============================================================================
// Step 2: Index an array, which makes it multikey
// PART A -- MULTIKEY INDEXES (an index on an array field)
// =============================================================================
// There is no "createMultikeyIndex". You index the field, and MongoDB makes
// the index multikey BY ITSELF the moment it meets an array value.
db.students.createIndex({ subjects: 1 })
db.students.find({ subjects: "DS" }) // matches if the ARRAY CONTAINS it
db.students.find({ subjects: { $all: ["DS", "Python"] } }) // contains BOTH
db.students.find({ subjects: { $size: 3 } }) // exactly three -- NOT indexed
db.students.find({ "subjects.0": "DS" }) // DS is the FIRST element
// One index ENTRY per array element. A student with 3 subjects contributes 3
// entries pointing at the same document, which is why multikey indexes are
// larger than they look, and why an array of 1,000 elements is a bad idea.
// Step 3: Meet the restrictions
// 1. A compound index may contain AT MOST ONE array field.
db.students.createIndex({ subjects: 1, dept: 1 }) // OK -- one array
// db.students.createIndex({ subjects: 1, tags: 1 }) // ERROR if BOTH arrays
// 2. A multikey index cannot be a shard key.
// 3. $size is never served by an index -- it must scan. Store a length field
// alongside the array if you need to query on it:
db.students.updateMany({}, [ { $set: { nSubjects: { $size: "$subjects" } } } ])
db.students.createIndex({ nSubjects: 1 })
// Step 4: Index an array of sub-documents, and fall into the trap
db.transcripts.drop()
db.transcripts.insertOne({ _id: 21, name: "Asha", enrollments: [
{ course: "DSC301", grade: "B" }, { course: "STA302", grade: "A" } ] })
db.transcripts.createIndex({ "enrollments.grade": 1 }) // also multikey
// The trap from experiment 9, restated: without $elemMatch the two conditions
// may be satisfied by DIFFERENT elements of the array.
db.transcripts.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" }) // Asha -- wrongly
db.transcripts.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } }) // nobody
// [Corrected: these queried db.students, whose documents have no enrollments,
// so both found nothing and the trap did not show. Asha's B in DSC301 and A in
// STA302 are two different elements, and only $elemMatch tells them apart.]
// =============================================================================
// Step 5: Create a text index
// PART B -- TEXT INDEXES
// =============================================================================
db.articles.drop()
db.articles.insertMany([
{ _id: 1, title: "Introduction to MongoDB",
body: "MongoDB is a document database that stores data in BSON." },
{ _id: 2, title: "Aggregation pipelines explained",
body: "The aggregation framework processes documents through stages." },
{ _id: 3, title: "Indexing strategy in MongoDB",
body: "An index is a B-tree. Aggregation queries benefit from indexes too." },
{ _id: 4, title: "Relational databases",
body: "SQL databases use tables, rows and joins." }
])
// Weights make a hit in the title count ten times a hit in the body.
db.articles.createIndex({ title: "text", body: "text" },
{ weights: { title: 10, body: 1 },
name: "article_text",
default_language: "english" })
// Step 6: Search
db.articles.find({ $text: { $search: "mongodb" } })
db.articles.find({ $text: { $search: "mongodb aggregation" } }) // OR, not AND
db.articles.find({ $text: { $search: "\"aggregation framework\"" } }) // PHRASE
db.articles.find({ $text: { $search: "mongodb -relational" } }) // EXCLUDE
// Step 7: Rank by relevance
db.articles.find({ $text: { $search: "mongodb aggregation" } },
{ score: { $meta: "textScore" }, title: 1 }).sort({ score: { $meta: "textScore" } })
// [Corrected: the .sort(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find(), unsorted, without it, then rejected ".sort(...)" as an invalid command.]
// The sort is NOT optional. $text returns matches in no particular order; the
// score exists only if you project it, and only sorts if you sort by it.
// --- the rules, all examinable ----------------------------------------------
// 1. ONE text index per collection. It may span many fields -- even every
// field, via { "$**": "text" } -- but you cannot have two.
db.articles.dropIndex("article_text")
db.articles.createIndex({ "$**": "text" }) // a WILDCARD text index
db.articles.dropIndex("$**_text")
db.articles.createIndex({ title: "text", body: "text" },
{ weights: { title: 10, body: 1 }, name: "article_text" })
// 2. $text searches WORDS, not substrings. "mongo" does not match "MongoDB".
// For substrings and prefixes you need a regex, or Atlas Search.
// 3. Search is case-insensitive and diacritic-insensitive by default.
// 4. Stemming and stop words follow default_language: searching "stores" also
// matches "store" and "storing"; "the" and "is" are ignored entirely.
// 5. Only ONE $text expression per query, and it cannot appear inside $or with
// a non-text clause.
db.articles.getIndexes()
In Python, through mongomock, 14_text_multikey.py:
"""Experiment 14 — Text search and multikey indexes.
Two halves, and they run differently.
MULTIKEY is fully executed: mongomock matches array fields the way a real
server does, so every assertion here is a real one.
TEXT SEARCH is not. mongomock accepts createIndex([("x", "text")]) but raises
NotImplementedError on $text -- that is asserted below rather than glossed
over, and the ranking and stemming rules are then set out as data. The runnable
substitute is a regex scan, which is also the honest answer to "what do I do
when I have no text index?".
"""
import mongomock
from fixtures import fresh_db, names
# =============================================================================
# PART A -- multikey indexes: fully executed
# =============================================================================
def an_index_on_an_array_is_multikey_automatically():
"""You never ask for a multikey index. MongoDB decides."""
db = fresh_db()
db.students.create_index("subjects")
# The index exists and looks like any other single-field index.
assert "subjects_1" in {i["name"] for i in db.students.list_indexes()}
# Matching is CONTAINS, not equals: no student's subjects field IS "DS".
assert names(db.students.find({"subjects": "DS"})) == ["Asha", "Kiran", "Ravi"]
assert names(db.students.find({"subjects": "R"})) == ["Meena"]
# And the whole array still matches as a whole, if you give it exactly.
assert names(db.students.find({"subjects": ["DS", "Python"]})) == ["Ravi"]
assert names(db.students.find({"subjects": ["Python", "DS"]})) == [], \
"the whole-array form is ORDER SENSITIVE; the contains form is not"
print(" { subjects: 'DS' } matched Asha, Kiran, Ravi -- CONTAINS, not equals")
print(" { subjects: ['Python','DS'] } matched nobody -- whole-array match")
print(" is order sensitive, and Ravi's array is ['DS','Python']")
def one_index_entry_per_element():
"""The cost model: a document with n array values costs n index entries."""
db = fresh_db()
entries = sum(len(d["subjects"]) for d in db.students.find())
docs = db.students.count_documents({})
assert docs == 5
assert entries == 9, entries # 3 + 2 + 2 + 1 + 1
print(f" {docs} documents -> {entries} index entries "
f"({entries / docs:.1f} per document)")
print(" an array of 1,000 elements means 1,000 entries for ONE")
print(" document -- multikey indexes are bigger than they look")
def all_and_size_and_positional():
db = fresh_db()
# $all -- contains ALL of these (order irrelevant)
assert names(db.students.find({"subjects": {"$all": ["DS", "Python"]}})) \
== ["Asha", "Ravi"]
# bare {a: x, ...} on an array is OR-ish across elements; $all is AND
assert names(db.students.find({"subjects": {"$all": ["Stats", "R"]}})) == ["Meena"]
# $size -- exact length. NEVER served by an index.
assert names(db.students.find({"subjects": {"$size": 3}})) == ["Asha"]
assert names(db.students.find({"subjects": {"$size": 1}})) == ["Bhanu", "Kiran"]
# positional: the FIRST element specifically
assert names(db.students.find({"subjects.0": "DS"})) == ["Asha", "Kiran", "Ravi"]
assert names(db.students.find({"subjects.0": "Stats"})) == ["Bhanu", "Meena"]
print(" $all ['DS','Python'] -> Asha, Ravi (contains BOTH)")
print(" $size 3 -> Asha (never uses the index)")
print(" subjects.0 'Stats' -> Bhanu, Meena (FIRST element only)")
def store_the_length_if_you_query_it():
"""The fix for $size's collection scan: a field you CAN index."""
db = fresh_db()
for d in db.students.find():
db.students.update_one({"_id": d["_id"]},
{"$set": {"nSubjects": len(d["subjects"])}})
db.students.create_index("nSubjects")
assert names(db.students.find({"nSubjects": 3})) == ["Asha"]
assert names(db.students.find({"nSubjects": {"$gte": 2}})) \
== ["Asha", "Meena", "Ravi"]
print(" nSubjects >= 2 -> Asha, Meena, Ravi -- and unlike $size this one is")
print(" indexable, and supports RANGES, which $size cannot express")
def the_elemmatch_trap_again():
"""Two conditions on an array of sub-documents. The classic wrong answer."""
db = fresh_db()
db.students.update_one(
{"_id": 21},
{"$set": {"enrollments": [{"course": "DSC301", "grade": "B"},
{"course": "STA302", "grade": "A"}]}})
db.students.create_index("enrollments.grade")
# WRONG: satisfied by two DIFFERENT elements -- Asha got B in DSC301.
wrong = names(db.students.find({"enrollments.course": "DSC301",
"enrollments.grade": "A"}))
assert wrong == ["Asha"], wrong
# RIGHT: both conditions on the SAME element.
right = names(db.students.find(
{"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))
assert right == [], right
print(" Asha: DSC301->B, STA302->A")
print(" without $elemMatch -> ['Asha'] WRONG (two different elements)")
print(" with $elemMatch -> [] RIGHT")
def the_compound_restriction():
"""At most ONE array field in a compound index. Stated, not executed."""
rules = [
("{ subjects: 1 }", "OK", "one array field"),
("{ subjects: 1, dept: 1 }", "OK", "one array, one scalar"),
("{ dept: 1, subjects: 1 }", "OK", "order does not change the rule"),
("{ subjects: 1, tags: 1 }", "ERROR",
"two array fields -- cannot compute the cross product"),
]
assert sum(1 for _, v, _ in rules if v == "ERROR") == 1
print(" compound indexes containing arrays:")
for spec, verdict, why in rules:
print(f" {spec:28s} {verdict:6s} {why}")
print(" the reason: indexing both would need EVERY pair, so a")
print(" document with 10 and 10 would need 100 index entries")
# =============================================================================
# PART B -- text search: mongomock cannot run it, so say so and prove it
# =============================================================================
ARTICLES = [
{"_id": 1, "title": "Introduction to MongoDB",
"body": "MongoDB is a document database that stores data in BSON."},
{"_id": 2, "title": "Aggregation pipelines explained",
"body": "The aggregation framework processes documents through stages."},
{"_id": 3, "title": "Indexing strategy in MongoDB",
"body": "An index is a B-tree. Aggregation queries benefit from indexes too."},
{"_id": 4, "title": "Relational databases",
"body": "SQL databases use tables, rows and joins."},
]
def text_search_is_not_implemented_here():
"""Asserted, so this file can never quietly start claiming to test $text."""
db = mongomock.MongoClient().collegeDB
db.articles.insert_many([dict(a) for a in ARTICLES])
db.articles.create_index([("title", "text"), ("body", "text")])
try:
list(db.articles.find({"$text": {"$search": "mongodb"}}))
raise SystemExit("mongomock now implements $text -- rewrite this file "
"to assert results instead of documenting them")
except NotImplementedError as exc:
message = str(exc)
assert "$text" in message, message
print(" mongomock accepted the text INDEX and then raised")
print(f" NotImplementedError: {message}")
print(" so nothing below is a test result -- it is documentation")
def what_a_real_server_would_return():
"""The expected results, worked out by hand from the four articles."""
expected = [
('"mongodb"', [1, 3], "the word, in title or body"),
('"mongodb aggregation"', [1, 2, 3], "OR of the two terms, NOT and"),
('"\\"aggregation framework\\""', [2], "a PHRASE -- adjacent words"),
('"mongodb -relational"', [1, 3], "- excludes; 4 never matched anyway"),
('"mongo"', [], "WORDS, not substrings"),
('"stores"', [1], "stemming: matches 'stores'/'storing'"),
('"the"', [], "a stop word -- ignored entirely"),
]
print(" on a real server, $text: { $search: ... } would return:")
for term, ids, why in expected:
print(f" {term:32s} -> {str(ids):10s} {why}")
print(" 'mongo' returning NOTHING is the one that catches people:")
print(" a text index stores WORDS. For prefixes use a regex anchored")
print(" with ^, or Atlas Search; a bare /mongo/ scans the collection")
def regex_is_the_runnable_substitute():
"""What you actually do without a text index -- and it IS executed."""
db = mongomock.MongoClient().collegeDB
db.articles.insert_many([dict(a) for a in ARTICLES])
hits = sorted(d["_id"] for d in
db.articles.find({"title": {"$regex": "mongo", "$options": "i"}}))
assert hits == [1, 3], hits
# And here is where regex BEATS $text: substrings.
sub = sorted(d["_id"] for d in
db.articles.find({"body": {"$regex": "aggregat", "$options": "i"}}))
assert sub == [2, 3], sub
# And where it loses: no stemming, no ranking.
stem = sorted(d["_id"] for d in
db.articles.find({"body": {"$regex": "store$", "$options": "i"}}))
assert stem == [], "regex has no stemming -- 'stores' does not match 'store$'"
print(" regex /mongo/i on title -> [1, 3] (substring: $text would miss)")
print(" regex /aggregat/i on body-> [2, 3] (substring again)")
print(" regex /store$/i on body -> [] (no stemming, no ranking,")
print(" and an unanchored regex cannot use an index -- COLLSCAN)")
def the_text_index_rules():
rules = [
("How many per collection?", "ONE",
"it may span many fields, but you cannot have two"),
("Every field?", '{ "$**": "text" }',
"a wildcard text index -- convenient, and large"),
("Weights", "{ title: 10, body: 1 }",
"a title hit scores ten times a body hit"),
("Getting the score", '{ $meta: "textScore" }',
"must be PROJECTED to exist"),
("Ordering by it", '.sort({ score: { $meta: "textScore" } })',
"$text does NOT sort by relevance on its own"),
("Case / accents", "insensitive by default", "both, unless configured"),
("Inside $or", "not with a non-text clause", "one $text per query"),
]
assert len(rules) == 7
print(" text index rules:")
for q, a, why in rules:
print(f" {q:24s} {a:34s} {why}")
def main():
print("Experiment 14 -- Text search and multikey indexes")
print(" PART A -- multikey: EXECUTED and asserted")
# Step 1: See an index on an array become multikey
an_index_on_an_array_is_multikey_automatically()
# Step 2: Count one entry per element
one_index_entry_per_element()
# Step 3: Use $all, $size and the positional operator
all_and_size_and_positional()
# Step 4: Store the length to query it
store_the_length_if_you_query_it()
# Step 5: Fall into the $elemMatch trap
the_elemmatch_trap_again()
# Step 6: State the compound restriction
the_compound_restriction()
print(" PART B -- text search: NOT executable here")
# Step 7: Note that mongomock has no $text
text_search_is_not_implemented_here()
# Step 8: Work out what a real server returns
what_a_real_server_would_return()
# Step 9: Search by regex instead
regex_is_the_runnable_substitute()
# Step 10: State the text index rules
the_text_index_rules()
if __name__ == "__main__":
main()
In mongosh, 14_text_multikey.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.createIndex({ subjects: 1 })
subjects_1
collegeDB> db.students.find({ subjects: "DS" }) // matches if the ARRAY CONTAINS it
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.find({ subjects: { $all: ["DS", "Python"] } }) // contains BOTH
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
}
]
collegeDB> db.students.find({ subjects: { $size: 3 } }) // exactly three -- NOT indexed
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
}
]
collegeDB> db.students.find({ "subjects.0": "DS" }) // DS is the FIRST element
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: [ 'DS', 'Python' ],
age: 21,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false
}
]
collegeDB> db.students.createIndex({ subjects: 1, dept: 1 }) // OK -- one array
subjects_1_dept_1
collegeDB> db.students.updateMany({}, [ { $set: { nSubjects: { $size: "$subjects" } } } ])
{
acknowledged: true,
insertedId: null,
matchedCount: 5,
modifiedCount: 5,
upsertedCount: 0
}
collegeDB> db.students.createIndex({ nSubjects: 1 })
nSubjects_1
collegeDB> db.transcripts.drop()
true
collegeDB> db.transcripts.insertOne({ _id: 21, name: "Asha", enrollments: [
... { course: "DSC301", grade: "B" }, { course: "STA302", grade: "A" } ] })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.transcripts.createIndex({ "enrollments.grade": 1 }) // also multikey
enrollments.grade_1
collegeDB> db.transcripts.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" }) // Asha -- wrongly
[
{
_id: 21,
name: 'Asha',
enrollments: [ { course: 'DSC301', grade: 'B' }, { course: 'STA302', grade: 'A' } ]
}
]
collegeDB> db.transcripts.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } }) // nobody
collegeDB> db.articles.drop()
true
collegeDB> db.articles.insertMany([
... { _id: 1, title: "Introduction to MongoDB",
... body: "MongoDB is a document database that stores data in BSON." },
... { _id: 2, title: "Aggregation pipelines explained",
... body: "The aggregation framework processes documents through stages." },
... { _id: 3, title: "Indexing strategy in MongoDB",
... body: "An index is a B-tree. Aggregation queries benefit from indexes too." },
... { _id: 4, title: "Relational databases",
... body: "SQL databases use tables, rows and joins." }
... ])
{ acknowledged: true, insertedIds: { '0': 1, '1': 2, '2': 3, '3': 4 } }
collegeDB> db.articles.createIndex({ title: "text", body: "text" },
... { weights: { title: 10, body: 1 },
... name: "article_text",
... default_language: "english" })
article_text
collegeDB> db.articles.find({ $text: { $search: "mongodb" } })
[
{
_id: 1,
title: 'Introduction to MongoDB',
body: 'MongoDB is a document database that stores data in BSON.'
},
{
_id: 3,
title: 'Indexing strategy in MongoDB',
body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
}
]
collegeDB> db.articles.find({ $text: { $search: "mongodb aggregation" } }) // OR, not AND
[
{
_id: 2,
title: 'Aggregation pipelines explained',
body: 'The aggregation framework processes documents through stages.'
},
{
_id: 3,
title: 'Indexing strategy in MongoDB',
body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
},
{
_id: 1,
title: 'Introduction to MongoDB',
body: 'MongoDB is a document database that stores data in BSON.'
}
]
collegeDB> db.articles.find({ $text: { $search: "\"aggregation framework\"" } }) // PHRASE
[
{
_id: 2,
title: 'Aggregation pipelines explained',
body: 'The aggregation framework processes documents through stages.'
}
]
collegeDB> db.articles.find({ $text: { $search: "mongodb -relational" } }) // EXCLUDE
[
{
_id: 1,
title: 'Introduction to MongoDB',
body: 'MongoDB is a document database that stores data in BSON.'
},
{
_id: 3,
title: 'Indexing strategy in MongoDB',
body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
}
]
collegeDB> db.articles.find({ $text: { $search: "mongodb aggregation" } },
... { score: { $meta: "textScore" }, title: 1 }).sort({ score: { $meta: "textScore" } })
[
{ _id: 1, title: 'Introduction to MongoDB', score: 8.083333333333334 },
{
_id: 2,
title: 'Aggregation pipelines explained',
score: 7.266666666666666
},
{
_id: 3,
title: 'Indexing strategy in MongoDB',
score: 7.238095238095237
}
]
collegeDB> db.articles.dropIndex("article_text")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.articles.createIndex({ "$**": "text" }) // a WILDCARD text index
$**_text
collegeDB> db.articles.dropIndex("$**_text")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.articles.createIndex({ title: "text", body: "text" },
... { weights: { title: 10, body: 1 }, name: "article_text" })
article_text
collegeDB> db.articles.getIndexes()
[
{ v: 2, key: { _id: 1 }, name: '_id_' },
{
v: 2,
key: { _fts: 'text', _ftsx: 1 },
name: 'article_text',
weights: { body: 1, title: 10 },
default_language: 'english',
language_override: 'language',
textIndexVersion: 3
}
]
In Python, through mongomock, 14_text_multikey.py:
OUTPUT
Experiment 14 -- Text search and multikey indexes
PART A -- multikey: EXECUTED and asserted
{ subjects: 'DS' } matched Asha, Kiran, Ravi -- CONTAINS, not equals
{ subjects: ['Python','DS'] } matched nobody -- whole-array match
is order sensitive, and Ravi's array is ['DS','Python']
5 documents -> 9 index entries (1.8 per document)
an array of 1,000 elements means 1,000 entries for ONE
document -- multikey indexes are bigger than they look
$all ['DS','Python'] -> Asha, Ravi (contains BOTH)
$size 3 -> Asha (never uses the index)
subjects.0 'Stats' -> Bhanu, Meena (FIRST element only)
nSubjects >= 2 -> Asha, Meena, Ravi -- and unlike $size this one is
indexable, and supports RANGES, which $size cannot express
Asha: DSC301->B, STA302->A
without $elemMatch -> ['Asha'] WRONG (two different elements)
with $elemMatch -> [] RIGHT
compound indexes containing arrays:
{ subjects: 1 } OK one array field
{ subjects: 1, dept: 1 } OK one array, one scalar
{ dept: 1, subjects: 1 } OK order does not change the rule
{ subjects: 1, tags: 1 } ERROR two array fields -- cannot compute the cross product
the reason: indexing both would need EVERY pair, so a
document with 10 and 10 would need 100 index entries
PART B -- text search: NOT executable here
mongomock accepted the text INDEX and then raised
NotImplementedError: The $text operator is not implemented in mongomock yet
so nothing below is a test result -- it is documentation
on a real server, $text: { $search: ... } would return:
"mongodb" -> [1, 3] the word, in title or body
"mongodb aggregation" -> [1, 2, 3] OR of the two terms, NOT and
"\"aggregation framework\"" -> [2] a PHRASE -- adjacent words
"mongodb -relational" -> [1, 3] - excludes; 4 never matched anyway
"mongo" -> [] WORDS, not substrings
"stores" -> [1] stemming: matches 'stores'/'storing'
"the" -> [] a stop word -- ignored entirely
'mongo' returning NOTHING is the one that catches people:
a text index stores WORDS. For prefixes use a regex anchored
with ^, or Atlas Search; a bare /mongo/ scans the collection
regex /mongo/i on title -> [1, 3] (substring: $text would miss)
regex /aggregat/i on body-> [2, 3] (substring again)
regex /store$/i on body -> [] (no stemming, no ranking,
and an unanchored regex cannot use an index -- COLLSCAN)
text index rules:
How many per collection? ONE it may span many fields, but you cannot have two
Every field? { "$**": "text" } a wildcard text index -- convenient, and large
Weights { title: 10, body: 1 } a title hit scores ten times a body hit
Getting the score { $meta: "textScore" } must be PROJECTED to exist
Ordering by it .sort({ score: { $meta: "textScore" } }) $text does NOT sort by relevance on its own
Case / accents insensitive by default both, unless configured
Inside $or not with a non-text clause one $text per query
Corrected: the $elemMatch trap queried db.students, whose documents have no
enrolments, so both queries found nothing; it now has a document to find, in which Asha has a B
in DSC301 and an A in STA302. And the ranking's .sort(...) began its own line, so the search ran
unsorted and the shell rejected the sort.
RESULT
The array index is multikey; $text matches whole words and ranks by score; without $elemMatch, Asha is found for an A in DSC301 she does not have.
Aggregate documents with $match, $group, $project and $sort.
Build an aggregation pipeline, and see each stage's SQL counterpart.
In mongosh, 15_aggregation.js:
In Python, through mongomock, 15_aggregation.py:
THE POINT
Shown twice, and the pair is the point. With the active: true
filter, Kiran (DS, 71, inactive) is excluded, so DS is (88+65)/2 = 76.5
over 2 and Stats (94+52)/2 = 73 over 2. Without it — Unit 4's
Problem 1(a) — DS is (88+65+71)/3 = 74.667 over 3. Same grouping, and the
$match before it moves the DS average up.
$match before and after $group are WHERE and HAVING — the same stage in different
positions. mongomock lacks $round and $stdDevPop, so the Python half computes those in Python
and says so; the server has both.
In mongosh, 15_aggregation.js:
// Experiment 15 -- The aggregation pipeline: $match, $group, $project, $sort.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 15_aggregation.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// Step 2: $match, $group, $sort: the whole pipeline
// SELECT dept, ROUND(AVG(maths),2) AS avg, COUNT(*) AS n
// FROM students
// WHERE active = true -- $match BEFORE $group
// GROUP BY dept
// HAVING COUNT(*) > 1 -- $match AFTER $group
// ORDER BY avg DESC;
db.students.aggregate([
{ $match: { active: true } },
{ $group: { _id: "$dept", avg: { $avg: "$marks.maths" },
n: { $sum: 1 } } },
{ $match: { n: { $gt: 1 } } },
{ $sort: { avg: -1 } },
{ $project: { _id: 0, dept: "$_id", avg: { $round: ["$avg", 2] }, n: 1 } }
])
// WHERE and HAVING are THE SAME STAGE in different positions. Say that in the
// viva; it is the sentence that shows you understand pipelines.
// Step 3: $group without a filter
db.students.aggregate([
{ $group: { _id: "$dept", avgMaths: { $avg: "$marks.maths" },
n: { $sum: 1 } } },
{ $sort: { _id: 1 } }
])
// _id is MANDATORY in $group. It is the grouping key. And $group returns its
// groups in NO fixed order -- sort them whenever the order is shown.
// [Changed: the $sort was added; the groups came back in different orders.]
db.students.aggregate([ { $group: { _id: null, avg: { $avg: "$marks.maths" },
n: { $sum: 1 } } } ])
// _id: null groups EVERYTHING into one bucket -- a grand total.
// Step 4: The accumulators
db.students.aggregate([
{ $group: {
_id: "$dept",
n: { $sum: 1 }, // COUNT(*)
totMaths: { $sum: "$marks.maths" }, // SUM
avgMaths: { $avg: "$marks.maths" }, // AVG
best: { $max: "$marks.maths" }, // MAX
worst: { $min: "$marks.maths" }, // MIN
sd: { $stdDevPop: "$marks.maths" }, // Course 4's population sd
everyone: { $push: "$name" }, // ALL values, as an array
distinct: { $addToSet: "$name" }, // DISTINCT values
anyone: { $first: "$name" } // needs a $sort to be meaningful
} },
{ $set: { distinct: { $sortArray: { input: "$distinct", sortBy: 1 } } } },
{ $sort: { _id: 1 } }
])
// $addToSet keeps NO order, so the $set sorts that array before it is shown.
// [Changed: the $set and the $sort were added. The distinct names came back
// in a different order from one run to the next.]
// $push and $addToSet have no SQL equivalent, and are the reason MongoDB does
// not need GROUP_CONCAT.
// Step 5: $project: include, exclude, compute, rename
db.students.aggregate([
{ $project: {
_id: 0,
name: 1, // include
total: { $add: ["$marks.maths", "$marks.stats"] },
pct: { $round: [ { $divide: [ { $add: ["$marks.maths", "$marks.stats"] },
2 ] }, 1 ] },
dept: "$dept", // rename by re-assigning
band: { $switch: { branches: [
{ case: { $gte: ["$marks.maths", 75] }, then: "Distinction" },
{ case: { $gte: ["$marks.maths", 60] }, then: "First" },
{ case: { $gte: ["$marks.maths", 40] }, then: "Pass" } ],
default: "Fail" } }
} }
])
// $addFields (alias: $set) keeps everything and adds -- usually what you meant.
db.students.aggregate([
{ $addFields: { total: { $add: ["$marks.maths", "$marks.stats"] } } },
{ $sort: { total: -1 } },
{ $limit: 3 }
])
// Step 6: Put $match first
db.students.aggregate([ { $group: { _id: "$dept", n: { $sum: 1 } } },
{ $match: { _id: "DS" } } ]) // SLOW: groups all
db.students.aggregate([ { $match: { dept: "DS" } },
{ $group: { _id: "$dept", n: { $sum: 1 } } } ]) // fast
// Only the second can use an index on dept. Once documents have flowed through
// $group they are NEW documents, and no index describes them.
// Step 7: See what each stage does
db.students.aggregate([ { $match: { active: true } },
{ $group: { _id: "$dept", n: { $sum: 1 } } } ],
{ explain: true })
// Or truncate the pipeline and run the prefix -- the fastest way to find the
// stage that emptied your result. Compass's Aggregations tab does this for you.
// Step 8: Send the result to a collection
db.students.aggregate([
{ $group: { _id: "$dept", avg: { $avg: "$marks.maths" } } },
{ $merge: { into: "dept_summary", on: "_id",
whenMatched: "replace", whenNotMatched: "insert" } }
])
// $out replaces the whole target collection; $merge updates it incrementally.
// Both must be the LAST stage.
In Python, through mongomock, 15_aggregation.py:
"""Experiment 15 — $match, $group, $project, $sort.
Every figure printed here is computed by mongomock and asserted against the
hand-worked arithmetic in unit-4.md, so the notes and the pipeline check each
other. The one exception is $stdDevPop, which mongomock raises
NotImplementedError on; that is asserted as a limitation and the value is
computed in Python instead, clearly labelled.
"""
import statistics
import mongomock
from fixtures import fresh_db
def by(rows, key="_id"):
"""A pipeline result keyed for assertion. Aggregation order is not a promise."""
return {r[key]: r for r in rows}
def where_and_having_are_the_same_stage():
"""The whole pipeline, and the sentence that earns the marks."""
db = fresh_db()
rows = list(db.students.aggregate([
{"$match": {"active": True}},
{"$group": {"_id": "$dept", "avg": {"$avg": "$marks.maths"},
"n": {"$sum": 1}}},
{"$match": {"n": {"$gt": 1}}},
{"$sort": {"avg": -1}},
# $round would go here on a real server; mongomock does not have it,
# and unsupported_operators() below asserts that. Both averages are
# exact anyway, so nothing is hidden by leaving it out.
{"$project": {"_id": 0, "dept": "$_id", "avg": 1, "n": 1}},
]))
# active: true drops Kiran (DS, 71), so DS is Asha and Ravi only.
assert rows == [{"dept": "DS", "avg": 76.5, "n": 2},
{"dept": "Stats", "avg": 73.0, "n": 2}], rows
assert (88 + 65) / 2 == 76.5
assert (94 + 52) / 2 == 73.0
# $sort was honoured: descending by avg.
assert [r["avg"] for r in rows] == sorted((r["avg"] for r in rows), reverse=True)
print(" WHERE active=true, GROUP BY dept, HAVING n>1, ORDER BY avg DESC")
print(" DS (88+65)/2 = 76.5 n=2")
print(" Stats (94+52)/2 = 73.0 n=2")
print(" Kiran is DS with 71 but active:false, so the $match BEFORE")
print(" $group excludes him and the DS average RISES from 74.67")
def the_filter_changes_the_answer():
"""Same grouping, no $match: unit-4.md Problem 1(a)'s figures."""
db = fresh_db()
rows = by(db.students.aggregate([
{"$group": {"_id": "$dept", "avg": {"$avg": "$marks.maths"},
"n": {"$sum": 1}}}]))
assert rows["DS"]["n"] == 3 and rows["Stats"]["n"] == 2
assert round(rows["DS"]["avg"], 3) == 74.667, rows["DS"]["avg"]
assert (88 + 65 + 71) / 3 == rows["DS"]["avg"]
assert rows["Stats"]["avg"] == 73.0
print(" without the $match: DS (88+65+71)/3 = 74.667 over 3, Stats 73 over 2")
print(" these are unit-4.md Problem 1(a)'s numbers, and the pair of")
print(" results above is why 'which $match, and where' is the question")
def group_id_is_the_key_and_null_is_the_grand_total():
db = fresh_db()
total = list(db.students.aggregate([
{"$group": {"_id": None, "avg": {"$avg": "$marks.maths"},
"n": {"$sum": 1}}}]))
assert total == [{"_id": None, "avg": 74.0, "n": 5}], total
assert (88 + 65 + 94 + 71 + 52) / 5 == 74.0
# _id is not optional -- omitting it is an error, not a grand total.
# (A real server says "a group specification must include an _id";
# mongomock reaches for the key and raises KeyError. Either way: rejected.)
try:
list(db.students.aggregate([{"$group": {"n": {"$sum": 1}}}]))
raise SystemExit("$group without _id should not be accepted")
except KeyError as exc:
assert str(exc) == "'_id'", exc
print(" _id: null -> one bucket: avg 74.0 over 5 (the grand total)")
print(" _id omitted -> ERROR. It is the grouping KEY, not an option")
def the_accumulators():
db = fresh_db()
rows = by(db.students.aggregate([{"$group": {
"_id": "$dept",
"n": {"$sum": 1},
"totMaths": {"$sum": "$marks.maths"},
"avgMaths": {"$avg": "$marks.maths"},
"best": {"$max": "$marks.maths"},
"worst": {"$min": "$marks.maths"},
"everyone": {"$push": "$name"},
"distinct": {"$addToSet": "$dept"},
}}]))
ds = rows["DS"]
assert ds["n"] == 3
assert ds["totMaths"] == 224 == 88 + 65 + 71
assert round(ds["avgMaths"], 3) == 74.667
assert (ds["best"], ds["worst"]) == (88, 65)
assert sorted(ds["everyone"]) == ["Asha", "Kiran", "Ravi"]
assert ds["distinct"] == ["DS"], "addToSet de-duplicates; push does not"
st = rows["Stats"]
assert (st["n"], st["totMaths"], st["avgMaths"]) == (2, 146, 73.0)
assert (st["best"], st["worst"]) == (94, 52)
print(" dept n sum avg max min $push")
for d in ("DS", "Stats"):
r = rows[d]
print(f" {d:6s} {r['n']} {r['totMaths']:3d} {r['avgMaths']:7.3f} "
f"{r['best']:3d} {r['worst']:3d} {r['everyone']}")
print(" $push and $addToSet have NO SQL equivalent -- they are why")
print(" MongoDB never needed GROUP_CONCAT")
def unsupported_operators():
"""Two operators this pipeline would use on a real server, and cannot here.
Asserted rather than commented, so the day mongomock gains them this file
fails and gets rewritten to test them instead of describing them.
"""
db = fresh_db()
try:
list(db.students.aggregate([
{"$group": {"_id": "$dept", "sd": {"$stdDevPop": "$marks.maths"}}}]))
raise SystemExit("mongomock now implements $stdDevPop -- assert it here")
except NotImplementedError as exc:
assert "$stdDevPop" in str(exc), exc
try:
list(db.students.aggregate([
{"$project": {"m": {"$round": ["$marks.maths", 1]}}}]))
raise SystemExit("mongomock now implements $round -- use it above")
except Exception as exc:
assert "$round" in str(exc), exc
print(" not available in mongomock (both work on a real server):")
print(" $round OperationFailure: Unrecognized expression '$round'")
print(" $stdDevPop NotImplementedError")
print(" so $stdDevPop is computed in Python here instead:")
for dept in ("DS", "Stats"):
vals = [d["marks"]["maths"] for d in db.students.find({"dept": dept})]
pop = statistics.pstdev(vals)
samp = statistics.stdev(vals)
print(f" {dept:6s} {vals} $stdDevPop {pop:7.4f} $stdDevSamp {samp:7.4f}")
print(" Course 4's distinction, unchanged: $stdDevPop divides by n,")
print(" $stdDevSamp by n-1. MongoDB makes you choose, as R does")
def project_computes_and_renames():
db = fresh_db()
rows = by(db.students.aggregate([{"$project": {
"_id": 0, "name": 1,
"total": {"$add": ["$marks.maths", "$marks.stats"]},
"band": {"$switch": {"branches": [
{"case": {"$gte": ["$marks.maths", 75]}, "then": "Distinction"},
{"case": {"$gte": ["$marks.maths", 60]}, "then": "First"},
{"case": {"$gte": ["$marks.maths", 40]}, "then": "Pass"}],
"default": "Fail"}},
}}], ), key="name")
assert rows["Asha"] == {"name": "Asha", "total": 179, "band": "Distinction"}
assert rows["Meena"]["total"] == 183 and rows["Meena"]["band"] == "Distinction"
assert rows["Ravi"]["band"] == "First" and rows["Kiran"]["band"] == "First"
assert rows["Bhanu"]["band"] == "Pass"
assert all("dept" not in r for r in rows.values()), \
"$project is EXCLUSIVE: naming any field drops the rest"
print(" name total band")
for n in ("Meena", "Asha", "Kiran", "Ravi", "Bhanu"):
print(f" {n:6s} {rows[n]['total']:5d} {rows[n]['band']}")
print(" dept is GONE -- $project keeps only what you name. Use")
print(" $addFields when you meant 'everything, plus this'")
def addfields_keeps_everything():
db = fresh_db()
rows = list(db.students.aggregate([
{"$addFields": {"total": {"$add": ["$marks.maths", "$marks.stats"]}}},
{"$sort": {"total": -1}},
{"$limit": 3},
]))
assert [(r["name"], r["total"]) for r in rows] == \
[("Meena", 183), ("Asha", 179), ("Kiran", 137)], rows
assert all("dept" in r and "subjects" in r for r in rows), \
"$addFields adds; it never removes"
print(" top 3 by total, via $addFields: Meena 183, Asha 179, Kiran 137")
print(" dept and subjects survived -- that is the difference")
def match_first_or_pay_for_it():
"""The optimiser helps, but only when it can. State the rule, not the hope."""
db = fresh_db()
late = list(db.students.aggregate([
{"$group": {"_id": "$dept", "n": {"$sum": 1}}},
{"$match": {"_id": "DS"}}]))
early = list(db.students.aggregate([
{"$match": {"dept": "DS"}},
{"$group": {"_id": "$dept", "n": {"$sum": 1}}}]))
assert late == early == [{"_id": "DS", "n": 3}], (late, early)
print(" both orders return [{_id: 'DS', n: 3}] -- SAME ANSWER, different cost")
print(" $match first can use an index on dept and groups 3 documents;")
print(" $match last groups all 5 and then filters GROUPS, which no")
print(" index describes, because they are new documents")
print(" On 5 documents this is invisible. On 5,000,000 it is the")
print(" whole query -- and that is what explain() shows you")
def main():
print("Experiment 15 -- $match, $group, $project, $sort")
# Step 1: $match then $group: WHERE and HAVING
where_and_having_are_the_same_stage()
# Step 2: See the filter change the answer
the_filter_changes_the_answer()
# Step 3: Group by a key, or by null for a total
group_id_is_the_key_and_null_is_the_grand_total()
# Step 4: Use the accumulators
the_accumulators()
# Step 5: Note the operators mongomock lacks
unsupported_operators()
# Step 6: $project: compute and rename
project_computes_and_renames()
# Step 7: $addFields: keep everything
addfields_keeps_everything()
# Step 8: $match first
match_first_or_pay_for_it()
if __name__ == "__main__":
main()
In mongosh, 15_aggregation.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.aggregate([
... { $match: { active: true } },
... { $group: { _id: "$dept", avg: { $avg: "$marks.maths" },
... n: { $sum: 1 } } },
... { $match: { n: { $gt: 1 } } },
... { $sort: { avg: -1 } },
... { $project: { _id: 0, dept: "$_id", avg: { $round: ["$avg", 2] }, n: 1 } }
... ])
[ { n: 2, dept: 'DS', avg: 76.5 }, { n: 2, dept: 'Stats', avg: 73 } ]
collegeDB> db.students.aggregate([
... { $group: { _id: "$dept", avgMaths: { $avg: "$marks.maths" },
... n: { $sum: 1 } } },
... { $sort: { _id: 1 } }
... ])
[
{ _id: 'DS', avgMaths: 74.66666666666667, n: 3 },
{ _id: 'Stats', avgMaths: 73, n: 2 }
]
collegeDB> db.students.aggregate([ { $group: { _id: null, avg: { $avg: "$marks.maths" },
... n: { $sum: 1 } } } ])
[ { _id: null, avg: 74, n: 5 } ]
collegeDB> db.students.aggregate([
... { $group: {
... _id: "$dept",
... n: { $sum: 1 }, // COUNT(*)
... totMaths: { $sum: "$marks.maths" }, // SUM
... avgMaths: { $avg: "$marks.maths" }, // AVG
... best: { $max: "$marks.maths" }, // MAX
... worst: { $min: "$marks.maths" }, // MIN
... sd: { $stdDevPop: "$marks.maths" }, // Course 4's population sd
... everyone: { $push: "$name" }, // ALL values, as an array
... distinct: { $addToSet: "$name" }, // DISTINCT values
... anyone: { $first: "$name" } // needs a $sort to be meaningful
... } },
... { $set: { distinct: { $sortArray: { input: "$distinct", sortBy: 1 } } } },
... { $sort: { _id: 1 } }
... ])
[
{
_id: 'DS',
n: 3,
totMaths: 224,
avgMaths: 74.66666666666667,
best: 88,
worst: 65,
sd: 9.741092797468305,
everyone: [ 'Asha', 'Ravi', 'Kiran' ],
distinct: [ 'Asha', 'Kiran', 'Ravi' ],
anyone: 'Asha'
},
{
_id: 'Stats',
n: 2,
totMaths: 146,
avgMaths: 73,
best: 94,
worst: 52,
sd: 21,
everyone: [ 'Meena', 'Bhanu' ],
distinct: [ 'Bhanu', 'Meena' ],
anyone: 'Meena'
}
]
collegeDB> db.students.aggregate([
... { $project: {
... _id: 0,
... name: 1, // include
... total: { $add: ["$marks.maths", "$marks.stats"] },
... pct: { $round: [ { $divide: [ { $add: ["$marks.maths", "$marks.stats"] },
... 2 ] }, 1 ] },
... dept: "$dept", // rename by re-assigning
... band: { $switch: { branches: [
... { case: { $gte: ["$marks.maths", 75] }, then: "Distinction" },
... { case: { $gte: ["$marks.maths", 60] }, then: "First" },
... { case: { $gte: ["$marks.maths", 40] }, then: "Pass" } ],
... default: "Fail" } }
... } }
... ])
[
{
name: 'Asha',
total: 179,
pct: 89.5,
dept: 'DS',
band: 'Distinction'
},
{ name: 'Ravi', total: 123, pct: 61.5, dept: 'DS', band: 'First' },
{
name: 'Meena',
total: 183,
pct: 91.5,
dept: 'Stats',
band: 'Distinction'
},
{ name: 'Kiran', total: 137, pct: 68.5, dept: 'DS', band: 'First' },
{ name: 'Bhanu', total: 99, pct: 49.5, dept: 'Stats', band: 'Pass' }
]
collegeDB> db.students.aggregate([
... { $addFields: { total: { $add: ["$marks.maths", "$marks.stats"] } } },
... { $sort: { total: -1 } },
... { $limit: 3 }
... ])
[
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: [ 'Stats', 'R' ],
age: 20,
active: true,
total: 183
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: [ 'DS', 'Stats', 'Python' ],
age: 20,
active: true,
total: 179
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: [ 'DS' ],
age: 22,
active: false,
total: 137
}
]
collegeDB> db.students.aggregate([ { $group: { _id: "$dept", n: { $sum: 1 } } },
... { $match: { _id: "DS" } } ]) // SLOW: groups all
[ { _id: 'DS', n: 3 } ]
collegeDB> db.students.aggregate([ { $match: { dept: "DS" } },
... { $group: { _id: "$dept", n: { $sum: 1 } } } ]) // fast
[ { _id: 'DS', n: 3 } ]
collegeDB> db.students.aggregate([ { $match: { active: true } },
... { $group: { _id: "$dept", n: { $sum: 1 } } } ],
... { explain: true })
{
explainVersion: '2',
queryPlanner: {
namespace: 'collegeDB.students',
parsedQuery: { active: { '$eq': true } },
indexFilterSet: false,
queryHash: '38BCAAF9',
planCacheShapeHash: '38BCAAF9',
planCacheKey: 'FE63AA7F',
optimizationTimeMillis: 0,
optimizedPipeline: true,
maxIndexedOrSolutionsReached: false,
maxIndexedAndSolutionsReached: false,
maxScansToExplodeReached: false,
prunedSimilarIndexes: false,
winningPlan: {
isCached: false,
queryPlan: {
stage: 'GROUP',
planNodeId: 3,
inputStage: {
stage: 'COLLSCAN',
planNodeId: 1,
filter: { active: { '$eq': true } },
nss: 'collegeDB.students',
direction: 'forward'
}
},
slotBasedPlan: {
slots: '$$RESULT=s9 env: { }',
stages: '[3] project [s9 = newBsonObj("_id", s7, "n", s8)] \n' +
'[3] project [s8 = (convert ( s6, int32) ?: s6)] \n' +
'[3] group [s7] [s6 = count()] spillSlots[s5] mergingExprs[sum(s5)] \n' +
'[3] project [s7 = (s3 ?: null)] \n' +
'[1] filter {traverseF(s4, lambda(l2.0) { ((move(l2.0) == true) ?: false) }, false)} \n' +
'[1] scan generic [s1 = record, s2 = recordId] [s3 = dept, s4 = active] @"ac85a8ec-c420-4d7a-9134-2c217094af59" '
}
},
rejectedPlans: []
},
executionStats: {
executionSuccess: true,
nReturned: 2,
executionTimeMillis: 8,
totalKeysExamined: 0,
totalDocsExamined: 5,
executionStages: {
stage: 'project',
planNodeId: 3,
nReturned: 2,
executionTimeMillisEstimate: 6,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
projections: { '9': 'newBsonObj("_id", s7, "n", s8) ' },
inputStage: {
stage: 'project',
planNodeId: 3,
nReturned: 2,
executionTimeMillisEstimate: 6,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
projections: { '8': '(convert ( s6, int32) ?: s6) ' },
inputStage: {
stage: 'group',
planNodeId: 3,
nReturned: 2,
executionTimeMillisEstimate: 6,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
groupBySlots: [ Long('7') ],
expressions: {
'6': 'count() ',
initExprs: { '6': null, mergingExprs: { '5': 'sum(s5) ' } }
},
usedDisk: true,
spills: 2,
spilledBytes: 36,
spilledRecords: 2,
spilledDataStorageSize: 4096,
peakTrackedMemBytes: 50,
inputStage: {
stage: 'project',
planNodeId: 3,
nReturned: 4,
executionTimeMillisEstimate: 0,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
projections: { '7': '(s3 ?: null) ' },
inputStage: {
stage: 'filter',
planNodeId: 1,
nReturned: 4,
executionTimeMillisEstimate: 0,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
numTested: 5,
filter: 'traverseF(s4, lambda(l2.0) { ((move(l2.0) == true) ?: false) }, false) ',
inputStage: {
stage: 'scan',
planNodeId: 1,
nReturned: 5,
executionTimeMillisEstimate: 0,
opens: 1,
closes: 1,
saveState: 0,
restoreState: 0,
isEOF: 1,
numReads: 5,
recordSlot: 1,
recordIdSlot: 2,
scanFieldNames: [ 'dept', 'active' ],
scanFieldSlots: [ Long('3'), Long('4') ]
}
}
}
}
}
},
allPlansExecution: []
},
queryShapeHash: 'A17918C996816EA2A7275E096C6198D8C70A925310205C2EF51E25DFDC7B2BD2',
peakTrackedMemBytes: Long('50'),
command: {
aggregate: 'students',
pipeline: [
{ '$match': { active: true } },
{ '$group': { _id: '$dept', n: { '$sum': 1 } } }
],
cursor: {},
'$db': 'collegeDB'
},
serverInfo: {
host: 'vm',
port: 27017,
version: '8.3.7',
gitVersion: 'nogitversion'
},
serverParameters: {
internalQueryFacetBufferSizeBytes: 104857600,
internalDocumentSourceGroupMaxMemoryBytes: 104857600,
internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
internalQueryFacetMaxOutputDocSizeBytes: 104857600,
internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
internalQueryProhibitBlockingMergeOnMongoS: 0,
internalQueryMaxAddToSetBytes: 104857600,
internalQueryFrameworkControl: 'trySbeRestricted',
internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
},
ok: 1
}
collegeDB> db.students.aggregate([
... { $group: { _id: "$dept", avg: { $avg: "$marks.maths" } } },
... { $merge: { into: "dept_summary", on: "_id",
... whenMatched: "replace", whenNotMatched: "insert" } }
... ])
In Python, through mongomock, 15_aggregation.py:
OUTPUT
Experiment 15 -- $match, $group, $project, $sort
WHERE active=true, GROUP BY dept, HAVING n>1, ORDER BY avg DESC
DS (88+65)/2 = 76.5 n=2
Stats (94+52)/2 = 73.0 n=2
Kiran is DS with 71 but active:false, so the $match BEFORE
$group excludes him and the DS average RISES from 74.67
without the $match: DS (88+65+71)/3 = 74.667 over 3, Stats 73 over 2
these are unit-4.md Problem 1(a)'s numbers, and the pair of
results above is why 'which $match, and where' is the question
_id: null -> one bucket: avg 74.0 over 5 (the grand total)
_id omitted -> ERROR. It is the grouping KEY, not an option
dept n sum avg max min $push
DS 3 224 74.667 88 65 ['Asha', 'Ravi', 'Kiran']
Stats 2 146 73.000 94 52 ['Meena', 'Bhanu']
$push and $addToSet have NO SQL equivalent -- they are why
MongoDB never needed GROUP_CONCAT
not available in mongomock (both work on a real server):
$round OperationFailure: Unrecognized expression '$round'
$stdDevPop NotImplementedError
so $stdDevPop is computed in Python here instead:
DS [88, 65, 71] $stdDevPop 9.7411 $stdDevSamp 11.9304
Stats [94, 52] $stdDevPop 21.0000 $stdDevSamp 29.6985
Course 4's distinction, unchanged: $stdDevPop divides by n,
$stdDevSamp by n-1. MongoDB makes you choose, as R does
name total band
Meena 183 Distinction
Asha 179 Distinction
Kiran 137 First
Ravi 123 First
Bhanu 99 Pass
dept is GONE -- $project keeps only what you name. Use
$addFields when you meant 'everything, plus this'
top 3 by total, via $addFields: Meena 183, Asha 179, Kiran 137
dept and subjects survived -- that is the difference
both orders return [{_id: 'DS', n: 3}] -- SAME ANSWER, different cost
$match first can use an index on dept and groups 3 documents;
$match last groups all 5 and then filters GROUPS, which no
index describes, because they are new documents
On 5 documents this is invisible. On 5,000,000 it is the
whole query -- and that is what explain() shows you
Changed: two $groups now end with a $sort, and the $addToSet list of names is
sorted with $sortArray. $group and $addToSet keep no order, and the names came back in a
different order from one run to the next.
RESULT
With the active filter DS averages 76.5 over 2 and Stats 73 over 2; without it, DS averages 74.67 over 3.
Aggregate across collections and arrays with $lookup, $unwind and $bucket.
Unwind arrays, join collections, and bucket values into ranges.
In mongosh, 16_advanced_agg.js:
In Python, through mongomock, 16_advanced_agg.py:
THE POINT
Subject counts DS 3, Stats 3, Python 2, R 1; and the bucket
distribution — with the top boundary at 101, not 100, because buckets are
[lower, upper) and 100 as the boundary would lose a perfect scorer to
default.
$unwind silently drops empty and missing arrays, and
preserveNullAndEmptyArrays: true keeps them. That one is worth seeing fail.
In mongosh, 16_advanced_agg.js:
// Experiment 16 -- $lookup, $unwind and $bucket.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 16_advanced_agg.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.
// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
// =============================================================================
// $unwind -- one output document per array element
// =============================================================================
// Step 2: $unwind, and count the subjects
db.students.aggregate([ { $unwind: "$subjects" } ])
// 5 students with 3+2+2+1+1 subjects -> 9 documents out.
// The only way to count array CONTENTS:
db.students.aggregate([
{ $unwind: "$subjects" },
{ $group: { _id: "$subjects", n: { $sum: 1 },
who: { $push: "$name" } } },
{ $sort: { n: -1, _id: 1 } }
])
// DS 3, Stats 3, Python 2, R 1
// Step 3: See $unwind drop empty and missing arrays
db.students.insertOne({ _id: 26, name: "Latha", dept: "DS", subjects: [] })
db.students.insertOne({ _id: 27, name: "Mohan", dept: "DS" }) // no field
db.students.aggregate([ { $unwind: "$subjects" },
{ $count: "rows" } ]) // Latha and Mohan are GONE
db.students.aggregate([
{ $unwind: { path: "$subjects", preserveNullAndEmptyArrays: true } },
{ $count: "rows" }
]) // both come back, subjects unset
// includeArrayIndex gives you the position, which $unwind otherwise loses:
db.students.aggregate([
{ $unwind: { path: "$subjects", includeArrayIndex: "pos" } },
{ $match: { pos: 0 } } // each student's FIRST subject
])
// $unwind on a NON-array behaves as if it were a one-element array -- it does
// NOT error. That is why a typo'd path silently returns nothing instead.
// =============================================================================
// $lookup -- the left outer join
// =============================================================================
db.enrollments.aggregate([
{ $lookup: { from: "students", localField: "student_id",
foreignField: "_id", as: "student" } },
{ $lookup: { from: "courses", localField: "course_id",
foreignField: "_id", as: "course" } },
{ $unwind: "$student" },
{ $unwind: "$course" },
{ $project: { _id: 0, name: "$student.name",
title: "$course.title", grade: 1 } }
])
// as: ALWAYS an array, even for a one-to-one match -- hence the $unwind.
// LEFT OUTER: an unmatched document keeps its row with an EMPTY array, which
// is exactly why the $unwind after it silently deletes the unmatched rows.
// If you want them, preserveNullAndEmptyArrays: true.
// Step 4: $lookup, and count enrolments per course
db.courses.aggregate([
{ $lookup: { from: "enrollments", localField: "_id",
foreignField: "course_id", as: "enrolled" } },
{ $project: { _id: 0, title: 1,
n: { $size: "$enrolled" } } }, // $size, no $unwind needed
{ $sort: { n: -1 } }
])
// $size on the joined array beats $unwind + $group when you only want a count.
// Step 5: Filter the joined side first, with $lookup's pipeline
db.courses.aggregate([
{ $lookup: {
from: "enrollments",
let: { cid: "$_id" },
pipeline: [
{ $match: { $expr: { $and: [ { $eq: ["$course_id", "$$cid"] },
{ $eq: ["$grade", "A"] } ] } } },
{ $project: { _id: 0, student_id: 1 } }
],
as: "aGrades" } }
])
// $$cid is the OUTER variable; $course_id the inner field. Two dollars means
// "from let". This form is how you avoid dragging 10,000 rows in to keep 3.
// =============================================================================
// $bucket and $bucketAuto -- histograms
// =============================================================================
// Step 6: $bucket and $facet
db.students.aggregate([
{ $bucket: {
groupBy: "$marks.maths",
boundaries: [0, 40, 60, 75, 101],
default: "Other",
output: { count: { $sum: 1 }, names: { $push: "$name" } } } }
])
// Boundaries are [lower, upper) -- CLOSED below, OPEN above.
// The last one is 101, NOT 100: with 100 as the top boundary a student who
// scored exactly 100 falls outside every bucket and lands in "Other".
// Without a `default`, an out-of-range value is an ERROR, not a silent drop.
db.students.aggregate([
{ $bucketAuto: { groupBy: "$marks.maths", buckets: 3 } }
])
// $bucketAuto picks the boundaries to even out the COUNTS. Good for
// exploration; useless for a report, because the boundaries move when the
// data does and yesterday's chart is not comparable with today's.
// $facet runs several pipelines over the SAME input, in one pass:
db.students.aggregate([
{ $facet: {
byDept: [ { $group: { _id: "$dept", n: { $sum: 1 } } }, { $sort: { _id: 1 } } ],
byBand: [ { $bucket: { groupBy: "$marks.maths",
boundaries: [0, 40, 60, 75, 101],
default: "Other",
output: { n: { $sum: 1 } } } } ],
topThree: [ { $sort: { "marks.maths": -1 } }, { $limit: 3 },
{ $project: { _id: 0, name: 1 } } ] } }
])
// [Changed: the $sort in byDept was added. $group returns its groups in no fixed
// order, and they came back in a different order from one run to the next.]
In Python, through mongomock, 16_advanced_agg.py:
"""Experiment 16 — $lookup, $unwind and $bucket.
Executed and asserted: $unwind (including the empty-array trap), $lookup in
both directions, $bucket, and $facet.
Not available in mongomock, and asserted as unavailable rather than described
as if tested: $bucketAuto, and $lookup's let/pipeline form.
"""
from pymongo.errors import OperationFailure
from fixtures import fresh_db
def by(rows, key="_id"):
return {r[key]: r for r in rows}
# =============================================================================
# $unwind
# =============================================================================
def one_document_per_element():
db = fresh_db()
before = db.students.count_documents({})
after = len(list(db.students.aggregate([{"$unwind": "$subjects"}])))
assert before == 5
assert after == 9 == 3 + 2 + 2 + 1 + 1, after
print(f" {before} students -> {after} documents (3+2+2+1+1 subjects)")
def counting_array_contents():
"""The whole reason $unwind exists. $group alone cannot do this."""
db = fresh_db()
rows = by(db.students.aggregate([
{"$unwind": "$subjects"},
{"$group": {"_id": "$subjects", "n": {"$sum": 1},
"who": {"$push": "$name"}}},
{"$sort": {"n": -1, "_id": 1}}]))
assert {k: v["n"] for k, v in rows.items()} == \
{"DS": 3, "Stats": 3, "Python": 2, "R": 1}
assert sorted(rows["DS"]["who"]) == ["Asha", "Kiran", "Ravi"]
assert sorted(rows["Stats"]["who"]) == ["Asha", "Bhanu", "Meena"]
assert rows["R"]["who"] == ["Meena"]
assert sum(v["n"] for v in rows.values()) == 9, "every element counted once"
print(" subject counts:")
for s in ("DS", "Stats", "Python", "R"):
print(f" {s:7s} {rows[s]['n']} {sorted(rows[s]['who'])}")
def unwind_silently_drops_empty_and_missing():
"""The one worth seeing fail. Two students vanish from a count."""
db = fresh_db()
db.students.insert_one({"_id": 26, "name": "Latha", "dept": "DS",
"subjects": []})
db.students.insert_one({"_id": 27, "name": "Mohan", "dept": "DS"})
assert db.students.count_documents({}) == 7
dropped = list(db.students.aggregate([{"$unwind": "$subjects"}]))
assert len(dropped) == 9, len(dropped)
assert "Latha" not in {d["name"] for d in dropped}
assert "Mohan" not in {d["name"] for d in dropped}
kept = list(db.students.aggregate([
{"$unwind": {"path": "$subjects",
"preserveNullAndEmptyArrays": True}}]))
assert len(kept) == 11, len(kept)
survivors = {d["name"] for d in kept}
assert "Latha" in survivors and "Mohan" in survivors
assert all("subjects" not in d for d in kept
if d["name"] in ("Latha", "Mohan")), \
"they come back with the field UNSET, not with an empty array"
print(" 7 students, two with no subjects (empty array / missing field):")
print(f" plain $unwind -> {len(dropped)} docs, both LOST")
print(f" preserveNullAndEmptyArrays: true -> {len(kept)} docs, both kept")
print(" 'my count is short and I cannot see why' is nearly always this")
def includearrayindex_and_non_arrays():
db = fresh_db()
firsts = list(db.students.aggregate([
{"$unwind": {"path": "$subjects", "includeArrayIndex": "pos"}},
{"$match": {"pos": 0}},
{"$project": {"_id": 0, "name": 1, "subjects": 1}}]))
assert len(firsts) == 5, firsts
assert {d["name"]: d["subjects"] for d in firsts} == \
{"Asha": "DS", "Ravi": "DS", "Meena": "Stats",
"Kiran": "DS", "Bhanu": "Stats"}
# $unwind on a NON-array is not an error: it acts as a one-element array.
scalar = list(db.students.aggregate([{"$unwind": "$name"}]))
assert len(scalar) == 5, "unwinding a string yields the same 5 documents"
# A typo'd path is therefore SILENT -- it just returns nothing.
typo = list(db.students.aggregate([{"$unwind": "$subject"}]))
assert typo == [], "no error, no rows -- the commonest silent failure"
print(" includeArrayIndex 'pos', $match pos:0 -> each student's FIRST subject")
print(" $unwind on a string -> 5 documents (treated as one element)")
print(" $unwind on '$subject' -> 0 documents (a TYPO, and it does not error)")
# =============================================================================
# $lookup
# =============================================================================
def lookup_returns_an_array_and_is_a_left_outer_join():
db = fresh_db()
db.enrollments.insert_one({"student_id": 21, "course_id": "GONE404",
"grade": "F"})
joined = list(db.enrollments.aggregate([
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}}]))
assert len(joined) == 6, "every enrolment survives -- LEFT outer"
sizes = {d["course_id"]: len(d["course"]) for d in joined}
assert sizes["GONE404"] == 0, "no match -> EMPTY ARRAY, not a missing row"
assert all(v == 1 for k, v in sizes.items() if k != "GONE404"), sizes
assert all(isinstance(d["course"], list) for d in joined), \
"as: is ALWAYS an array, even one-to-one -- that is why $unwind follows"
# And now the consequence: the $unwind after it deletes the orphan.
unwound = list(db.enrollments.aggregate([
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": "$course"}]))
assert len(unwound) == 5, "the orphan's empty array was dropped by $unwind"
print(" 6 enrolments, one pointing at a course that does not exist:")
print(" after $lookup -> 6 rows, the orphan has course: []")
print(" after $lookup + $unwind -> 5 rows, the orphan is GONE")
print(" $lookup is a LEFT outer join; the $unwind after it turns it")
print(" into an inner one. Nothing warns you")
def joining_both_directions():
db = fresh_db()
# enrolments -> the student and the course behind each one
rows = list(db.enrollments.aggregate([
{"$lookup": {"from": "students", "localField": "student_id",
"foreignField": "_id", "as": "student"}},
{"$lookup": {"from": "courses", "localField": "course_id",
"foreignField": "_id", "as": "course"}},
{"$unwind": "$student"},
{"$unwind": "$course"},
{"$project": {"_id": 0, "name": "$student.name",
"title": "$course.title", "grade": 1}},
{"$sort": {"name": 1, "title": 1}}]))
assert len(rows) == 5
assert rows[0] == {"name": "Asha", "title": "Data Science with R",
"grade": "A"}, rows[0]
assert {r["name"] for r in rows} == {"Asha", "Ravi", "Meena", "Kiran"}
# courses -> how many enrolled. $size beats $unwind when you want a count.
counts = list(db.courses.aggregate([
{"$lookup": {"from": "enrollments", "localField": "_id",
"foreignField": "course_id", "as": "enrolled"}},
{"$project": {"_id": 0, "title": 1, "n": {"$size": "$enrolled"}}},
{"$sort": {"n": -1, "title": 1}}]))
assert counts == [{"title": "Data Science with R", "n": 2},
{"title": "Statistical Foundations", "n": 2},
{"title": "Web Technologies", "n": 1}], counts
print(" enrolments -> student + course:")
for r in rows:
print(f" {r['name']:6s} {r['title']:24s} {r['grade']}")
print(" courses -> enrolment counts, via $size on the joined array:")
for c in counts:
print(f" {c['title']:24s} {c['n']}")
print(" $size needs no $unwind and no $group -- one stage, one pass")
def the_let_pipeline_form_is_not_implemented_here():
"""Asserted, so this file cannot start claiming to test what it describes."""
db = fresh_db()
try:
list(db.courses.aggregate([{"$lookup": {
"from": "enrollments",
"let": {"cid": "$_id"},
"pipeline": [{"$match": {"$expr": {"$and": [
{"$eq": ["$course_id", "$$cid"]},
{"$eq": ["$grade", "A"]}]}}}],
"as": "aGrades"}}]))
raise SystemExit("mongomock now implements let/pipeline -- assert it")
except NotImplementedError as exc:
assert "let" in str(exc), exc
# The runnable equivalent: join everything, then filter. Same answer,
# more work -- which is exactly the point the let form makes.
rows = list(db.courses.aggregate([
{"$lookup": {"from": "enrollments", "localField": "_id",
"foreignField": "course_id", "as": "e"}},
{"$unwind": "$e"},
{"$match": {"e.grade": "A"}},
{"$group": {"_id": "$title", "n": {"$sum": 1}}},
{"$sort": {"_id": 1}}]))
assert rows == [{"_id": "Data Science with R", "n": 1},
{"_id": "Statistical Foundations", "n": 1}], rows
print(" $lookup with let/pipeline: NotImplementedError in mongomock")
print(" the join-then-filter equivalent DOES run, and gives:")
for r in rows:
print(f" {r['_id']:24s} {r['n']} grade-A enrolment(s)")
print(" same answer, and on real data far more expensive: it drags")
print(" every enrolment in and then throws most of them away.")
print(" $$cid is the OUTER let variable, $course_id the inner field")
# =============================================================================
# $bucket
# =============================================================================
def bucket_boundaries_are_closed_below_and_open_above():
db = fresh_db()
rows = by(db.students.aggregate([{"$bucket": {
"groupBy": "$marks.maths",
"boundaries": [0, 40, 60, 75, 101],
"default": "Other",
"output": {"count": {"$sum": 1}, "names": {"$push": "$name"}}}}]))
assert {k: v["count"] for k, v in rows.items()} == {40: 1, 60: 2, 75: 2}
assert rows[40]["names"] == ["Bhanu"] # 52
assert sorted(rows[60]["names"]) == ["Kiran", "Ravi"] # 71, 65
assert sorted(rows[75]["names"]) == ["Asha", "Meena"] # 88, 94
assert 0 not in rows, "the 0-40 bucket is EMPTY and is simply not emitted"
print(" bucket count names")
for lo, hi in ((0, 40), (40, 60), (60, 75), (75, 101)):
r = rows.get(lo)
label = f"[{lo:3d},{hi:4d})"
if r:
print(f" {label} {r['count']:5d} {sorted(r['names'])}")
else:
print(f" {label} {'--':>5s} (empty buckets are NOT emitted)")
def why_the_top_boundary_is_101():
"""A perfect score is the test case, and 100 as the boundary loses it."""
db = fresh_db()
db.students.insert_one({"_id": 28, "name": "Perfect", "dept": "DS",
"marks": {"maths": 100, "stats": 100},
"subjects": ["DS"], "age": 20, "active": True})
with_101 = by(db.students.aggregate([{"$bucket": {
"groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 101],
"default": "Other", "output": {"names": {"$push": "$name"}}}}]))
assert "Perfect" in with_101[75]["names"], with_101
with_100 = by(db.students.aggregate([{"$bucket": {
"groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 100],
"default": "Other", "output": {"names": {"$push": "$name"}}}}]))
assert with_100["Other"]["names"] == ["Perfect"], with_100
assert "Perfect" not in with_100[75]["names"]
# And without a default, an out-of-range value is an ERROR.
try:
list(db.students.aggregate([{"$bucket": {
"groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 100],
"output": {"n": {"$sum": 1}}}}]))
raise SystemExit("$bucket should reject an out-of-range value with no default")
except OperationFailure as exc:
assert "no default was specified" in str(exc), exc
print(" a student who scored exactly 100:")
print(" boundaries [..., 75, 101] -> lands in [75,101) CORRECT")
print(" boundaries [..., 75, 100] -> lands in 'Other' WRONG")
print(" boundaries [..., 75, 100], no default -> ERROR")
print(" buckets are [lower, upper): closed below, OPEN above. The top")
print(" boundary must exceed the maximum, so it is max+1, not max")
def bucketauto_is_not_implemented_here():
db = fresh_db()
try:
list(db.students.aggregate([
{"$bucketAuto": {"groupBy": "$marks.maths", "buckets": 3}}]))
raise SystemExit("mongomock now implements $bucketAuto -- assert it")
except NotImplementedError as exc:
assert "$bucketAuto" in str(exc), exc
print(" $bucketAuto: NotImplementedError in mongomock")
print(" it picks boundaries to even out the COUNTS, so you never")
print(" state them. Good for a first look at unfamiliar data; wrong")
print(" for a report, because the boundaries MOVE when the data does")
print(" and last month's chart is no longer comparable with this one")
def facet_runs_several_pipelines_in_one_pass():
db = fresh_db()
out = list(db.students.aggregate([{"$facet": {
"byDept": [{"$group": {"_id": "$dept", "n": {"$sum": 1}}},
{"$sort": {"_id": 1}}],
"topThree": [{"$sort": {"marks.maths": -1}}, {"$limit": 3},
{"$project": {"_id": 0, "name": 1}}],
}}]))
assert len(out) == 1, "$facet emits exactly ONE document"
result = out[0]
assert result["byDept"] == [{"_id": "DS", "n": 3},
{"_id": "Stats", "n": 2}], result["byDept"]
assert [d["name"] for d in result["topThree"]] == ["Meena", "Asha", "Kiran"]
print(" $facet -> ONE document holding both results:")
print(f" byDept {result['byDept']}")
print(f" topThree {[d['name'] for d in result['topThree']]}")
print(" one pass over the collection instead of two queries, which")
print(" is how a dashboard gets all its panels in a single round trip")
def main():
print("Experiment 16 -- $lookup, $unwind, $bucket")
# Step 1: $unwind: one document per element
one_document_per_element()
# Step 2: Count what the arrays hold
counting_array_contents()
# Step 3: See $unwind drop empty and missing arrays
unwind_silently_drops_empty_and_missing()
# Step 4: includeArrayIndex, and fields that are not arrays
includearrayindex_and_non_arrays()
# Step 5: $lookup: an array, by a left outer join
lookup_returns_an_array_and_is_a_left_outer_join()
# Step 6: Join both ways
joining_both_directions()
# Step 7: Note that mongomock lacks $lookup's pipeline form
the_let_pipeline_form_is_not_implemented_here()
# Step 8: $bucket: closed below, open above
bucket_boundaries_are_closed_below_and_open_above()
# Step 9: See why the top boundary is 101
why_the_top_boundary_is_101()
# Step 10: Note that mongomock lacks $bucketAuto
bucketauto_is_not_implemented_here()
# Step 11: $facet: several pipelines in one pass
facet_runs_several_pipelines_in_one_pass()
if __name__ == "__main__":
main()
In mongosh, 16_advanced_agg.js:
OUTPUT
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.aggregate([ { $unwind: "$subjects" } ])
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: 'DS',
age: 20,
active: true
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: 'Stats',
age: 20,
active: true
},
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: 'Python',
age: 20,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: 'DS',
age: 21,
active: true
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: 'Python',
age: 21,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: 'Stats',
age: 20,
active: true
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: 'R',
age: 20,
active: true
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: 'DS',
age: 22,
active: false
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: 'Stats',
age: 21,
active: true
}
]
collegeDB> db.students.aggregate([
... { $unwind: "$subjects" },
... { $group: { _id: "$subjects", n: { $sum: 1 },
... who: { $push: "$name" } } },
... { $sort: { n: -1, _id: 1 } }
... ])
[
{ _id: 'DS', n: 3, who: [ 'Asha', 'Ravi', 'Kiran' ] },
{ _id: 'Stats', n: 3, who: [ 'Asha', 'Meena', 'Bhanu' ] },
{ _id: 'Python', n: 2, who: [ 'Asha', 'Ravi' ] },
{ _id: 'R', n: 1, who: [ 'Meena' ] }
]
collegeDB> db.students.insertOne({ _id: 26, name: "Latha", dept: "DS", subjects: [] })
{ acknowledged: true, insertedId: 26 }
collegeDB> db.students.insertOne({ _id: 27, name: "Mohan", dept: "DS" }) // no field
{ acknowledged: true, insertedId: 27 }
collegeDB> db.students.aggregate([ { $unwind: "$subjects" },
... { $count: "rows" } ]) // Latha and Mohan are GONE
[ { rows: 9 } ]
collegeDB> db.students.aggregate([
... { $unwind: { path: "$subjects", preserveNullAndEmptyArrays: true } },
... { $count: "rows" }
... ]) // both come back, subjects unset
[ { rows: 11 } ]
collegeDB> db.students.aggregate([
... { $unwind: { path: "$subjects", includeArrayIndex: "pos" } },
... { $match: { pos: 0 } } // each student's FIRST subject
... ])
[
{
_id: 21,
name: 'Asha',
dept: 'DS',
marks: { maths: 88, stats: 91 },
subjects: 'DS',
age: 20,
active: true,
pos: Long('0')
},
{
_id: 22,
name: 'Ravi',
dept: 'DS',
marks: { maths: 65, stats: 58 },
subjects: 'DS',
age: 21,
active: true,
pos: Long('0')
},
{
_id: 23,
name: 'Meena',
dept: 'Stats',
marks: { maths: 94, stats: 89 },
subjects: 'Stats',
age: 20,
active: true,
pos: Long('0')
},
{
_id: 24,
name: 'Kiran',
dept: 'DS',
marks: { maths: 71, stats: 66 },
subjects: 'DS',
age: 22,
active: false,
pos: Long('0')
},
{
_id: 25,
name: 'Bhanu',
dept: 'Stats',
marks: { maths: 52, stats: 47 },
subjects: 'Stats',
age: 21,
active: true,
pos: Long('0')
}
]
collegeDB> db.enrollments.aggregate([
... { $lookup: { from: "students", localField: "student_id",
... foreignField: "_id", as: "student" } },
... { $lookup: { from: "courses", localField: "course_id",
... foreignField: "_id", as: "course" } },
... { $unwind: "$student" },
... { $unwind: "$course" },
... { $project: { _id: 0, name: "$student.name",
... title: "$course.title", grade: 1 } }
... ])
[
{ grade: 'A', name: 'Asha', title: 'Data Science with R' },
{ grade: 'B', name: 'Asha', title: 'Statistical Foundations' },
{ grade: 'C', name: 'Ravi', title: 'Data Science with R' },
{ grade: 'A', name: 'Meena', title: 'Statistical Foundations' },
{ grade: 'B', name: 'Kiran', title: 'Web Technologies' }
]
collegeDB> db.courses.aggregate([
... { $lookup: { from: "enrollments", localField: "_id",
... foreignField: "course_id", as: "enrolled" } },
... { $project: { _id: 0, title: 1,
... n: { $size: "$enrolled" } } }, // $size, no $unwind needed
... { $sort: { n: -1 } }
... ])
[
{ title: 'Data Science with R', n: 2 },
{ title: 'Statistical Foundations', n: 2 },
{ title: 'Web Technologies', n: 1 }
]
collegeDB> db.courses.aggregate([
... { $lookup: {
... from: "enrollments",
... let: { cid: "$_id" },
... pipeline: [
... { $match: { $expr: { $and: [ { $eq: ["$course_id", "$$cid"] },
... { $eq: ["$grade", "A"] } ] } } },
... { $project: { _id: 0, student_id: 1 } }
... ],
... as: "aGrades" } }
... ])
[
{
_id: 'DSC301',
title: 'Data Science with R',
credits: 4,
instructor: 'Dr. Rao',
aGrades: [ { student_id: 21 } ]
},
{
_id: 'STA302',
title: 'Statistical Foundations',
credits: 3,
instructor: 'Dr. Devi',
aGrades: [ { student_id: 23 } ]
},
{
_id: 'WEB303',
title: 'Web Technologies',
credits: 3,
instructor: 'Dr. Kumar',
aGrades: []
}
]
collegeDB> db.students.aggregate([
... { $bucket: {
... groupBy: "$marks.maths",
... boundaries: [0, 40, 60, 75, 101],
... default: "Other",
... output: { count: { $sum: 1 }, names: { $push: "$name" } } } }
... ])
[
{ _id: 40, count: 1, names: [ 'Bhanu' ] },
{ _id: 60, count: 2, names: [ 'Ravi', 'Kiran' ] },
{ _id: 75, count: 2, names: [ 'Asha', 'Meena' ] },
{ _id: 'Other', count: 2, names: [ 'Latha', 'Mohan' ] }
]
collegeDB> db.students.aggregate([
... { $bucketAuto: { groupBy: "$marks.maths", buckets: 3 } }
... ])
[
{ _id: { min: null, max: 52 }, count: 2 },
{ _id: { min: 52, max: 71 }, count: 2 },
{ _id: { min: 71, max: 94 }, count: 3 }
]
collegeDB> db.students.aggregate([
... { $facet: {
... byDept: [ { $group: { _id: "$dept", n: { $sum: 1 } } }, { $sort: { _id: 1 } } ],
... byBand: [ { $bucket: { groupBy: "$marks.maths",
... boundaries: [0, 40, 60, 75, 101],
... default: "Other",
... output: { n: { $sum: 1 } } } } ],
... topThree: [ { $sort: { "marks.maths": -1 } }, { $limit: 3 },
... { $project: { _id: 0, name: 1 } } ] } }
... ])
[
{
byDept: [ { _id: 'DS', n: 5 }, { _id: 'Stats', n: 2 } ],
byBand: [
{ _id: 40, n: 1 },
{ _id: 60, n: 2 },
{ _id: 75, n: 2 },
{ _id: 'Other', n: 2 }
],
topThree: [ { name: 'Meena' }, { name: 'Asha' }, { name: 'Kiran' } ]
}
]
In Python, through mongomock, 16_advanced_agg.py:
OUTPUT
Experiment 16 -- $lookup, $unwind, $bucket
5 students -> 9 documents (3+2+2+1+1 subjects)
subject counts:
DS 3 ['Asha', 'Kiran', 'Ravi']
Stats 3 ['Asha', 'Bhanu', 'Meena']
Python 2 ['Asha', 'Ravi']
R 1 ['Meena']
7 students, two with no subjects (empty array / missing field):
plain $unwind -> 9 docs, both LOST
preserveNullAndEmptyArrays: true -> 11 docs, both kept
'my count is short and I cannot see why' is nearly always this
includeArrayIndex 'pos', $match pos:0 -> each student's FIRST subject
$unwind on a string -> 5 documents (treated as one element)
$unwind on '$subject' -> 0 documents (a TYPO, and it does not error)
6 enrolments, one pointing at a course that does not exist:
after $lookup -> 6 rows, the orphan has course: []
after $lookup + $unwind -> 5 rows, the orphan is GONE
$lookup is a LEFT outer join; the $unwind after it turns it
into an inner one. Nothing warns you
enrolments -> student + course:
Asha Data Science with R A
Asha Statistical Foundations B
Kiran Web Technologies B
Meena Statistical Foundations A
Ravi Data Science with R C
courses -> enrolment counts, via $size on the joined array:
Data Science with R 2
Statistical Foundations 2
Web Technologies 1
$size needs no $unwind and no $group -- one stage, one pass
$lookup with let/pipeline: NotImplementedError in mongomock
the join-then-filter equivalent DOES run, and gives:
Data Science with R 1 grade-A enrolment(s)
Statistical Foundations 1 grade-A enrolment(s)
same answer, and on real data far more expensive: it drags
every enrolment in and then throws most of them away.
$$cid is the OUTER let variable, $course_id the inner field
bucket count names
[ 0, 40) -- (empty buckets are NOT emitted)
[ 40, 60) 1 ['Bhanu']
[ 60, 75) 2 ['Kiran', 'Ravi']
[ 75, 101) 2 ['Asha', 'Meena']
a student who scored exactly 100:
boundaries [..., 75, 101] -> lands in [75,101) CORRECT
boundaries [..., 75, 100] -> lands in 'Other' WRONG
boundaries [..., 75, 100], no default -> ERROR
buckets are [lower, upper): closed below, OPEN above. The top
boundary must exceed the maximum, so it is max+1, not max
$bucketAuto: NotImplementedError in mongomock
it picks boundaries to even out the COUNTS, so you never
state them. Good for a first look at unfamiliar data; wrong
for a report, because the boundaries MOVE when the data does
and last month's chart is no longer comparable with this one
$facet -> ONE document holding both results:
byDept [{'n': 3, '_id': 'DS'}, {'n': 2, '_id': 'Stats'}]
topThree ['Meena', 'Asha', 'Kiran']
one pass over the collection instead of two queries, which
is how a dashboard gets all its panels in a single round trip
Latha (an empty subjects array) and Mohan (none at all) have no maths mark either, so
$bucket puts them in Other. Changed: the $facet's count by department now ends with a
$sort, for the same reason as Experiment 15.
RESULT
Subject counts are DS 3, Stats 3, Python 2, R 1; $unwind drops Latha and Mohan; the buckets use 101 as the top boundary so a perfect score is kept.
Set up a replica set, and watch it replicate and fail over.
Initiate a three-member replica set, read from a secondary, force a failover, and configure the members.
WHAT TO DEMONSTRATE
What to demonstrate: rs.status() showing one PRIMARY and two SECONDARY;
rs.stepDown() triggering an election; writes failing during it; and w: "majority" versus
w: 1. The theory is Unit 5 §5.7 — and the question "why an odd number of members?" is asked
every year.
_drive_17_replication.py starts three mongod processes on ports 27017–27019 and a fourth on
27020, as section 0's route without Docker describes, and types each part of the script into a
shell on the member it names. It waits where you would wait — for the election after
rs.initiate(), and for a new primary after rs.stepDown() — and asserts what the experiment
shows. An election, the oplog and every time in rs.status() differ on every run, so this is
one run's output, recorded.
To do this on your own machine, docker compose with three mongo services is the easiest
route, or use Atlas — its free tier is a three-member replica set.
// Experiment 17 -- Replication: setting up and observing a replica set.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server:
// _drive_17_replication.py starts three mongod processes, as section 0
// describes, and runs each part of this file on the member it names. What each
// line printed is on the lab page, and
// tools/data-science/capture_lab_outputs.py runs it again. There is no .py
// half: mongomock is a library, not a server, and nothing about replication
// could stand in for three of them. (Until October 2026 mongod could not be
// installed where these labs are checked, and this file was desk-checked only.)
// =============================================================================
// Step 1: Get three servers
// 0. Getting three servers
// =============================================================================
// EASIEST -- MongoDB Atlas. The free tier IS a three-member replica set, which
// a local install is not. Connect with mongodb+srv:// and rs.status() works.
//
// LOCAL -- docker compose, three services on one network:
//
// services:
// mongo1: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27017:27017"] }
// mongo2: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27018:27017"] }
// mongo3: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27019:27017"] }
//
// docker compose up -d
// docker compose exec mongo1 mongosh
//
// WITHOUT DOCKER -- three mongod processes, three data directories, three ports:
//
// mkdir -p /data/rs0-{1,2,3}
// mongod --replSet rs0 --port 27017 --dbpath /data/rs0-1 --bind_ip localhost &
// mongod --replSet rs0 --port 27018 --dbpath /data/rs0-2 --bind_ip localhost &
// mongod --replSet rs0 --port 27019 --dbpath /data/rs0-3 --bind_ip localhost &
// =============================================================================
// Step 2: Initiate the set, and read its status
// 1. Initiating the set -- run ONCE, on ONE member
// =============================================================================
rs.initiate({
_id: "rs0",
members: [
{ _id: 0, host: "localhost:27017" },
{ _id: 1, host: "localhost:27018" },
{ _id: 2, host: "localhost:27019" }
]
})
// With docker compose, the hosts are the service names: mongo1:27017,
// mongo2:27017 and mongo3:27017.
// [Changed: the hosts were mongo1, mongo2 and mongo3, which exist only on the
// docker compose network. These are the three processes of the route without
// Docker, which is how this file is run.]
// The shell prompt changes: rs0 [direct: other] > ... then rs0 [primary] >
// Election takes a second or two. Until it finishes there is no primary and
// no writes are accepted.
rs.status() // members[], each with stateStr, health, optimeDate
rs.conf() // the configuration, with each member's votes and priority
rs.isMaster() // or db.hello() -- who is primary right now?
// WHAT TO SHOW THE EXAMINER in rs.status():
// * exactly ONE member with stateStr "PRIMARY"
// * two with "SECONDARY"
// * health: 1 on all three
// * optimeDate close together -- that closeness IS replication lag
// =============================================================================
// Step 3: Write on the primary, and read on a secondary
// 2. Watching data replicate
// =============================================================================
// On the PRIMARY:
use collegeDB
db.students.insertOne({ _id: 21, name: "Asha", dept: "DS" })
// On a SECONDARY -- a second terminal, connected to port 27018:
// mongosh --port 27018
use collegeDB
db.students.find() // Asha is there: the write has replicated
db.getMongo().getReadPref() // mode 'primary' -- and yet it read a secondary
// A shell connected to ONE member reads from that member, whatever the read
// preference says. Connected to the SET --
// mongosh "mongodb://localhost:27017,localhost:27018/?replicaSet=rs0"
// -- it sends every read to the primary, unless you say otherwise:
db.getMongo().setReadPref("secondaryPreferred")
// A secondary may be behind the primary, so reading from one is something
// you choose, by read preference, when slightly stale data is acceptable.
// [Corrected: this said the first find() fails, "not primary and
// secondaryOk=false", until setReadPref() is called. That was the old mongo
// shell. In mongosh 2, connected straight to a secondary, it reads, as above.
// It also had no use collegeDB: a new shell starts in test, where there is no
// Asha to find.]
// =============================================================================
// Step 4: Look at the oplog
// 3. The oplog -- how replication actually works
// =============================================================================
use local
db.oplog.rs.find().sort({ $natural: -1 }).limit(5)
db.oplog.rs.stats().maxSize // the CAP, in bytes
rs.printReplicationInfo() // oplog size and its time window
// The oplog is a CAPPED collection of idempotent operations. Secondaries tail
// it and replay it. Two consequences worth stating in the viva:
// * idempotent, so replaying an entry twice is safe -- which is what makes
// recovery after a crash possible at all
// * capped, so if a secondary falls further behind than the oplog's time
// window, it can no longer catch up and needs a FULL resync
// =============================================================================
// Step 5: Step the primary down
// 4. Failover -- the demonstration that earns the marks
// =============================================================================
rs.printSecondaryReplicationInfo() // lag per secondary, in seconds
rs.stepDown(60) // primary steps down for 60 seconds
// Watch: an election starts, a secondary becomes PRIMARY, and for roughly
// 10-30 seconds there is NO primary and every write fails. Show that gap.
// Or pull the plug, which is more convincing:
// docker compose stop mongo1
// rs.status() // mongo1 health 0, stateStr "(not reachable/healthy)"
// docker compose start mongo1 // it rejoins as a SECONDARY, not primary
// =============================================================================
// Step 6: Set the write concern and the read concern
// 5. Write concern and read concern -- the durability dial
// =============================================================================
// On the NEW primary: after the step-down, connect to whichever member
// rs.status() now shows as PRIMARY.
use collegeDB
// [Corrected: this use collegeDB was missing. After section 3's use local, the
// inserts below went into the local database, which is never replicated.]
db.students.insertOne({ _id: 22, name: "Ravi" },
{ writeConcern: { w: 1 } })
// Acknowledged by the PRIMARY only. Fast. Lost if the primary dies before the
// secondaries have it -- a "rollback".
db.students.insertOne({ _id: 23, name: "Meena" },
{ writeConcern: { w: "majority", j: true, wtimeout: 5000 } })
// Acknowledged by a MAJORITY, and on disk (j: true). Slower, and survives the
// loss of any one member. ALWAYS set wtimeout, or a stalled member hangs you.
db.students.find().readConcern("majority") // only data a majority holds
db.students.find().readConcern("local") // the default: may be rolled back
// | w | acknowledged by | survives primary loss? |
// |----------|---------------------|------------------------|
// | 0 | nobody (fire and forget) | no |
// | 1 | the primary | NO |
// | majority | 2 of 3 | yes |
// =============================================================================
// Step 7: Make a member hidden and delayed, and add an arbiter
// 6. Priority, hidden members and arbiters
// =============================================================================
cfg = rs.conf()
cfg.members[2].priority = 0 // never becomes primary
cfg.members[2].hidden = true // and clients never see it
cfg.members[2].secondaryDelaySecs = 3600 // an HOUR behind -- a live undo button
// [Corrected: this was slaveDelay, which MongoDB 5.0 renamed. MongoDB 8 refuses
// the old name: "BSON field 'MemberConfig.slaveDelay' is an unknown field".]
rs.reconfig(cfg)
// A delayed hidden member is the answer to "someone ran deleteMany({})".
// It has the data as it was an hour ago.
db.adminCommand({ setDefaultRWConcern: 1, defaultWriteConcern: { w: "majority" } })
rs.addArb("localhost:27020") // an ARBITER: votes, stores no data
// An arbiter changes what "majority" means without holding any data, so
// MongoDB 5 and later refuse to add one until the default write concern has
// been set explicitly -- the line before it.
// [Corrected: the setDefaultRWConcern line was missing, and without it
// rs.addArb() fails: "Reconfig attempted to install a config that would change
// the implicit default write concern".]
// =============================================================================
// WHY AN ODD NUMBER OF MEMBERS? -- asked every year
// =============================================================================
// A primary must be elected by a STRICT MAJORITY of votes.
//
// 3 members: majority 2, tolerates 1 failure
// 4 members: majority 3, tolerates 1 failure <-- no better than 3
// 5 members: majority 3, tolerates 2 failures
//
// The fourth member buys NO extra fault tolerance and adds a machine, network
// traffic and a chance of a tie. So: odd numbers.
//
// The majority rule also prevents SPLIT BRAIN. If the network partitions 3-2,
// only the side of 3 can elect a primary; the side of 2 has no majority and
// steps down to secondary. Two primaries accepting conflicting writes is
// impossible by construction, not by convention.
//
// This is Unit 5 §5.7, and it is CAP in practice: MongoDB chooses CONSISTENCY
// over availability, so the minority side refuses writes rather than diverge.
OUTPUT
[mongosh connected to 127.0.0.1:27017, not yet in a set]
test> rs.initiate({
... _id: "rs0",
... members: [
... { _id: 0, host: "localhost:27017" },
... { _id: 1, host: "localhost:27018" },
... { _id: 2, host: "localhost:27019" }
... ]
... })
{
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104251, i: 1 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104251, i: 1 })
}
[waited for the election: PRIMARY,SECONDARY,SECONDARY]
[mongosh connected to 127.0.0.1:27019, the primary]
rs0 [direct: primary] test> rs.status() // members[], each with stateStr, health, optimeDate
{
set: 'rs0',
date: ISODate('2026-10-04T08:57:46.891Z'),
myState: 1,
term: Long('1'),
syncSourceHost: '',
syncSourceId: -1,
heartbeatIntervalMillis: Long('2000'),
majorityVoteCount: 2,
writeMajorityCount: 2,
votingMembersCount: 3,
writableVotingMembersCount: 3,
optimes: {
lastCommittedOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
lastCommittedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
readConcernMajorityOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
appliedOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
durableOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
writtenOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z')
},
lastStableRecoveryTimestamp: Timestamp({ t: 1791104251, i: 1 }),
electionCandidateMetrics: {
lastElectionReason: 'electionTimeout',
lastElectionDate: ISODate('2026-10-04T08:57:42.079Z'),
electionTerm: Long('1'),
lastCommittedOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
lastSeenWrittenOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
lastSeenOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
numVotesNeeded: 2,
priorityAtElection: 1,
electionTimeoutMillis: Long('10000'),
numCatchUpOps: Long('0'),
newTermStartDate: ISODate('2026-10-04T08:57:42.106Z'),
wMajorityWriteAvailabilityDate: ISODate('2026-10-04T08:57:42.598Z')
},
members: [
{
_id: 0,
name: 'localhost:27017',
health: 1,
state: 2,
stateStr: 'SECONDARY',
uptime: 15,
optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeDurable: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
optimeDurableDate: ISODate('2026-10-04T08:57:42.000Z'),
optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastHeartbeat: ISODate('2026-10-04T08:57:46.098Z'),
lastHeartbeatRecv: ISODate('2026-10-04T08:57:46.599Z'),
pingMs: Long('0'),
lastHeartbeatMessage: '',
syncSourceHost: 'localhost:27019',
syncSourceId: 2,
infoMessage: '',
configVersion: 1,
configTerm: 1
},
{
_id: 1,
name: 'localhost:27018',
health: 1,
state: 2,
stateStr: 'SECONDARY',
uptime: 15,
optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeDurable: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
optimeDurableDate: ISODate('2026-10-04T08:57:42.000Z'),
optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastHeartbeat: ISODate('2026-10-04T08:57:46.092Z'),
lastHeartbeatRecv: ISODate('2026-10-04T08:57:45.094Z'),
pingMs: Long('0'),
lastHeartbeatMessage: '',
syncSourceHost: 'localhost:27019',
syncSourceId: 2,
infoMessage: '',
configVersion: 1,
configTerm: 1
},
{
_id: 2,
name: 'localhost:27019',
health: 1,
state: 1,
stateStr: 'PRIMARY',
uptime: 18,
optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
syncSourceHost: '',
syncSourceId: -1,
infoMessage: 'Could not find member to sync from',
electionTime: Timestamp({ t: 1791104262, i: 1 }),
electionDate: ISODate('2026-10-04T08:57:42.000Z'),
configVersion: 1,
configTerm: 1,
self: true,
lastHeartbeatMessage: ''
}
],
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104262, i: 17 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104262, i: 17 })
}
rs0 [direct: primary] test> rs.conf() // the configuration, with each member's votes and priority
{
_id: 'rs0',
version: 1,
term: 1,
members: [
{
_id: 0,
host: 'localhost:27017',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
},
{
_id: 1,
host: 'localhost:27018',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
},
{
_id: 2,
host: 'localhost:27019',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
}
],
protocolVersion: Long('1'),
writeConcernMajorityJournalDefault: true,
settings: {
chainingAllowed: true,
heartbeatIntervalMillis: 2000,
heartbeatTimeoutSecs: 10,
electionTimeoutMillis: 10000,
catchUpTimeoutMillis: -1,
catchUpTakeoverDelayMillis: 30000,
getLastErrorModes: {},
getLastErrorDefaults: { w: 1, wtimeout: 0 },
replicaSetId: ObjectId('6ac214fbd195cc8f74280fe8')
}
}
rs0 [direct: primary] test> rs.isMaster() // or db.hello() -- who is primary right now?
{
topologyVersion: { processId: ObjectId('6ac214f81bed474ff801f3e6'), counter: Long('6') },
hosts: [ 'localhost:27017', 'localhost:27018', 'localhost:27019' ],
setName: 'rs0',
setVersion: 1,
ismaster: true,
secondary: false,
primary: 'localhost:27019',
me: 'localhost:27019',
electionId: ObjectId('7fffffff0000000000000001'),
lastWrite: {
opTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
lastWriteDate: ISODate('2026-10-04T08:57:42.000Z'),
majorityOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
majorityWriteDate: ISODate('2026-10-04T08:57:42.000Z')
},
maxBsonObjectSize: 16777216,
maxMessageSizeBytes: 48000000,
maxWriteBatchSize: 100000,
localTime: ISODate('2026-10-04T08:57:46.953Z'),
logicalSessionTimeoutMinutes: 30,
connectionId: 26,
minWireVersion: 0,
maxWireVersion: 28,
readOnly: false,
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104262, i: 17 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104262, i: 17 }),
isWritablePrimary: true
}
rs0 [direct: primary] test> use collegeDB
switched to db collegeDB
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 21, name: "Asha", dept: "DS" })
{ acknowledged: true, insertedId: 21 }
[mongosh connected to 127.0.0.1:27018, a secondary]
rs0 [direct: secondary] test> use collegeDB
switched to db collegeDB
rs0 [direct: secondary] collegeDB> db.students.find() // Asha is there: the write has replicated
[ { _id: 21, name: 'Asha', dept: 'DS' } ]
rs0 [direct: secondary] collegeDB> db.getMongo().getReadPref() // mode 'primary' -- and yet it read a secondary
ReadPreference {
mode: 'primary',
tags: undefined,
hedge: undefined,
maxStalenessSeconds: undefined
}
rs0 [direct: secondary] collegeDB> db.getMongo().setReadPref("secondaryPreferred")
[mongosh connected to 127.0.0.1:27019, the primary]
rs0 [direct: primary] test> use local
switched to db local
rs0 [direct: primary] local> db.oplog.rs.find().sort({ $natural: -1 }).limit(5)
[
{
lsid: {
id: UUID('5d38bf63-2894-4ec9-9e62-a15c41016796'),
uid: Binary.createFromBase64('47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=', 0)
},
txnNumber: Long('1'),
op: 'i',
ns: 'collegeDB.students',
ui: UUID('bbcd39f2-8897-449a-9be1-e8bd273a7ff0'),
o: { _id: 21, name: 'Asha', dept: 'DS' },
o2: { _id: 21 },
stmtId: 0,
ts: Timestamp({ t: 1791104267, i: 2 }),
t: Long('1'),
v: Long('2'),
wall: ISODate('2026-10-04T08:57:47.089Z'),
prevOpTime: { ts: Timestamp({ t: 0, i: 0 }), t: Long('-1') }
},
{
op: 'c',
ns: 'collegeDB.$cmd',
ui: UUID('bbcd39f2-8897-449a-9be1-e8bd273a7ff0'),
o: {
create: 'students',
idIndex: { v: 2, key: { _id: 1 }, name: '_id_' }
},
o2: {
catalogId: Long('21'),
ident: 'a1566e98-ef1a-4d0a-adb2-031960a74fde',
idIndexIdent: 'ff62dc33-5d23-426b-9110-5af4399f2413'
},
versionContext: { OFCV: '8.3' },
ts: Timestamp({ t: 1791104267, i: 1 }),
t: Long('1'),
v: Long('2'),
wall: ISODate('2026-10-04T08:57:47.089Z')
},
{
op: 'c',
ns: 'config.$cmd',
ui: UUID('c69231f3-5a46-44c1-89e6-fb542ea67c9f'),
o: {
createIndexes: 'sampledQueriesDiff',
v: 2,
key: { expireAt: 1 },
name: 'SampledQueriesDiffTTLIndex',
expireAfterSeconds: 0
},
o2: { indexIdent: 'bf6df23a-ba1e-45c1-ba39-8e77f78c5ae1' },
ts: Timestamp({ t: 1791104262, i: 17 }),
t: Long('1'),
v: Long('2'),
wall: ISODate('2026-10-04T08:57:42.140Z')
},
{
op: 'c',
ns: 'config.$cmd',
ui: UUID('c69231f3-5a46-44c1-89e6-fb542ea67c9f'),
o: {
create: 'sampledQueriesDiff',
idIndex: { v: 2, key: { _id: 1 }, name: '_id_' }
},
o2: {
catalogId: Long('20'),
ident: 'f5ea0985-ff6c-4433-afa0-e359fa0d57a5',
idIndexIdent: 'eed90427-f9ea-4074-a290-de4b7efc9bbd'
},
versionContext: { OFCV: '8.3' },
ts: Timestamp({ t: 1791104262, i: 16 }),
t: Long('1'),
v: Long('2'),
wall: ISODate('2026-10-04T08:57:42.140Z')
},
{
op: 'c',
ns: 'config.$cmd',
ui: UUID('96827267-ad25-4d3c-9d1b-670dc16c0bf5'),
o: {
createIndexes: 'sampledQueries',
v: 2,
key: { expireAt: 1 },
name: 'SampledQueriesTTLIndex',
expireAfterSeconds: 0
},
o2: { indexIdent: '71e3a6c9-faf8-4928-9f86-babfe77fdbb3' },
ts: Timestamp({ t: 1791104262, i: 15 }),
t: Long('1'),
v: Long('2'),
wall: ISODate('2026-10-04T08:57:42.133Z')
}
]
rs0 [direct: primary] local> db.oplog.rs.stats().maxSize // the CAP, in bytes
1038090240
rs0 [direct: primary] local> rs.printReplicationInfo() // oplog size and its time window
actual oplog size
'990 MB'
---
configured oplog size
'990 MB'
---
log length start to end
'16 secs (0 hrs)'
---
oplog first event time
'Sun Oct 04 2026 08:57:31 GMT+0000 (Coordinated Universal Time)'
---
oplog last event time
'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)'
---
now
'Sun Oct 04 2026 08:57:53 GMT+0000 (Coordinated Universal Time)'
rs0 [direct: primary] local> rs.printSecondaryReplicationInfo() // lag per secondary, in seconds
source: localhost:27017
{
syncedTo: 'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)',
replLag: '0 secs (0 hrs) behind the primary '
}
---
source: localhost:27018
{
syncedTo: 'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)',
replLag: '0 secs (0 hrs) behind the primary '
}
rs0 [direct: primary] local> rs.stepDown(60) // primary steps down for 60 seconds
{
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104267, i: 2 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104267, i: 2 })
}
[waited for the election after the step-down: the primary is now 127.0.0.1:27017]
[mongosh connected to 127.0.0.1:27017, the new primary]
rs0 [direct: primary] test> use collegeDB
switched to db collegeDB
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 22, name: "Ravi" },
... { writeConcern: { w: 1 } })
{ acknowledged: true, insertedId: 22 }
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 23, name: "Meena" },
... { writeConcern: { w: "majority", j: true, wtimeout: 5000 } })
{ acknowledged: true, insertedId: 23 }
rs0 [direct: primary] collegeDB> db.students.find().readConcern("majority") // only data a majority holds
[
{ _id: 21, name: 'Asha', dept: 'DS' },
{ _id: 22, name: 'Ravi' },
{ _id: 23, name: 'Meena' }
]
rs0 [direct: primary] collegeDB> db.students.find().readConcern("local") // the default: may be rolled back
[
{ _id: 21, name: 'Asha', dept: 'DS' },
{ _id: 22, name: 'Ravi' },
{ _id: 23, name: 'Meena' }
]
rs0 [direct: primary] collegeDB> cfg = rs.conf()
{
_id: 'rs0',
version: 1,
term: 2,
members: [
{
_id: 0,
host: 'localhost:27017',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
},
{
_id: 1,
host: 'localhost:27018',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
},
{
_id: 2,
host: 'localhost:27019',
arbiterOnly: false,
buildIndexes: true,
hidden: false,
priority: 1,
tags: {},
secondaryDelaySecs: Long('0'),
votes: 1
}
],
protocolVersion: Long('1'),
writeConcernMajorityJournalDefault: true,
settings: {
chainingAllowed: true,
heartbeatIntervalMillis: 2000,
heartbeatTimeoutSecs: 10,
electionTimeoutMillis: 10000,
catchUpTimeoutMillis: -1,
catchUpTakeoverDelayMillis: 30000,
getLastErrorModes: {},
getLastErrorDefaults: { w: 1, wtimeout: 0 },
replicaSetId: ObjectId('6ac214fbd195cc8f74280fe8')
}
}
rs0 [direct: primary] collegeDB> cfg.members[2].priority = 0 // never becomes primary
0
rs0 [direct: primary] collegeDB> cfg.members[2].hidden = true // and clients never see it
true
rs0 [direct: primary] collegeDB> cfg.members[2].secondaryDelaySecs = 3600 // an HOUR behind -- a live undo button
3600
rs0 [direct: primary] collegeDB> rs.reconfig(cfg)
{
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104275, i: 3 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104275, i: 3 })
}
rs0 [direct: primary] collegeDB> db.adminCommand({ setDefaultRWConcern: 1, defaultWriteConcern: { w: "majority" } })
{
defaultReadConcern: { level: 'local' },
defaultWriteConcern: { w: 'majority', wtimeout: 0 },
updateOpTime: Timestamp({ t: 1791104275, i: 3 }),
updateWallClockTime: ISODate('2026-10-04T08:57:55.610Z'),
defaultWriteConcernSource: 'global',
defaultReadConcernSource: 'implicit',
localUpdateWallClockTime: ISODate('2026-10-04T08:57:55.621Z'),
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104275, i: 5 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104275, i: 5 })
}
rs0 [direct: primary] collegeDB> rs.addArb("localhost:27020") // an ARBITER: votes, stores no data
{
ok: 1,
'$clusterTime': {
clusterTime: Timestamp({ t: 1791104280, i: 1 }),
signature: {
hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
keyId: Long('0')
}
},
operationTime: Timestamp({ t: 1791104280, i: 1 })
}
[rs.status() at the end: 27017 PRIMARY, 27018 SECONDARY, 27019 SECONDARY, 27020 ARBITER]
checked: one PRIMARY and two SECONDARY members; Asha replicated to the secondary, which read her with the read preference 'primary'; after rs.stepDown() the primary moved from 27019 to 27017; the reconfiguration and the arbiter were accepted
Four corrections, each found by running it:
The secondary read. The script said db.students.find() on a secondary fails, "not primary
and secondaryOk=false", until setReadPref(). That was the old mongo shell: mongosh 2,
connected straight to a secondary, reads from it whatever the read preference says, as the
session shows. It is a connection to the whole set that sends reads to the primary. The
secondary's session also lacked use collegeDB, so it looked in test.
use collegeDB before section 5. After section 3's use local, the inserts went into the
local database, which is never replicated.
slaveDelay is now secondaryDelaySecs. MongoDB 5.0 renamed it, and MongoDB 8 rejects the
old name.
rs.addArb() needs setDefaultRWConcern first, in MongoDB 5 and later; without it the
arbiter is refused.
And the hosts: rs.initiate named mongo1–mongo3, which exist only on the Docker network; it
now names the local route's three processes, with the Docker form in a comment.
RESULT
One PRIMARY and two SECONDARY; Asha replicated to the secondary; after rs.stepDown() a new primary was elected; the hidden delayed member and the arbiter were accepted.
Store and retrieve a large file with GridFS.
Put a file into GridFS, see how it is stored, and get it back.
WHAT TO DEMONSTRATE
What to demonstrate: that a file appears as one document in fs.files
and many in fs.chunks, and that chunks == ceil(bytes / 261120). The
point to state: GridFS is for files over 16 MB, or where you need to read
ranges — for smaller files, object storage is usually better.
_drive_18_gridfs.py makes a 10 MB file and runs section 1's mongofiles commands for real —
put, list, and get to a copy, compared byte for byte — then types sections 2 to 4 into
mongosh, and deletes the file with mongofiles delete, as section 5 recommends.
// Experiment 18 -- GridFS: storing and retrieving large files.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server:
// _drive_18_gridfs.py puts a 10 MB file into GridFS with mongofiles, as section
// 1 describes, then types the rest of this file into mongosh. What each line
// printed is on the lab page, and tools/data-science/capture_lab_outputs.py
// runs it again. There is no .py half: mongomock does not implement GridFS.
// (Until October 2026 mongod could not be installed where these labs are
// checked, and this file was desk-checked only.)
// =============================================================================
// WHY GridFS EXISTS
// =============================================================================
// A BSON document is capped at 16 MB. GridFS gets round that by splitting the
// file into chunks of 255 KB and storing each chunk as its own document.
//
// fs.files -- ONE metadata document per file
// fs.chunks -- MANY documents, each a 255 KB slice, with files_id and n
//
// The 16 MB cap is not arbitrary: it is what keeps a single document cheap to
// move around in memory and over the wire. GridFS does not remove the cap; it
// works within it.
// =============================================================================
// Step 1: Put a file in GridFS with mongofiles
// 1. From the command line: mongofiles
// =============================================================================
// mongofiles -d collegeDB put lecture.mp4
// mongofiles -d collegeDB list
// mongofiles -d collegeDB get lecture.mp4
// mongofiles -d collegeDB delete lecture.mp4
// mongofiles -d collegeDB --local ./copy.mp4 get lecture.mp4
//
// mongofiles ships with the MongoDB Database Tools, a SEPARATE download from
// the server. That trips people up.
// =============================================================================
// Step 2: Look at what it stored, and count the chunks
// 2. Looking at what it stored
// =============================================================================
use collegeDB
db.fs.files.find().pretty()
// { _id, length, chunkSize: 261120, uploadDate, filename, metadata }
db.fs.chunks.find({}, { data: 0 }).sort({ n: 1 }).limit(3)
// { _id, files_id, n: 0 }, { ..., n: 1 }, ... -- n is the ORDER
const file = db.fs.files.findOne({ filename: "lecture.mp4" })
db.fs.chunks.countDocuments({ files_id: file._id }) // 41 for a 10 MB file
// [Corrected: this read files_id: <the _id from fs.files>, a placeholder, which
// mongosh rejects as a syntax error. The line before it looks the _id up.]
// --- THE ARITHMETIC, which is what gets asked ---------------------------------
// chunkSize is 255 KB = 255 * 1024 = 261120 bytes.
//
// chunks = ceil(length / 261120)
//
// 1 MB = 1048576 B -> ceil(1048576/261120) = 5 chunks
// 10 MB = 10485760 B -> ceil(10485760/261120) = 41 chunks
// 100 MB = 104857600 -> ceil(104857600/261120)= 402 chunks
// 1 GB = 1073741824 -> ceil(1073741824/261120)= 4113 chunks
//
// The LAST chunk is short -- GridFS does not pad. So the stored size is the
// file's size plus a little metadata, not a multiple of 255 KB.
// --- the indexes GridFS creates for itself -----------------------------------
db.fs.chunks.getIndexes() // { files_id: 1, n: 1 }, UNIQUE
db.fs.files.getIndexes() // { filename: 1, uploadDate: 1 }
// The unique compound index on (files_id, n) is what guarantees the chunks
// reassemble in the right order and cannot be duplicated.
// =============================================================================
// Step 3: Do it from a driver
// 3. From the shell / a driver
// =============================================================================
// mongosh has no built-in put; you use a driver. In Node:
//
// const bucket = new GridFSBucket(db, { bucketName: "lectures" });
// fs.createReadStream("lecture.mp4").pipe(bucket.openUploadStream("lecture.mp4",
// { metadata: { course: "DSC301", week: 3 } }));
// bucket.openDownloadStreamByName("lecture.mp4").pipe(res);
//
// A custom bucketName gives lectures.files / lectures.chunks instead of fs.*.
//
// STREAMING IS THE POINT. openDownloadStream can start at any byte:
// bucket.openDownloadStreamByName("lecture.mp4", { start: 5_000_000 })
// which is how a video seeks without downloading the whole file first.
// =============================================================================
// Step 4: Query by metadata
// 4. Querying by metadata -- what a filesystem cannot do
// =============================================================================
db.fs.files.find({ "metadata.course": "DSC301" })
db.fs.files.find({ length: { $gt: 50 * 1024 * 1024 } })
db.fs.files.aggregate([
{ $group: { _id: "$metadata.course", n: { $sum: 1 },
totalBytes: { $sum: "$length" } } },
{ $sort: { _id: 1 } }
])
// [Changed: the $sort was added. $group returns its groups in no fixed order.]
db.fs.files.createIndex({ "metadata.course": 1, uploadDate: -1 })
// =============================================================================
// Step 5: Delete it, with its chunks
// 5. Deleting -- the one thing to be careful about
// =============================================================================
// WRONG: this orphans every chunk of the file.
// db.fs.files.deleteOne({ filename: "lecture.mp4" })
//
// RIGHT: use the driver's bucket.delete(id) or mongofiles delete, which
// removes the metadata document AND its chunks. GridFS is two collections
// kept consistent by the DRIVER, not by the database -- there is no cascade.
// =============================================================================
// WHEN NOT TO USE GridFS -- worth a mark, and usually the right answer
// =============================================================================
// USE IT when:
// * files exceed 16 MB
// * you need RANGE reads (video seeking)
// * you want the file and its metadata under one backup and one replica set
// * atomic-ish updates matter more than throughput
//
// DO NOT use it when:
// * files are small and numerous -- the chunk documents cost more than they
// save, and a plain BinData field under 16 MB is simpler
// * you are serving them over HTTP at volume -- S3 or a CDN is faster,
// cheaper, and does not put the read load on your database
//
// The honest summary: GridFS is a good answer when the files must live with
// the data. It is rarely the best answer for a website's static assets.
OUTPUT
made lecture.mp4: 10,485,760 bytes
$ mongofiles -d collegeDB put lecture.mp4
(nothing printed)
$ mongofiles -d collegeDB list
lecture.mp4 10485760
$ mongofiles -d collegeDB --local ./copy.mp4 get lecture.mp4
(nothing printed)
copy.mp4 is the same as lecture.mp4, byte for byte: True
test> use collegeDB
switched to db collegeDB
collegeDB> db.fs.files.find().pretty()
[
{
_id: ObjectId('6ac2154da083e669e2dec979'),
length: Long('10485760'),
chunkSize: 261120,
uploadDate: ISODate('2026-10-04T08:58:53.727Z'),
filename: 'lecture.mp4',
metadata: {}
}
]
collegeDB> db.fs.chunks.find({}, { data: 0 }).sort({ n: 1 }).limit(3)
[
{
_id: ObjectId('6ac2154da083e669e2dec97a'),
files_id: ObjectId('6ac2154da083e669e2dec979'),
n: 0
},
{
_id: ObjectId('6ac2154da083e669e2dec97b'),
files_id: ObjectId('6ac2154da083e669e2dec979'),
n: 1
},
{
_id: ObjectId('6ac2154da083e669e2dec97c'),
files_id: ObjectId('6ac2154da083e669e2dec979'),
n: 2
}
]
collegeDB> const file = db.fs.files.findOne({ filename: "lecture.mp4" })
collegeDB> db.fs.chunks.countDocuments({ files_id: file._id }) // 41 for a 10 MB file
41
collegeDB> db.fs.chunks.getIndexes() // { files_id: 1, n: 1 }, UNIQUE
[
{ v: 2, key: { _id: 1 }, name: '_id_' },
{
v: 2,
key: { files_id: 1, n: 1 },
name: 'files_id_1_n_1',
unique: true
}
]
collegeDB> db.fs.files.getIndexes() // { filename: 1, uploadDate: 1 }
[
{ v: 2, key: { _id: 1 }, name: '_id_' },
{
v: 2,
key: { filename: 1, uploadDate: 1 },
name: 'filename_1_uploadDate_1'
}
]
collegeDB> db.fs.files.find({ "metadata.course": "DSC301" })
collegeDB> db.fs.files.find({ length: { $gt: 50 * 1024 * 1024 } })
collegeDB> db.fs.files.aggregate([
... { $group: { _id: "$metadata.course", n: { $sum: 1 },
... totalBytes: { $sum: "$length" } } },
... { $sort: { _id: 1 } }
... ])
[ { _id: null, n: 1, totalBytes: Long('10485760') } ]
collegeDB> db.fs.files.createIndex({ "metadata.course": 1, uploadDate: -1 })
metadata.course_1_uploadDate_-1
$ mongofiles -d collegeDB delete lecture.mp4
(nothing printed)
after mongofiles delete: 0 files, 0 chunks -- the metadata document AND its chunks
mongofiles cannot attach metadata, so section 4's queries on metadata.course find
nothing, as they would after any mongofiles put; a driver upload, as in section 3, sets it.
Corrected: the chunk count read { files_id: <the _id from fs.files> }, a placeholder that
mongosh rejects as a syntax error; the line before it now looks the _id up. Changed: the
metadata $group ends with a $sort.
RESULT
The 10 MB file is one document in fs.files and 41 in fs.chunks, as ceil(10485760 / 261120) says; the copy back is identical; mongofiles delete removes both.
Make a multi-document transaction, committed and aborted.
Transfer money between two accounts in a transaction, and see an abort undo everything.
WHAT TO DEMONSTRATE
Transactions are unavailable on a standalone mongod — they depend on the
oplog and majority commit, so a replica set is required. That fact is itself a
five-mark answer. This experiment runs on a one-member replica set, which is enough.
What to demonstrate: a transfer that commits, and one that aborts midway leaving both balances unchanged. And the point from Unit 5 §5.9: a schema that needs transactions for its common operations is usually one that should have embedded.
// Experiment 19 -- Multi-document ACID transactions.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. There is no .py half: transactions need a replica set, which mongomock
// is not. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
// =============================================================================
// THE PREREQUISITE, which is itself a five-mark answer
// =============================================================================
// A standalone mongod REFUSES to start a transaction. In mongosh 2, against
// MongoDB 8, the first write inside one fails with:
//
// MongoServerError: This MongoDB deployment does not support retryable
// writes. Please add retryWrites=false to your connection string.
//
// -- and with retryWrites=false in the connection string, still with that.
// [Corrected: this quoted "Transaction numbers are only allowed on a replica
// set member or mongos", which MongoDB 8 with mongosh 2 does not print. The
// message above is what a standalone gives, tried both ways.]
//
// WHY: a transaction's commit must be durable and visible atomically, and
// MongoDB implements that on the oplog with a majority write concern. A
// standalone has no oplog to write to and no majority to reach. So: set up
// experiment 17 first, or use Atlas, whose free tier is a replica set.
//
// Single-DOCUMENT operations have always been atomic, replica set or not.
// That is the point most students miss, and it is why most well-modelled
// MongoDB applications never need this feature at all.
// =============================================================================
// Step 1: Set up two accounts
// 1. Setting up something worth a transaction
// =============================================================================
use bankDB
db.accounts.drop()
db.accounts.insertMany([
{ _id: "A", holder: "Asha", balance: 5000 },
{ _id: "B", holder: "Ravi", balance: 3000 }
])
// A transfer touches TWO documents. Nothing about the document model makes
// that atomic, and here the two balances genuinely belong to two owners --
// so this is the case where embedding is not the answer.
// =============================================================================
// Step 2: Transfer, and commit
// 2. The core API -- a transfer that COMMITS
// =============================================================================
const session = db.getMongo().startSession()
const accounts = session.getDatabase("bankDB").accounts
session.startTransaction({
readConcern: { level: "snapshot" },
writeConcern: { w: "majority" }
})
try {
accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
accounts.updateOne({ _id: "B" }, { $inc: { balance: 500 } })
session.commitTransaction()
print("committed: A 4500, B 3500")
} catch (e) {
session.abortTransaction()
print("aborted: " + e)
throw e
} finally {
session.endSession()
}
// IMPORTANT: reads and writes must go through session.getDatabase(...).
// db.accounts.updateOne(...) inside the block is NOT in the transaction --
// it commits immediately, and nothing warns you. That is the single commonest
// mistake with this API.
// =============================================================================
// Step 3: Overdraw, and abort
// 3. The demonstration that matters: an ABORT leaves BOTH unchanged
// =============================================================================
// Show the balances before, run this, show them after. Nothing moved.
const s2 = db.getMongo().startSession()
const acc2 = s2.getDatabase("bankDB").accounts
s2.startTransaction()
try {
acc2.updateOne({ _id: "A" }, { $inc: { balance: -9999 } }) // overdraws
const a = acc2.findOne({ _id: "A" })
if (a.balance < 0) throw new Error("insufficient funds")
acc2.updateOne({ _id: "B" }, { $inc: { balance: 9999 } })
s2.commitTransaction()
} catch (e) {
s2.abortTransaction() // A's -9999 is UNDONE
print("aborted, both balances unchanged: " + e.message)
} finally { s2.endSession() }
db.accounts.find() // A 4500, B 3500: the first transfer, and nothing of the second
// [Added: the balances were only printed as text above. This reads them back.]
// Read A from ANOTHER shell while the transaction is open: you see the OLD
// balance. Uncommitted writes are invisible outside the session -- snapshot
// isolation, and the visible proof that this is a real transaction.
// =============================================================================
// Step 4: Retry on a transient error
// 4. Retrying -- required, not optional
// =============================================================================
// A transaction can fail with a TRANSIENT error (a write conflict, a failover
// mid-commit). The error carries a label, and the caller is expected to retry:
//
// e.hasErrorLabel("TransientTransactionError") -> retry the WHOLE thing
// e.hasErrorLabel("UnknownTransactionCommitResult") -> retry the COMMIT only
//
// Drivers wrap this: session.withTransaction(fn) retries for you and is what
// you should actually use.
//
// session.withTransaction(() => {
// accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
// accounts.updateOne({ _id: "B" }, { $inc: { balance: 500 } })
// })
//
// The callback must be IDEMPOTENT, because it may run more than once.
// =============================================================================
// 5. The limits
// =============================================================================
// * default 60-second time limit (transactionLifetimeLimitSeconds); a
// transaction that exceeds it is aborted
// * 16 MB of oplog entries per transaction
// * they hold locks, so long transactions block other writers
// * a sharded transaction is slower again -- it coordinates across shards
// * DDL inside a transaction is restricted: no createIndex, no dropDatabase;
// collection creation is allowed only from MongoDB 4.4 onward
// =============================================================================
// THE POINT, from Unit 5 §5.9
// =============================================================================
// Transactions arrived in MongoDB 4.0 and are genuinely ACID. But:
//
// IF YOUR COMMON OPERATIONS NEED THEM, YOUR SCHEMA IS PROBABLY WRONG.
//
// A student and their address, an order and its lines, a post and its
// comments -- embed those and every update is a single-document write, atomic
// with no transaction at all. Transactions are for the genuine cross-entity
// case: a bank transfer, where the two balances belong to different people and
// no amount of remodelling puts them in one document.
//
// Compare with Course 5: in SQL, transactions are how you do ordinary work.
// Here they are the escape hatch for the case the document model does not
// cover -- and reaching for it often is the signal to reconsider the model,
// or to ask whether the data was relational all along.
OUTPUT
rs0 [direct: primary] test> use bankDB
switched to db bankDB
rs0 [direct: primary] bankDB> db.accounts.drop()
true
rs0 [direct: primary] bankDB> db.accounts.insertMany([
... { _id: "A", holder: "Asha", balance: 5000 },
... { _id: "B", holder: "Ravi", balance: 3000 }
... ])
{ acknowledged: true, insertedIds: { '0': 'A', '1': 'B' } }
rs0 [direct: primary] bankDB> const session = db.getMongo().startSession()
rs0 [direct: primary] bankDB> const accounts = session.getDatabase("bankDB").accounts
rs0 [direct: primary] bankDB> session.startTransaction({
... readConcern: { level: "snapshot" },
... writeConcern: { w: "majority" }
... })
rs0 [direct: primary] bankDB> try {
... accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
... accounts.updateOne({ _id: "B" }, { $inc: { balance: 500 } })
... session.commitTransaction()
... print("committed: A 4500, B 3500")
... } catch (e) {
... session.abortTransaction()
... print("aborted: " + e)
... throw e
... } finally {
... session.endSession()
... }
committed: A 4500, B 3500
rs0 [direct: primary] bankDB> const s2 = db.getMongo().startSession()
rs0 [direct: primary] bankDB> const acc2 = s2.getDatabase("bankDB").accounts
rs0 [direct: primary] bankDB> s2.startTransaction()
rs0 [direct: primary] bankDB> try {
... acc2.updateOne({ _id: "A" }, { $inc: { balance: -9999 } }) // overdraws
... const a = acc2.findOne({ _id: "A" })
... if (a.balance < 0) throw new Error("insufficient funds")
... acc2.updateOne({ _id: "B" }, { $inc: { balance: 9999 } })
... s2.commitTransaction()
... } catch (e) {
... s2.abortTransaction() // A's -9999 is UNDONE
... print("aborted, both balances unchanged: " + e.message)
... } finally { s2.endSession() }
aborted, both balances unchanged: insufficient funds
rs0 [direct: primary] bankDB> db.accounts.find() // A 4500, B 3500: the first transfer, and nothing of the second
[
{ _id: 'A', holder: 'Asha', balance: 4500 },
{ _id: 'B', holder: 'Ravi', balance: 3500 }
]
Corrected: the script quoted a standalone server's refusal as "Transaction numbers
are only allowed on a replica set member or mongos". MongoDB 8 with mongosh 2 says instead "This
MongoDB deployment does not support retryable writes. Please add retryWrites=false to your
connection string" — and says it again with retryWrites=false. Added: a find() at the end
reads the balances back; both were only printed as text.
RESULT
The transfer committed, A 4500 and B 3500; the overdraft aborted, and both balances are unchanged, read back from the database.
Build a mini-application: a library management system.
Implement the library schema, issue and return books, run the reports, and check the stock counts stay true.
In mongosh, 20_case_study.js:
In Python, through mongomock, 20_case_study.py:
THE POINT
A library management system — the schema designed in practice.md Section C question 1 — exercising CRUD, aggregation and indexing together:
Issue a book: insert a loan and decrement availableCopies
conditionally (availableCopies: { $gt: 0 }), so a third copy of a
two-copy book cannot be lent.
Return it: set returned, compute any fine, increment the count back.
{ returned: null, due: { $lt: asAt } }.$match first.availableCopies is consistent with the count of unreturned
loans.That last check is the point of the experiment: the computed pattern speeds up the hottest query and introduces a value that can drift, and the only defence is to check it. The Python half asserts it at every step.
In mongosh, 20_case_study.js:
// Experiment 20 -- Case study: a library management system.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 20_case_study.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// The schema is the one designed in practice.md Section C question 1. Read
// that answer first: it justifies every embed and every reference, and this
// script only implements it.
use libraryDB
db.books.drop(); db.members.drop(); db.loans.drop()
// =============================================================================
// Step 1: Seed the books and members
// 1. SEED
// =============================================================================
db.books.insertMany([
{ _id: "978-1491954461", title: "MongoDB: The Definitive Guide",
authors: ["Shannon Bradshaw", "Kristina Chodorow"],
publisher: { name: "O'Reilly", year: 2019 },
subjects: ["databases", "nosql"], totalCopies: 5, availableCopies: 5 },
{ _id: "978-0134685991", title: "Effective Java",
authors: ["Joshua Bloch"], publisher: { name: "Addison-Wesley", year: 2018 },
subjects: ["programming", "java"], totalCopies: 2, availableCopies: 2 },
{ _id: "978-1449355739", title: "Learning Python",
authors: ["Mark Lutz"], publisher: { name: "O'Reilly", year: 2013 },
subjects: ["programming", "python"], totalCopies: 3, availableCopies: 3 }
])
db.members.insertMany([
{ _id: "M2026001", name: "Asha Kumari", email: "asha@nri.ac.in",
phones: ["9876543210"],
address: { city: "Vijayawada", state: "AP", pin: "520010" },
joined: ISODate("2026-07-01"), active: true, currentLoanCount: 0 },
{ _id: "M2026002", name: "Ravi Teja", email: "ravi@nri.ac.in",
phones: ["9876500000"],
address: { city: "Guntur", state: "AP", pin: "522002" },
joined: ISODate("2026-07-05"), active: true, currentLoanCount: 0 }
])
// =============================================================================
// Step 2: Create the indexes
// 2. INDEXES -- practice.md Step 4
// =============================================================================
db.books.createIndex({ title: "text", authors: "text" })
db.books.createIndex({ subjects: 1 }) // multikey
db.loans.createIndex({ member_id: 1, returned: 1 }) // query 2
db.loans.createIndex({ returned: 1, due: 1 }) // query 4, ESR
db.loans.createIndex({ isbn: 1, issued: -1 }) // query 5
db.members.createIndex({ email: 1 }, { unique: true })
// =============================================================================
// Step 3: Issue books, with a conditional decrement
// 3. ISSUE -- two writes, and a CONDITIONAL decrement
// =============================================================================
// A fixed "today", as in 20_case_study.py, so that a loan can be late.
const TODAY = ISODate("2026-08-26")
const DAY = 24 * 60 * 60 * 1000
// [Changed: TODAY was added, and issue() takes the day. Every loan was issued
// at new Date(), the moment the script ran, so no return could ever be late and
// the overdue report could never find anything.]
function issue(memberId, isbn, on = TODAY) {
const book = db.books.findOne({ _id: isbn })
const member = db.members.findOne({ _id: memberId })
// The guard is IN THE FILTER, not in an if. Checking availableCopies > 0 and
// then decrementing is two operations, and two concurrent borrowers both
// pass the check. This is one operation, so only one of them can match.
const dec = db.books.updateOne({ _id: isbn, availableCopies: { $gt: 0 } },
{ $inc: { availableCopies: -1 } })
if (dec.modifiedCount === 0) return { ok: false, why: "no copies available" }
const issued = on
const due = new Date(issued.getTime() + 14 * 24 * 60 * 60 * 1000)
db.loans.insertOne({
member_id: memberId, isbn,
book_title: book.title, // EXTENDED REFERENCE -- practice.md Step 3
member_name: member.name,
issued, due, returned: null, fine: 0 })
db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: 1 } })
return { ok: true }
}
issue("M2026001", "978-1491954461")
issue("M2026001", "978-0134685991")
issue("M2026002", "978-0134685991")
issue("M2026002", "978-1449355739")
// The third copy of a two-copy book:
issue("M2026001", "978-0134685991") // -> { ok: false, why: "no copies available" }
// =============================================================================
// Step 4: Return them, and charge the fines
// 4. RETURN -- set returned, compute the fine, put the copy back
// =============================================================================
function returnBook(memberId, isbn, on) {
const loan = db.loans.findOne({ member_id: memberId, isbn, returned: null })
if (!loan) return { ok: false, why: "no open loan" }
const daysLate = Math.max(0, Math.ceil((on - loan.due) / (24 * 60 * 60 * 1000)))
const fine = daysLate * 2 // Rs 2 per day
db.loans.updateOne({ _id: loan._id }, { $set: { returned: on, fine } })
db.books.updateOne({ _id: isbn }, { $inc: { availableCopies: 1 } })
db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: -1 } })
return { ok: true, daysLate, fine }
}
returnBook("M2026002", "978-1449355739", new Date(TODAY.getTime() + 10 * DAY)) // day 10 of 14
returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 20 * DAY)) // day 20: 6 late
returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 21 * DAY)) // again: refused
// [Changed: the return was at new Date(); the late return and the second
// return, which 20_case_study.py asserts, were added.]
// =============================================================================
// Step 5: Run the five reports
// 5. THE FIVE QUERIES -- practice.md Step 5
// =============================================================================
// 1. availability
db.books.findOne({ _id: "978-1491954461" }, { title: 1, availableCopies: 1 })
// 2. a member's current loans
db.loans.find({ member_id: "M2026001", returned: null })
// 4. overdue -- the extended reference pays for itself here: no $lookup
const asAt = new Date(TODAY.getTime() + 20 * DAY) // overdue as at day 20
db.loans.find({ returned: null, due: { $lt: asAt } }).sort({ due: 1 })
// [Corrected: the .sort(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find(), unsorted, without it, then rejected ".sort(...)" as an invalid command.]
// 5. most borrowed -- $match FIRST
db.loans.aggregate([
{ $match: { issued: { $gte: ISODate("2026-01-01") } } },
{ $group: { _id: "$isbn", title: { $first: "$book_title" },
times: { $sum: 1 } } },
{ $sort: { times: -1, _id: 1 } },
{ $limit: 10 }
])
// subject report -- multikey + $unwind
db.books.aggregate([
{ $unwind: "$subjects" },
{ $group: { _id: "$subjects", titles: { $push: "$title" },
n: { $sum: 1 } } },
{ $sort: { n: -1, _id: 1 } }
])
// =============================================================================
// Step 6: Check the stock counts against the loans
// 6. THE INTEGRITY CHECK -- the point of the whole experiment
// =============================================================================
// availableCopies is the COMPUTED pattern: it makes query 1 constant-time and
// introduces a number that can DRIFT. The only defence is to check it.
db.books.aggregate([
{ $lookup: {
from: "loans", localField: "_id", foreignField: "isbn", as: "loans" } },
{ $project: {
title: 1, totalCopies: 1, availableCopies: 1,
out: { $size: { $filter: { input: "$loans", as: "l",
cond: { $eq: ["$$l.returned", null] } } } } } },
{ $addFields: { expected: { $subtract: ["$totalCopies", "$out"] } } },
{ $match: { $expr: { $ne: ["$availableCopies", "$expected"] } } }
])
// This should return NOTHING. Anything it returns is a book whose stored count
// disagrees with its open loans -- run it nightly, and alert on any row.
In Python, through mongomock, 20_case_study.py:
"""Experiment 20 — Case study: a library management system.
The schema is the one designed in practice.md Section C question 1, and this
runs the whole workflow against it: seed, issue, return, the five reports, and
the integrity check.
The integrity check is the point of the experiment. availableCopies is the
COMPUTED pattern -- it makes the availability lookup constant-time and
introduces a number that can drift out of step with the loans. So it is
asserted after every single write, not once at the end. When the last section
deliberately breaks it, the same check catches it.
"""
import datetime as dt
import mongomock
# This experiment does NOT use fixtures.py. The other nineteen share the
# collegeDB sample data; this one designs its own schema from scratch, which
# is the exercise.
DAY = dt.timedelta(days=1)
LOAN_DAYS = 14
FINE_PER_DAY = 2
BOOKS = [
{"_id": "978-1491954461", "title": "MongoDB: The Definitive Guide",
"authors": ["Shannon Bradshaw", "Kristina Chodorow"],
"publisher": {"name": "O'Reilly", "year": 2019},
"subjects": ["databases", "nosql"], "totalCopies": 5, "availableCopies": 5},
{"_id": "978-0134685991", "title": "Effective Java",
"authors": ["Joshua Bloch"],
"publisher": {"name": "Addison-Wesley", "year": 2018},
"subjects": ["programming", "java"], "totalCopies": 2, "availableCopies": 2},
{"_id": "978-1449355739", "title": "Learning Python",
"authors": ["Mark Lutz"], "publisher": {"name": "O'Reilly", "year": 2013},
"subjects": ["programming", "python"], "totalCopies": 3, "availableCopies": 3},
]
MEMBERS = [
{"_id": "M2026001", "name": "Asha Kumari", "email": "asha@nri.ac.in",
"phones": ["9876543210"],
"address": {"city": "Vijayawada", "state": "AP", "pin": "520010"},
"joined": dt.datetime(2026, 7, 1), "active": True, "currentLoanCount": 0},
{"_id": "M2026002", "name": "Ravi Teja", "email": "ravi@nri.ac.in",
"phones": ["9876500000"],
"address": {"city": "Guntur", "state": "AP", "pin": "522002"},
"joined": dt.datetime(2026, 7, 5), "active": True, "currentLoanCount": 0},
]
# A fixed "today", so the overdue report is reproducible instead of drifting.
TODAY = dt.datetime(2026, 8, 26)
def seed():
db = mongomock.MongoClient().libraryDB
db.books.insert_many([dict(b) for b in BOOKS])
db.members.insert_many([dict(m) for m in MEMBERS])
db.books.create_index("subjects")
db.loans.create_index([("member_id", 1), ("returned", 1)])
db.loans.create_index([("returned", 1), ("due", 1)])
db.loans.create_index([("isbn", 1), ("issued", -1)])
db.members.create_index("email", unique=True)
return db
# =============================================================================
# The two operations
# =============================================================================
def issue(db, member_id, isbn, on=TODAY):
"""Insert a loan and decrement the stored count -- CONDITIONALLY."""
book = db.books.find_one({"_id": isbn})
member = db.members.find_one({"_id": member_id})
if book is None or member is None:
return {"ok": False, "why": "unknown book or member"}
# The guard lives IN THE FILTER. Reading availableCopies, testing it, then
# decrementing is two operations: two concurrent borrowers both pass the
# test and both decrement, and the count goes negative. One operation
# cannot -- only one of them matches.
dec = db.books.update_one({"_id": isbn, "availableCopies": {"$gt": 0}},
{"$inc": {"availableCopies": -1}})
if dec.modified_count == 0:
return {"ok": False, "why": "no copies available"}
db.loans.insert_one({
"member_id": member_id, "isbn": isbn,
"book_title": book["title"], # extended reference
"member_name": member["name"], # extended reference
"issued": on, "due": on + LOAN_DAYS * DAY,
"returned": None, "fine": 0})
db.members.update_one({"_id": member_id},
{"$inc": {"currentLoanCount": 1}})
return {"ok": True}
def return_book(db, member_id, isbn, on=TODAY):
loan = db.loans.find_one({"member_id": member_id, "isbn": isbn,
"returned": None})
if loan is None:
return {"ok": False, "why": "no open loan"}
days_late = max(0, (on - loan["due"]).days)
fine = days_late * FINE_PER_DAY
db.loans.update_one({"_id": loan["_id"]},
{"$set": {"returned": on, "fine": fine}})
db.books.update_one({"_id": isbn}, {"$inc": {"availableCopies": 1}})
db.members.update_one({"_id": member_id},
{"$inc": {"currentLoanCount": -1}})
return {"ok": True, "daysLate": days_late, "fine": fine}
# =============================================================================
# The integrity check -- run after EVERY write
# =============================================================================
def drifted(db):
"""Books whose stored availableCopies disagrees with their open loans.
Empty is the only acceptable answer. This is the whole reason the computed
pattern is safe to use: it is cheap to verify, so verify it.
"""
bad = []
for b in db.books.find():
out = db.loans.count_documents({"isbn": b["_id"], "returned": None})
expected = b["totalCopies"] - out
if b["availableCopies"] != expected:
bad.append({"isbn": b["_id"], "title": b["title"],
"stored": b["availableCopies"], "expected": expected,
"openLoans": out})
return bad
def members_drifted(db):
"""The same check for currentLoanCount -- a second computed field."""
bad = []
for m in db.members.find():
out = db.loans.count_documents({"member_id": m["_id"], "returned": None})
if m["currentLoanCount"] != out:
bad.append({"member": m["_id"], "stored": m["currentLoanCount"],
"expected": out})
return bad
def consistent(db, step):
assert drifted(db) == [], (step, drifted(db))
assert members_drifted(db) == [], (step, members_drifted(db))
# =============================================================================
# The workflow
# =============================================================================
def the_happy_path(db):
consistent(db, "after seeding")
assert issue(db, "M2026001", "978-1491954461")["ok"]
consistent(db, "after issue 1")
assert issue(db, "M2026001", "978-0134685991")["ok"]
consistent(db, "after issue 2")
assert issue(db, "M2026002", "978-0134685991")["ok"]
consistent(db, "after issue 3")
assert issue(db, "M2026002", "978-1449355739")["ok"]
consistent(db, "after issue 4")
assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 0
assert db.members.find_one({"_id": "M2026001"})["currentLoanCount"] == 2
assert db.loans.count_documents({"returned": None}) == 4
print(" 4 issues, and availableCopies / currentLoanCount agree with the")
print(" loans collection after every single one:")
for b in db.books.find().sort("_id", 1):
print(f" {b['title']:34s} {b['availableCopies']}/{b['totalCopies']} available")
def the_sixth_copy_of_a_five_copy_book(db):
"""The conditional decrement, and why it is not an if statement."""
# Effective Java has 2 copies and both are out.
before = db.books.find_one({"_id": "978-0134685991"})["availableCopies"]
assert before == 0
result = issue(db, "M2026001", "978-0134685991")
assert result == {"ok": False, "why": "no copies available"}, result
assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 0, \
"the count must not go NEGATIVE"
assert db.loans.count_documents({"isbn": "978-0134685991"}) == 2, \
"and no loan row was written for the refused issue"
consistent(db, "after a refused issue")
print(" a third issue of a 2-copy book -> refused, count stayed at 0,")
print(" no loan row written")
print(" the guard is { _id: isbn, availableCopies: { $gt: 0 } } in the")
print(" FILTER. As an if-then-decrement it is two operations, and two")
print(" concurrent borrowers both pass the test. As one update, they")
print(" cannot: the second one matches nothing")
def returning_on_time_and_late(db):
# On time: issued TODAY, due TODAY+14, returned TODAY+10.
on_time = return_book(db, "M2026002", "978-1449355739", TODAY + 10 * DAY)
assert on_time == {"ok": True, "daysLate": 0, "fine": 0}, on_time
consistent(db, "after an on-time return")
assert db.books.find_one({"_id": "978-1449355739"})["availableCopies"] == 3
# Late: returned TODAY+20, six days past the due date.
late = return_book(db, "M2026002", "978-0134685991", TODAY + 20 * DAY)
assert late == {"ok": True, "daysLate": 6, "fine": 12}, late
assert 6 * FINE_PER_DAY == 12
consistent(db, "after a late return")
assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 1
# Returning something not on loan changes nothing.
again = return_book(db, "M2026002", "978-0134685991", TODAY + 21 * DAY)
assert again == {"ok": False, "why": "no open loan"}, again
consistent(db, "after a duplicate return")
print(" returned day 10 of a 14-day loan -> 0 days late, fine 0")
print(" returned day 20 of a 14-day loan -> 6 days late, fine Rs 12")
print(" returning it a second time -> refused, nothing changed")
print(" that last one matters: without the returned: null in the")
print(" filter, a double return increments availableCopies twice and")
print(" the library thinks it owns a copy it does not have")
def the_five_reports(db):
# 1. availability -- one document, no join
b = db.books.find_one({"_id": "978-1491954461"},
{"title": 1, "availableCopies": 1})
assert b == {"_id": "978-1491954461",
"title": "MongoDB: The Definitive Guide",
"availableCopies": 4}, b
# 2. a member's open loans
asha = list(db.loans.find({"member_id": "M2026001", "returned": None}))
assert len(asha) == 2
assert {l["isbn"] for l in asha} == {"978-1491954461", "978-0134685991"}
# 4. overdue, as at TODAY + 20 days
as_at = TODAY + 20 * DAY
overdue = list(db.loans.find({"returned": None, "due": {"$lt": as_at}})
.sort("due", 1))
assert len(overdue) == 2, overdue
# The extended reference is what makes this report joinless.
assert all("book_title" in l and "member_name" in l for l in overdue)
assert {l["member_name"] for l in overdue} == {"Asha Kumari"}
# 5. most borrowed
top = list(db.loans.aggregate([
{"$match": {"issued": {"$gte": dt.datetime(2026, 1, 1)}}},
{"$group": {"_id": "$isbn", "title": {"$first": "$book_title"},
"times": {"$sum": 1}}},
{"$sort": {"times": -1, "_id": 1}},
{"$limit": 10}]))
assert [(t["title"], t["times"]) for t in top] == [
("Effective Java", 2),
("Learning Python", 1),
("MongoDB: The Definitive Guide", 1)], top
# subjects -- multikey plus $unwind
subs = list(db.books.aggregate([
{"$unwind": "$subjects"},
{"$group": {"_id": "$subjects", "n": {"$sum": 1}}},
{"$sort": {"n": -1, "_id": 1}}]))
assert {s["_id"]: s["n"] for s in subs} == \
{"programming": 2, "databases": 1, "java": 1, "nosql": 1, "python": 1}
print(f" 1. availability MongoDB Definitive Guide {b['availableCopies']}/5")
print(f" 2. Asha's loans {len(asha)} open")
print(f" 4. overdue at {as_at.date()} {len(overdue)}, both Asha's, no $lookup")
print(" 5. most borrowed " +
", ".join(f"{t['title'].split(':')[0]} x{t['times']}" for t in top))
print(" report 4 reads ONE collection because book_title and")
print(" member_name were copied onto the loan -- the extended")
print(" reference pattern paying for itself")
def when_it_drifts_the_check_catches_it(db):
"""Break it on purpose. A check that has never failed is not a check."""
assert drifted(db) == []
# Exactly the bug the transaction in practice.md Step 5 prevents: the loan
# was written and the decrement was not.
db.loans.insert_one({"member_id": "M2026002", "isbn": "978-1491954461",
"book_title": "MongoDB: The Definitive Guide",
"member_name": "Ravi Teja",
"issued": TODAY, "due": TODAY + LOAN_DAYS * DAY,
"returned": None, "fine": 0})
bad = drifted(db)
assert len(bad) == 1, bad
assert bad[0]["isbn"] == "978-1491954461"
assert bad[0]["stored"] == 4 and bad[0]["expected"] == 3, bad
assert members_drifted(db) == [{"member": "M2026002",
"stored": 0, "expected": 1}], \
members_drifted(db)
print(" a loan written with the decrement missing -- a half-done issue:")
print(f" {bad[0]['title']}: stored {bad[0]['stored']}, "
f"expected {bad[0]['expected']} ({bad[0]['openLoans']} open loans)")
print(" M2026002: currentLoanCount 0, expected 1")
print(" nothing errored. The reports still ran. Only the check found")
print(" it -- which is why it runs nightly, and why practice.md puts")
print(" the two writes in a TRANSACTION in the first place")
# Repair, and confirm.
db.books.update_one({"_id": "978-1491954461"},
{"$inc": {"availableCopies": -1}})
db.members.update_one({"_id": "M2026002"},
{"$inc": {"currentLoanCount": 1}})
consistent(db, "after repair")
print(" repaired, and both checks are clean again")
def main():
print("Experiment 20 -- Library management case study")
print(" schema: practice.md Section C question 1")
# ONE database, carried through the whole workflow -- each stage builds on
# the state the last one left, exactly as the real application would.
# Step 1: Seed the books, members and loans
db = seed()
# Step 2: Issue and return a book
the_happy_path(db)
# Step 3: Refuse a sixth copy of a five-copy book
the_sixth_copy_of_a_five_copy_book(db)
# Step 4: Return on time and late
returning_on_time_and_late(db)
# Step 5: Run the five reports
the_five_reports(db)
# Step 6: Break the stock count, and catch it
when_it_drifts_the_check_catches_it(db)
if __name__ == "__main__":
main()
In mongosh, 20_case_study.js:
OUTPUT
test> use libraryDB
switched to db libraryDB
libraryDB> db.books.drop(); db.members.drop(); db.loans.drop()
true
libraryDB> db.books.insertMany([
... { _id: "978-1491954461", title: "MongoDB: The Definitive Guide",
... authors: ["Shannon Bradshaw", "Kristina Chodorow"],
... publisher: { name: "O'Reilly", year: 2019 },
... subjects: ["databases", "nosql"], totalCopies: 5, availableCopies: 5 },
... { _id: "978-0134685991", title: "Effective Java",
... authors: ["Joshua Bloch"], publisher: { name: "Addison-Wesley", year: 2018 },
... subjects: ["programming", "java"], totalCopies: 2, availableCopies: 2 },
... { _id: "978-1449355739", title: "Learning Python",
... authors: ["Mark Lutz"], publisher: { name: "O'Reilly", year: 2013 },
... subjects: ["programming", "python"], totalCopies: 3, availableCopies: 3 }
... ])
{
acknowledged: true,
insertedIds: { '0': '978-1491954461', '1': '978-0134685991', '2': '978-1449355739' }
}
libraryDB> db.members.insertMany([
... { _id: "M2026001", name: "Asha Kumari", email: "asha@nri.ac.in",
... phones: ["9876543210"],
... address: { city: "Vijayawada", state: "AP", pin: "520010" },
... joined: ISODate("2026-07-01"), active: true, currentLoanCount: 0 },
... { _id: "M2026002", name: "Ravi Teja", email: "ravi@nri.ac.in",
... phones: ["9876500000"],
... address: { city: "Guntur", state: "AP", pin: "522002" },
... joined: ISODate("2026-07-05"), active: true, currentLoanCount: 0 }
... ])
{
acknowledged: true,
insertedIds: { '0': 'M2026001', '1': 'M2026002' }
}
libraryDB> db.books.createIndex({ title: "text", authors: "text" })
title_text_authors_text
libraryDB> db.books.createIndex({ subjects: 1 }) // multikey
subjects_1
libraryDB> db.loans.createIndex({ member_id: 1, returned: 1 }) // query 2
member_id_1_returned_1
libraryDB> db.loans.createIndex({ returned: 1, due: 1 }) // query 4, ESR
returned_1_due_1
libraryDB> db.loans.createIndex({ isbn: 1, issued: -1 }) // query 5
isbn_1_issued_-1
libraryDB> db.members.createIndex({ email: 1 }, { unique: true })
email_1
libraryDB> const TODAY = ISODate("2026-08-26")
libraryDB> const DAY = 24 * 60 * 60 * 1000
libraryDB> function issue(memberId, isbn, on = TODAY) {
... const book = db.books.findOne({ _id: isbn })
... const member = db.members.findOne({ _id: memberId })
...
... // The guard is IN THE FILTER, not in an if. Checking availableCopies > 0 and
... // then decrementing is two operations, and two concurrent borrowers both
... // pass the check. This is one operation, so only one of them can match.
... const dec = db.books.updateOne({ _id: isbn, availableCopies: { $gt: 0 } },
... { $inc: { availableCopies: -1 } })
... if (dec.modifiedCount === 0) return { ok: false, why: "no copies available" }
...
... const issued = on
... const due = new Date(issued.getTime() + 14 * 24 * 60 * 60 * 1000)
... db.loans.insertOne({
... member_id: memberId, isbn,
... book_title: book.title, // EXTENDED REFERENCE -- practice.md Step 3
... member_name: member.name,
... issued, due, returned: null, fine: 0 })
... db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: 1 } })
... return { ok: true }
... }
[Function: issue]
libraryDB> issue("M2026001", "978-1491954461")
{ ok: true }
libraryDB> issue("M2026001", "978-0134685991")
{ ok: true }
libraryDB> issue("M2026002", "978-0134685991")
{ ok: true }
libraryDB> issue("M2026002", "978-1449355739")
{ ok: true }
libraryDB> issue("M2026001", "978-0134685991") // -> { ok: false, why: "no copies available" }
{ ok: false, why: 'no copies available' }
libraryDB> function returnBook(memberId, isbn, on) {
... const loan = db.loans.findOne({ member_id: memberId, isbn, returned: null })
... if (!loan) return { ok: false, why: "no open loan" }
...
... const daysLate = Math.max(0, Math.ceil((on - loan.due) / (24 * 60 * 60 * 1000)))
... const fine = daysLate * 2 // Rs 2 per day
...
... db.loans.updateOne({ _id: loan._id }, { $set: { returned: on, fine } })
... db.books.updateOne({ _id: isbn }, { $inc: { availableCopies: 1 } })
... db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: -1 } })
... return { ok: true, daysLate, fine }
... }
[Function: returnBook]
libraryDB> returnBook("M2026002", "978-1449355739", new Date(TODAY.getTime() + 10 * DAY)) // day 10 of 14
{ ok: true, daysLate: 0, fine: 0 }
libraryDB> returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 20 * DAY)) // day 20: 6 late
{ ok: true, daysLate: 6, fine: 12 }
libraryDB> returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 21 * DAY)) // again: refused
{ ok: false, why: 'no open loan' }
libraryDB> db.books.findOne({ _id: "978-1491954461" }, { title: 1, availableCopies: 1 })
{
_id: '978-1491954461',
title: 'MongoDB: The Definitive Guide',
availableCopies: 4
}
libraryDB> db.loans.find({ member_id: "M2026001", returned: null })
[
{
_id: ObjectId('6ac2155e99e081c7eead4d4c'),
member_id: 'M2026001',
isbn: '978-1491954461',
book_title: 'MongoDB: The Definitive Guide',
member_name: 'Asha Kumari',
issued: ISODate('2026-08-26T00:00:00.000Z'),
due: ISODate('2026-09-09T00:00:00.000Z'),
returned: null,
fine: 0
},
{
_id: ObjectId('6ac2155e99e081c7eead4d4d'),
member_id: 'M2026001',
isbn: '978-0134685991',
book_title: 'Effective Java',
member_name: 'Asha Kumari',
issued: ISODate('2026-08-26T00:00:00.000Z'),
due: ISODate('2026-09-09T00:00:00.000Z'),
returned: null,
fine: 0
}
]
libraryDB> const asAt = new Date(TODAY.getTime() + 20 * DAY) // overdue as at day 20
libraryDB> db.loans.find({ returned: null, due: { $lt: asAt } }).sort({ due: 1 })
[
{
_id: ObjectId('6ac2155e99e081c7eead4d4c'),
member_id: 'M2026001',
isbn: '978-1491954461',
book_title: 'MongoDB: The Definitive Guide',
member_name: 'Asha Kumari',
issued: ISODate('2026-08-26T00:00:00.000Z'),
due: ISODate('2026-09-09T00:00:00.000Z'),
returned: null,
fine: 0
},
{
_id: ObjectId('6ac2155e99e081c7eead4d4d'),
member_id: 'M2026001',
isbn: '978-0134685991',
book_title: 'Effective Java',
member_name: 'Asha Kumari',
issued: ISODate('2026-08-26T00:00:00.000Z'),
due: ISODate('2026-09-09T00:00:00.000Z'),
returned: null,
fine: 0
}
]
libraryDB> db.loans.aggregate([
... { $match: { issued: { $gte: ISODate("2026-01-01") } } },
... { $group: { _id: "$isbn", title: { $first: "$book_title" },
... times: { $sum: 1 } } },
... { $sort: { times: -1, _id: 1 } },
... { $limit: 10 }
... ])
[
{ _id: '978-0134685991', title: 'Effective Java', times: 2 },
{ _id: '978-1449355739', title: 'Learning Python', times: 1 },
{
_id: '978-1491954461',
title: 'MongoDB: The Definitive Guide',
times: 1
}
]
libraryDB> db.books.aggregate([
... { $unwind: "$subjects" },
... { $group: { _id: "$subjects", titles: { $push: "$title" },
... n: { $sum: 1 } } },
... { $sort: { n: -1, _id: 1 } }
... ])
[
{
_id: 'programming',
titles: [ 'Effective Java', 'Learning Python' ],
n: 2
},
{ _id: 'databases', titles: [ 'MongoDB: The Definitive Guide' ], n: 1 },
{ _id: 'java', titles: [ 'Effective Java' ], n: 1 },
{ _id: 'nosql', titles: [ 'MongoDB: The Definitive Guide' ], n: 1 },
{ _id: 'python', titles: [ 'Learning Python' ], n: 1 }
]
libraryDB> db.books.aggregate([
... { $lookup: {
... from: "loans", localField: "_id", foreignField: "isbn", as: "loans" } },
... { $project: {
... title: 1, totalCopies: 1, availableCopies: 1,
... out: { $size: { $filter: { input: "$loans", as: "l",
... cond: { $eq: ["$$l.returned", null] } } } } } },
... { $addFields: { expected: { $subtract: ["$totalCopies", "$out"] } } },
... { $match: { $expr: { $ne: ["$availableCopies", "$expected"] } } }
... ])
In Python, through mongomock, 20_case_study.py:
OUTPUT
Experiment 20 -- Library management case study
schema: practice.md Section C question 1
4 issues, and availableCopies / currentLoanCount agree with the
loans collection after every single one:
Effective Java 0/2 available
Learning Python 2/3 available
MongoDB: The Definitive Guide 4/5 available
a third issue of a 2-copy book -> refused, count stayed at 0,
no loan row written
the guard is { _id: isbn, availableCopies: { $gt: 0 } } in the
FILTER. As an if-then-decrement it is two operations, and two
concurrent borrowers both pass the test. As one update, they
cannot: the second one matches nothing
returned day 10 of a 14-day loan -> 0 days late, fine 0
returned day 20 of a 14-day loan -> 6 days late, fine Rs 12
returning it a second time -> refused, nothing changed
that last one matters: without the returned: null in the
filter, a double return increments availableCopies twice and
the library thinks it owns a copy it does not have
1. availability MongoDB Definitive Guide 4/5
2. Asha's loans 2 open
4. overdue at 2026-09-15 2, both Asha's, no $lookup
5. most borrowed Effective Java x2, Learning Python x1, MongoDB x1
report 4 reads ONE collection because book_title and
member_name were copied onto the loan -- the extended
reference pattern paying for itself
a loan written with the decrement missing -- a half-done issue:
MongoDB: The Definitive Guide: stored 4, expected 3 (2 open loans)
M2026002: currentLoanCount 0, expected 1
nothing errored. The reports still ran. Only the check found
it -- which is why it runs nightly, and why practice.md puts
the two writes in a TRANSACTION in the first place
repaired, and both checks are clean again
Changed: loans were issued at new Date(), the moment the script ran, so no return
could be late and the overdue report could never find anything. The script now issues on a fixed
day, 26 August 2026, as the Python half does, and returns one book on time and one six days late.
Corrected: the overdue report's .sort(...) began its own line, which mongosh rejected. And
this page said the refused loan was "the sixth copy of a five-copy book"; it is the third of
Effective Java's two copies.
RESULT
A third loan of a two-copy book is refused; returns on day 10 and day 20 are fined Rs 0 and Rs 12; two loans are overdue at day 20; the integrity check finds no drift.
An hour, a dataset, one experiment number, then a viva.
What costs marks:
updateOne where updateMany was meant — silent, and reports successreplaceOne destroying every other field$elemMatch for two conditions on an array of sub-documents$unwind before grouping on array contents$unwind after $lookup and getting an array$match after $group when it could have come first$bucket's top boundary to the maximum value and losing it{ f: null } matches missing fields too.sort(...) or .explain(...) at the shell: mongosh runs the line above
on its own, and rejects this oneWhat earns them:
Translate to SQL out loud. "This $group is a GROUP BY, and this second
$match is the HAVING." It shows you understand the pipeline rather than
having memorised it.
Run explain("executionStats") and quote
totalDocsExamined / nReturned. That ratio is the answer to "is this query
fast?", and wall-clock time on a five-document collection is not.
State the embed-or-reference decision and its cost. "I embedded the address because it is one-to-one and bounded; I referenced the courses because embedding would duplicate the title across 300 students, and renaming the instructor would then be 300 updates."
Say what is not enforced. Nothing stops a reference pointing at a deleted document. In Database Management Systems the database guaranteed it; here the application must.
When asked to demonstrate replication, GridFS or transactions, say what
they require — three mongod processes, and a replica set for
transactions. Knowing why transactions need one (they depend on the oplog
and majority commit) is worth more than a script you cannot run.
The same experiments, one page each, so a program can be reached by what it does rather than by its number.