Skip to the content
On this page
  1. The sample data
  2. Setting up for real
  3. Experiment 1 — Installation, Mongo Shell and Compass
  4. Experiment 2 — Databases, collections, inserting documents
  5. Experiment 3 — find() and comparison operators
  6. Experiment 4 — Logical operators
  7. Experiment 5 — Updating with $set, $unset, $inc, $rename
  8. Experiment 6 — Deleting
  9. Experiment 7 — Projection
  10. Experiment 8 — Sorting, limiting, skipping
  11. Experiment 9 — An embedded data model
  12. Experiment 10 — A normalized model with references
  13. Experiment 11 — One-to-one, one-to-many, many-to-many
  14. Experiment 12 — Schema validation with JSON Schema
  15. Experiment 13 — Single-field and compound indexes
  16. Experiment 14 — Text search and multikey indexes
  17. Experiment 15 — $match, $group, $project, $sort
  18. Experiment 16 — $lookup, $unwind, $bucket
  19. Experiment 17 — Replication with a replica set
  20. Experiment 18 — GridFS
  21. Experiment 19 — Transactions
  22. Experiment 20 — Case study: a mini-application
  23. Lab examination
  24. Each program, on its own page

20 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.

Code lives in labs/course-10-mongodb/.

NOTE

Both halves run. Every experiment has a mongosh script, NN_name.js — what the lab examiner will ask you to demonstrate — and it is run on MongoDB 8.3.7, typed into mongosh 2.12.0 line by line as you would at the prompt, on a fresh server each time. Under 5. Execution and Results is the session: each statement after its prompt, then what the shell printed. Sixteen also have a Python half, NN_name.py, the same query logic through mongomock, asserted by tools/data-science/run_mongo_labs.py.

Experiment 17 runs on a real three-member replica set (three mongod processes, as its section 0 describes, and a fourth as an arbiter), Experiment 18 with the real mongofiles, and Experiment 19 on a replica set, which transactions need.

Until October 2026 mongod could not be installed where these labs are checked, and every script said NOT EXECUTED. MongoDB's own download hosts are still blocked; conda-forge's builds of the same server can be reached, and tools/data-science/setup_mongodb.sh installs them. Running the scripts found eleven of them doing something other than what their comments said — a placeholder that is a syntax error, a variable never set, lines that start with a dot, statements on data that was not there, MongoDB 8 refusing two old forms — each corrected in its file, with a note, and listed under its experiment below.

tools/data-science/setup_mongodb.sh               # mongod and mongofiles, in /tmp/mongodb
npm --prefix tools/data-science install           # mongosh
python3 tools/data-science/run_mongo_labs.py      # both halves of every experiment

Two things in the output change from run to run: the ObjectIds, dates and UUIDs a server makes new each time, and everything about a replica set's election. capture_lab_outputs.py --check compares every other character of a session exactly; for Experiment 17 it reruns the replica set and checks what the experiment shows, as the driver's assertions.

The sample data

Every experiment from 3 onwards starts from the same five students, with the courses and enrolments Experiment 16 joins. 00_sample_data.js loads them: each script that needs them runs load("00_sample_data.js") straight after use collegeDB, so it can be run on its own, as often as you like. Start mongosh in the labs folder for load() to find it. The data is exactly that in fixtures.py, which the Python halves load, and run_mongo_labs.py checks that the two agree.

// The sample data every experiment from 3 onwards starts from: the five students
// of Experiment 2, with the courses and enrollments that Experiment 16 joins.
// It is the data in fixtures.py, which the Python halves load, exactly;
// tools/data-science/run_mongo_labs.py checks that the two agree.
//
// Each script loads it in its second line, with load("00_sample_data.js"), so
// it can be run on its own, as often as you like: start mongosh in this folder.
// It drops and refills only these three collections.

// Step 1: Switch to collegeDB
db = db.getSiblingDB("collegeDB")

// Step 2: Refill the students
db.students.drop()
db.students.insertMany([
  {"_id": 21, "name": "Asha", "dept": "DS", "marks": {"maths": 88, "stats": 91}, "subjects": ["DS", "Stats", "Python"], "age": 20, "active": true},
  {"_id": 22, "name": "Ravi", "dept": "DS", "marks": {"maths": 65, "stats": 58}, "subjects": ["DS", "Python"], "age": 21, "active": true},
  {"_id": 23, "name": "Meena", "dept": "Stats", "marks": {"maths": 94, "stats": 89}, "subjects": ["Stats", "R"], "age": 20, "active": true},
  {"_id": 24, "name": "Kiran", "dept": "DS", "marks": {"maths": 71, "stats": 66}, "subjects": ["DS"], "age": 22, "active": false},
  {"_id": 25, "name": "Bhanu", "dept": "Stats", "marks": {"maths": 52, "stats": 47}, "subjects": ["Stats"], "age": 21, "active": true}
])

// Step 3: Refill the courses
db.courses.drop()
db.courses.insertMany([
  {"_id": "DSC301", "title": "Data Science with R", "credits": 4, "instructor": "Dr. Rao"},
  {"_id": "STA302", "title": "Statistical Foundations", "credits": 3, "instructor": "Dr. Devi"},
  {"_id": "WEB303", "title": "Web Technologies", "credits": 3, "instructor": "Dr. Kumar"}
])

// Step 4: Refill the enrollments
db.enrollments.drop()
db.enrollments.insertMany([
  {"student_id": 21, "course_id": "DSC301", "grade": "A"},
  {"student_id": 21, "course_id": "STA302", "grade": "B"},
  {"student_id": 22, "course_id": "DSC301", "grade": "C"},
  {"student_id": 23, "course_id": "STA302", "grade": "A"},
  {"student_id": 24, "course_id": "WEB303", "grade": "B"}
])

// Step 5: Say what was loaded
print("sample data loaded: " + db.students.countDocuments() + " students, " +
      db.courses.countDocuments() + " courses, " + db.enrollments.countDocuments() + " enrollments")

Setting up for real

For the lab exam you need a real server. Three routes:

Route Command
MongoDB Atlas Free tier, no install — and it gives you a real replica set, which a local install does not
Docker docker run -d -p 27017:27017 --name mongo mongo
Local package apt install mongodb-org, or the platform installer
mongosh                                    # localhost:27017
mongosh "mongodb+srv://user:pass@cluster.mongodb.net/collegeDB"

Use Atlas or Docker. A local install commits you to managing a service, and Atlas is the only one of the three that gives you a replica set — which experiments 17 and 19 both require.


Experiment 1 — Installation, Mongo Shell and Compass

1. Question

Install MongoDB, and use the Mongo Shell and Compass.

2. Aim

Prove a server is running, find your way round mongosh, and know what Compass adds.

3. Steps

  1. Connect.
  2. Prove the install worked.
  3. Use the shell as a JavaScript REPL.
  4. Run the administrative commands.
  5. Clean up.

THE POINT

MongoDB Compass is the official GUI: browse collections, build queries without typing them, and read explain() output as a diagram rather than JSON. Worth installing for the explain visualiser alone.

Know for the viva: the default port is 27017; mongosh is a full JavaScript REPL, so loops and variables work in it; and show dbs will not list a database until something has been written to it.

4. Programme

// Experiment 1 -- Installing MongoDB, the Mongo shell and Compass.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. There is no .py half: these are server commands, with no query logic
// to run. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)

// Step 1: Connect
// From a terminal, NOT from inside mongosh:
//
//   mongosh                                          // localhost:27017
//   mongosh "mongodb://localhost:27017/collegeDB"    // straight into a db
//   mongosh "mongodb+srv://user:pass@cluster.mongodb.net/collegeDB"   // Atlas
//
// 27017 is the default port. Remember it -- it is asked in the viva.

// Step 2: Prove the install worked
db.version()                       // e.g. "7.0.14"
db.serverStatus().host             // hostname:port this shell is attached to
Math.floor(db.serverStatus().uptime / 60)   // whole minutes since mongod started
// [Changed: this was db.serverStatus().uptime, in seconds, which on a server
// started a moment ago reads 1 or 2 depending on the moment. In minutes it is 0.]
db.hostInfo().os                   // the OS the server is running on
db.hostInfo().system.numCores      // the cores it can see
// [Changed: this was db.hostInfo(), which prints some 350 lines, most of them
// disk counters that change by the second. These are the two lines that matter.]

show dbs                           // the databases that have been WRITTEN to
show collections                   // collections in the CURRENT database
db                                 // which database am I in?
db.getMongo()                      // the connection string

// Step 3: Use the shell as a JavaScript REPL
use collegeDB
load("00_sample_data.js")          // the five students, to have something to count
// This is the fact students most often miss, and it is worth demonstrating.
const depts = ["DS", "Stats", "CS"]
for (const d of depts) {
  print(`${d}: ${db.students.countDocuments({ dept: d })}`)
}

// Variables persist across statements; functions can be defined and reused.
function topper(dept) {
  return db.students.find({ dept }).sort({ "marks.maths": -1 }).limit(1).toArray()[0]
}
topper("DS")

// Load a script file from disk -- how you would run the rest of these labs:
//   load("02_create_insert.js")

// Step 4: Run the administrative commands
db.adminCommand({ listDatabases: 1 })
const s = db.stats()               // size, collection count, index count
({ collections: s.collections, objects: s.objects, indexes: s.indexes, dataSize: s.dataSize })
const c = db.students.stats()      // per-collection: documents, size, indexes
({ count: c.count, size: c.size, avgObjSize: c.avgObjSize, nindexes: c.nindexes })
// [Changed: these were db.stats() and db.students.stats() in full, which print
// the disk's free space and some 400 lines of storage-engine counters, both of
// which change from one run to the next. These are the figures the comments
// are about.]
db.getCollectionNames().sort()   // sorted: the server lists them in no fixed order

// Step 5: Clean up
use collegeDB                      // switches even if collegeDB does not exist
db.dropDatabase()                  // no confirmation, no undo

// --- MongoDB Compass ---------------------------------------------------------
// The official GUI (a separate download from the server).
//
//   * browse collections and documents without writing find()
//   * the Schema tab INFERS a schema from a sample -- the fastest way to see
//     what shape the documents in an inherited collection actually are
//   * the Explain Plan tab draws explain() output as a diagram instead of JSON,
//     which is worth the install on its own
//   * the Aggregations tab builds a pipeline stage by stage, showing the
//     intermediate documents after EACH stage -- exactly what you need when a
//     pipeline returns nothing and you cannot see which stage emptied it
//
// --- Know for the viva -------------------------------------------------------
//   * default port 27017
//   * mongosh is a full JavaScript REPL -- loops, variables, functions
//   * `show dbs` does NOT list a database until something has been written to
//     it: `use newdb` alone creates nothing
//   * the data directory defaults to /var/lib/mongodb (Linux); mongod refuses
//     to start if it does not exist or is not writable, which is the single
//     commonest install failure

5. Execution and Results

OUTPUT

test> db.version()                       // e.g. "7.0.14"
8.3.7
test> db.serverStatus().host             // hostname:port this shell is attached to
vm
test> Math.floor(db.serverStatus().uptime / 60)   // whole minutes since mongod started
0
test> db.hostInfo().os                   // the OS the server is running on
{ type: 'Linux', name: 'Ubuntu', version: '24.04' }
test> db.hostInfo().system.numCores      // the cores it can see
4
test> show dbs                           // the databases that have been WRITTEN to
admin    8.00 KiB
config  12.00 KiB
local    8.00 KiB
test> show collections                   // collections in the CURRENT database
test> db                                 // which database am I in?
test
test> db.getMongo()                      // the connection string
mongodb://127.0.0.1:27017/?directConnection=true&serverSelectionTimeoutMS=2000&appName=mongosh+2.12.0
test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")          // the five students, to have something to count
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> const depts = ["DS", "Stats", "CS"]
collegeDB> for (const d of depts) {
...   print(`${d}: ${db.students.countDocuments({ dept: d })}`)
... }
DS: 3
Stats: 2
CS: 0
collegeDB> function topper(dept) {
...   return db.students.find({ dept }).sort({ "marks.maths": -1 }).limit(1).toArray()[0]
... }
[Function: topper]
collegeDB> topper("DS")
{
  _id: 21,
  name: 'Asha',
  dept: 'DS',
  marks: { maths: 88, stats: 91 },
  subjects: [ 'DS', 'Stats', 'Python' ],
  age: 20,
  active: true
}
collegeDB> db.adminCommand({ listDatabases: 1 })
{
  databases: [
    { name: 'admin', sizeOnDisk: Long('8192'), empty: false },
    { name: 'collegeDB', sizeOnDisk: Long('24576'), empty: false },
    { name: 'config', sizeOnDisk: Long('12288'), empty: false },
    { name: 'local', sizeOnDisk: Long('8192'), empty: false }
  ],
  totalSize: Long('53248'),
  totalSizeMb: Long('0'),
  ok: 1
}
collegeDB> const s = db.stats()               // size, collection count, index count
collegeDB> ({ collections: s.collections, objects: s.objects, indexes: s.indexes, dataSize: s.dataSize })
{
  collections: Long('3'),
  objects: Long('13'),
  indexes: Long('3'),
  dataSize: 1296
}
collegeDB> const c = db.students.stats()      // per-collection: documents, size, indexes
collegeDB> ({ count: c.count, size: c.size, avgObjSize: c.avgObjSize, nindexes: c.nindexes })
{ count: 5, size: 660, avgObjSize: 132, nindexes: 1 }
collegeDB> db.getCollectionNames().sort()   // sorted: the server lists them in no fixed order
[ 'courses', 'enrollments', 'students' ]
collegeDB> use collegeDB                      // switches even if collegeDB does not exist
already on db collegeDB
collegeDB> db.dropDatabase()                  // no confirmation, no undo
{ ok: 1, dropped: 'collegeDB' }

db.version() reports the server, 8.3.7. Changed: db.hostInfo(), db.serverStatus().uptime, db.stats() and db.students.stats() printed hundreds of lines of machine and storage counters, which differ by the second; the script now asks each for the figures its comment is about. Experiment 1 has no Python half: there is no query logic in it.

RESULT

The server is MongoDB 8.3.7 on port 27017; mongosh runs JavaScript, so the loop and the function count and find as they should.

Experiment 2 — Databases, collections, inserting documents

1. Question

Create a database and a collection, and insert documents into it.

2. Aim

Insert one and many documents, and see what ordered and unordered inserts do on an error.

3. Steps

In mongosh, 02_create_insert.js:

  1. Switch to a database, which creates nothing yet.
  2. Create a collection.
  3. Insert one document.
  4. Insert many.
  5. See ordered stop at an error, and unordered carry on.
  6. Let MongoDB generate an ObjectId.
  7. Clean up.

In Python, through mongomock, 02_create_insert.py:

  1. See the database appear on the first write.
  2. Insert one and many.
  3. See ordered stop and unordered carry on.
  4. Look at a generated ObjectId.
  5. Manage the collections.

THE POINT

The two behaviours that are examinable, shown in the session and asserted by the Python half:

4. Programme

In mongosh, 02_create_insert.js:

// Experiment 2 -- Creating and using databases, creating collections,
// inserting documents.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 02_create_insert.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)

// Step 1: Switch to a database, which creates nothing yet
use collegeDB          // switches, but creates NOTHING yet
show dbs               // collegeDB is ABSENT until the first write

// Step 2: Create a collection
db.createCollection("students")        // only needed for OPTIONS
show collections

// Step 3: Insert one document
db.students.insertOne({
  _id: 21, name: "Asha", dept: "DS",
  marks: { maths: 88, stats: 91 },
  subjects: ["DS", "Stats", "Python"],
  age: 20, active: true
})
// -> { acknowledged: true, insertedId: 21 }

// Step 4: Insert many
db.students.insertMany([
  { _id: 22, name: "Ravi",  dept: "DS",    marks: { maths: 65, stats: 58 },
    subjects: ["DS", "Python"], age: 21, active: true },
  { _id: 23, name: "Meena", dept: "Stats", marks: { maths: 94, stats: 89 },
    subjects: ["Stats", "R"],   age: 20, active: true },
  { _id: 24, name: "Kiran", dept: "DS",    marks: { maths: 71, stats: 66 },
    subjects: ["DS"],           age: 22, active: false },
  { _id: 25, name: "Bhanu", dept: "Stats", marks: { maths: 52, stats: 47 },
    subjects: ["Stats"],        age: 21, active: true }
])

show dbs                          // NOW collegeDB appears
db.students.countDocuments()      // 5

// Step 5: See ordered stop at an error, and unordered carry on
db.students.insertMany([
  { _id: 30, name: "X" },
  { _id: 21, name: "DUPLICATE" },   // _id 21 exists -> error
  { _id: 31, name: "Y" }
])
// ordered (default): 30 inserted, 21 fails, 31 NEVER ATTEMPTED

db.students.insertMany([
  { _id: 40, name: "P" },
  { _id: 21, name: "DUPLICATE" },
  { _id: 41, name: "Q" }
], { ordered: false })
// unordered: 40 AND 41 inserted; only 21 fails

// Step 6: Let MongoDB generate an ObjectId
db.students.insertOne({ name: "Devi", dept: "Stats" })
db.students.findOne({ name: "Devi" })._id.getTimestamp()   // its creation time

// Step 7: Clean up
db.students.drop()
db.dropDatabase()

In Python, through mongomock, 02_create_insert.py:

"""Experiment 2 — Databases, collections, inserting documents.

Runs the same logic as 02_create_insert.js through mongomock and asserts it.
"""
import mongomock
from pymongo.errors import BulkWriteError, DuplicateKeyError
from fixtures import STUDENTS


def lazy_creation():
    """A database and a collection spring into existence on the first write."""
    client = mongomock.MongoClient()
    assert "collegeDB" not in client.list_database_names(), \
        "referencing a database does not create it"

    db = client.collegeDB
    assert "collegeDB" not in client.list_database_names(), \
        "nor does referencing it through the client"

    db.students.insert_one({"_id": 1, "name": "X"})
    assert "collegeDB" in client.list_database_names(), "NOW it exists"
    assert "students" in db.list_collection_names()

    print("  lazy creation: a database appears only after the first write")


def insert_one_and_many():
    db = mongomock.MongoClient().collegeDB

    r = db.students.insert_one(dict(STUDENTS[0]))
    assert r.inserted_id == 21
    assert db.students.count_documents({}) == 1

    r = db.students.insert_many([dict(d) for d in STUDENTS[1:]])
    assert r.inserted_ids == [22, 23, 24, 25]
    assert db.students.count_documents({}) == 5

    print(f"  insertOne -> insertedId 21; insertMany -> {r.inserted_ids}")


def ordered_stops_unordered_continues():
    """The examinable behaviour: ordered:true (the default) stops at the
    first error, so later documents are NEVER ATTEMPTED."""
    db = mongomock.MongoClient().collegeDB
    db.students.insert_many([dict(d) for d in STUDENTS])

    batch = [{"_id": 30, "name": "X"},
             {"_id": 21, "name": "DUPLICATE"},     # already exists
             {"_id": 31, "name": "Y"}]

    try:
        db.students.insert_many([dict(d) for d in batch])      # ordered=True
        raise AssertionError("expected a BulkWriteError")
    except BulkWriteError:
        pass

    assert db.students.count_documents({"_id": 30}) == 1, "inserted BEFORE the error"
    assert db.students.count_documents({"_id": 31}) == 0, \
        "NEVER ATTEMPTED -- ordered stops at the first failure"

    db2 = mongomock.MongoClient().collegeDB
    db2.students.insert_many([dict(d) for d in STUDENTS])
    try:
        db2.students.insert_many([dict(d) for d in batch], ordered=False)
        raise AssertionError("expected a BulkWriteError")
    except BulkWriteError:
        pass

    assert db2.students.count_documents({"_id": 30}) == 1
    assert db2.students.count_documents({"_id": 31}) == 1, \
        "unordered CONTINUES past the error"

    print("  ordered=True: 30 in, 21 fails, 31 never attempted")
    print("  ordered=False: 30 AND 31 in, only 21 fails")


def generated_object_id():
    db = mongomock.MongoClient().collegeDB
    r = db.students.insert_one({"name": "Devi", "dept": "Stats"})

    from bson import ObjectId
    assert isinstance(r.inserted_id, ObjectId)
    assert len(r.inserted_id.binary) == 12, "12 bytes"
    assert r.inserted_id.generation_time is not None, "it embeds its creation time"

    print(f"  omitting _id generates a 12-byte ObjectId carrying a timestamp")


def collection_management():
    db = mongomock.MongoClient().collegeDB
    db.students.insert_many([dict(d) for d in STUDENTS])

    assert db.students.count_documents({}) == 5
    assert db.students.count_documents({"dept": "DS"}) == 3
    assert sorted(db.students.distinct("dept")) == ["DS", "Stats"]

    db.students.drop()
    assert "students" not in db.list_collection_names()

    print("  countDocuments, distinct, drop -- all as documented")


def main():
    print("Experiment 2 -- Databases, collections, inserting")
    # Step 1: See the database appear on the first write
    lazy_creation()
    # Step 2: Insert one and many
    insert_one_and_many()
    # Step 3: See ordered stop and unordered carry on
    ordered_stops_unordered_continues()
    # Step 4: Look at a generated ObjectId
    generated_object_id()
    # Step 5: Manage the collections
    collection_management()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 02_create_insert.js:

OUTPUT

test> use collegeDB          // switches, but creates NOTHING yet
switched to db collegeDB
collegeDB> show dbs               // collegeDB is ABSENT until the first write
admin    8.00 KiB
config  12.00 KiB
local    8.00 KiB
collegeDB> db.createCollection("students")        // only needed for OPTIONS
{ ok: 1 }
collegeDB> show collections
students
collegeDB> db.students.insertOne({
...   _id: 21, name: "Asha", dept: "DS",
...   marks: { maths: 88, stats: 91 },
...   subjects: ["DS", "Stats", "Python"],
...   age: 20, active: true
... })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.students.insertMany([
...   { _id: 22, name: "Ravi",  dept: "DS",    marks: { maths: 65, stats: 58 },
...     subjects: ["DS", "Python"], age: 21, active: true },
...   { _id: 23, name: "Meena", dept: "Stats", marks: { maths: 94, stats: 89 },
...     subjects: ["Stats", "R"],   age: 20, active: true },
...   { _id: 24, name: "Kiran", dept: "DS",    marks: { maths: 71, stats: 66 },
...     subjects: ["DS"],           age: 22, active: false },
...   { _id: 25, name: "Bhanu", dept: "Stats", marks: { maths: 52, stats: 47 },
...     subjects: ["Stats"],        age: 21, active: true }
... ])
{
  acknowledged: true,
  insertedIds: { '0': 22, '1': 23, '2': 24, '3': 25 }
}
collegeDB> show dbs                          // NOW collegeDB appears
admin       8.00 KiB
collegeDB   8.00 KiB
config     12.00 KiB
local       8.00 KiB
collegeDB> db.students.countDocuments()      // 5
5
collegeDB> db.students.insertMany([
...   { _id: 30, name: "X" },
...   { _id: 21, name: "DUPLICATE" },   // _id 21 exists -> error
...   { _id: 31, name: "Y" }
... ])
Uncaught:

MongoBulkWriteError: E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }
Result: BulkWriteResult {
  insertedCount: 1,
  matchedCount: 0,
  modifiedCount: 0,
  deletedCount: 0,
  upsertedCount: 0,
  upsertedIds: {},
  insertedIds: { '0': 30 }
}
Write Errors: [
  WriteError {
    err: {
      index: 1,
      code: 11000,
      errmsg: 'E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }',
      errInfo: undefined,
      op: { _id: 21, name: 'DUPLICATE' }
    }
  }
]
collegeDB> db.students.insertMany([
...   { _id: 40, name: "P" },
...   { _id: 21, name: "DUPLICATE" },
...   { _id: 41, name: "Q" }
... ], { ordered: false })
Uncaught:

MongoBulkWriteError: E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }
Result: BulkWriteResult {
  insertedCount: 2,
  matchedCount: 0,
  modifiedCount: 0,
  deletedCount: 0,
  upsertedCount: 0,
  upsertedIds: {},
  insertedIds: { '0': 40, '2': 41 }
}
Write Errors: [
  WriteError {
    err: {
      index: 1,
      code: 11000,
      errmsg: 'E11000 duplicate key error collection: collegeDB.students index: _id_ dup key: { _id: 21 }',
      errInfo: undefined,
      op: { _id: 21, name: 'DUPLICATE' }
    }
  }
]
collegeDB> db.students.insertOne({ name: "Devi", dept: "Stats" })
{
  acknowledged: true,
  insertedId: ObjectId('6ac214abafbd22ad0b5b7bae')
}
collegeDB> db.students.findOne({ name: "Devi" })._id.getTimestamp()   // its creation time
ISODate('2026-10-04T08:56:11.000Z')
collegeDB> db.students.drop()
true
collegeDB> db.dropDatabase()
{ ok: 1, dropped: 'collegeDB' }

In Python, through mongomock, 02_create_insert.py:

OUTPUT

Experiment 2 -- Databases, collections, inserting
  lazy creation: a database appears only after the first write
  insertOne -> insertedId 21; insertMany -> [22, 23, 24, 25]
  ordered=True: 30 in, 21 fails, 31 never attempted
  ordered=False: 30 AND 31 in, only 21 fails
  omitting _id generates a 12-byte ObjectId carrying a timestamp
  countDocuments, distinct, drop -- all as documented

The ordered insert's error reports insertedCount: 1 — X went in, the duplicate stopped it, Y was never tried; the unordered one inserted both P and Q. The ObjectId and the time it carries are the server's, and differ every run.

RESULT

collegeDB appears in show dbs only after the first write; the ordered insert stops at the duplicate _id, and the unordered one carries on past it.

Experiment 3 — find() and comparison operators

1. Question

Query documents with find(), filtering with the comparison operators.

2. Aim

Filter with every comparison operator, and avoid the two traps.

3. Steps

In mongosh, 03_find_compare.js:

  1. Load the sample data.
  2. Find by equality, and findOne.
  3. Use the comparison operators.
  4. Query a sub-document with dot notation.
  5. See $ne match a missing field.
  6. Count, and find the distinct values.

In Python, through mongomock, 03_find_compare.py:

  1. Find by equality, and findOne.
  2. Use the comparison operators.
  3. Write a range as one object.
  4. Query a sub-document with dot notation.
  5. See $ne match a missing field.
  6. Count, and find the distinct values.

THE POINT

Every comparison operator, dot notation into a sub-document, and the two traps — a range must be one object, and $ne also matches documents where the field is missing.

4. Programme

In mongosh, 03_find_compare.js:

// Experiment 3 -- Basic queries using find(), filtering with comparison
// operators.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 03_find_compare.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: Find by equality, and findOne
db.students.find()                                  // everything
db.students.find({ dept: "DS" })                    // equality
db.students.find({ dept: "DS", age: 20 })           // implicit AND
db.students.findOne({ _id: 21 })                    // ONE document, or null

// Step 3: Use the comparison operators
db.students.find({ age: { $gt:  20 } })             // Ravi, Kiran, Bhanu
db.students.find({ age: { $gte: 21 } })
db.students.find({ age: { $lt:  21 } })             // Asha, Meena
db.students.find({ age: { $lte: 20 } })
db.students.find({ age: { $ne:  20 } })
db.students.find({ dept: { $in:  ["DS", "CS"] } })
db.students.find({ dept: { $nin: ["Stats"] } })

// A RANGE goes in ONE object. Written as two keys it is a JavaScript object
// with a duplicate key -- the first is SILENTLY DISCARDED.
db.students.find({ age: { $gte: 20, $lte: 21 } })   // correct
db.students.find({ age: { $gte: 20 }, age: { $lte: 21 } })   // WRONG, silently

// Step 4: Query a sub-document with dot notation
db.students.find({ "marks.maths": { $gte: 90 } })   // Meena
db.students.find({ "marks.maths": { $gt: 60, $lt: 90 } })

// Step 5: See $ne match a missing field
db.students.insertOne({ _id: 26, name: "NoDept" })
db.students.find({ dept: { $ne: "DS" } })           // Stats students AND _id 26
db.students.find({ dept: { $ne: "DS", $exists: true } })   // only real depts

// Step 6: Count, and find the distinct values
db.students.countDocuments({ dept: "DS" })
db.students.distinct("dept")

In Python, through mongomock, 03_find_compare.py:

"""Experiment 3 — find() and the comparison operators."""
from fixtures import fresh_db, names


def equality_and_findone():
    db = fresh_db()
    assert names(db.students.find({"dept": "DS"})) == ["Asha", "Kiran", "Ravi"]
    assert names(db.students.find({"dept": "DS", "age": 20})) == ["Asha"]

    one = db.students.find_one({"_id": 21})
    assert one["name"] == "Asha"
    assert db.students.find_one({"_id": 999}) is None, "findOne returns null"

    print("  equality, implicit AND, findOne -> a document or None")


def comparison_operators():
    db = fresh_db()
    cases = {
        "$gt 20":  ({"age": {"$gt": 20}},  ["Bhanu", "Kiran", "Ravi"]),
        "$gte 21": ({"age": {"$gte": 21}}, ["Bhanu", "Kiran", "Ravi"]),
        "$lt 21":  ({"age": {"$lt": 21}},  ["Asha", "Meena"]),
        "$lte 20": ({"age": {"$lte": 20}}, ["Asha", "Meena"]),
        "$ne 20":  ({"age": {"$ne": 20}},  ["Bhanu", "Kiran", "Ravi"]),
        "$in":     ({"dept": {"$in": ["DS", "CS"]}}, ["Asha", "Kiran", "Ravi"]),
        "$nin":    ({"dept": {"$nin": ["Stats"]}},   ["Asha", "Kiran", "Ravi"]),
    }
    for label, (q, want) in cases.items():
        got = names(db.students.find(q))
        assert got == want, f"{label}: {got} != {want}"

    print(f"  all seven comparison operators verified")


def a_range_is_one_object():
    """Written as two keys, the first is silently discarded."""
    db = fresh_db()

    correct = names(db.students.find({"age": {"$gte": 20, "$lte": 21}}))
    assert correct == ["Asha", "Bhanu", "Meena", "Ravi"], correct

    # In Python a dict literal with a duplicate key keeps the LAST -- the same
    # silent overwrite JavaScript performs. Only $lte survives.
    wrong = names(db.students.find({"age": {"$gte": 20}, "age": {"$lte": 21}}))
    assert wrong == ["Asha", "Bhanu", "Meena", "Ravi"] or "Kiran" not in wrong
    assert names(db.students.find({"age": {"$lte": 21}})) == wrong, \
        "only the LAST condition survived -- the $gte vanished"

    print("  a range must be ONE object; two keys silently drops one condition")


def dot_notation():
    db = fresh_db()
    assert names(db.students.find({"marks.maths": {"$gte": 90}})) == ["Meena"]
    assert names(db.students.find({"marks.maths": {"$gt": 60, "$lt": 90}})) == \
        ["Asha", "Kiran", "Ravi"]
    print("  dot notation reaches into sub-documents")


def ne_matches_missing_fields():
    """The trap: 'not equal to DS' is true of a field that does not exist."""
    db = fresh_db()
    db.students.insert_one({"_id": 26, "name": "NoDept"})

    loose = names(db.students.find({"dept": {"$ne": "DS"}}))
    assert "NoDept" in loose, "$ne ALSO matched the document with no dept"
    assert loose == ["Bhanu", "Meena", "NoDept"], loose

    tight = names(db.students.find({"dept": {"$ne": "DS", "$exists": True}}))
    assert tight == ["Bhanu", "Meena"], tight

    print("  $ne matched the document with NO dept field at all --")
    print("       combine with $exists: true when that matters")


def counting_and_distinct():
    db = fresh_db()
    assert db.students.count_documents({}) == 5
    assert db.students.count_documents({"dept": "DS"}) == 3
    assert sorted(db.students.distinct("dept")) == ["DS", "Stats"]
    assert sorted(db.students.distinct("subjects")) == ["DS", "Python", "R", "Stats"], \
        "distinct flattens ARRAY values"
    print("  distinct on an array field flattens it: DS, Python, R, Stats")


def main():
    print("Experiment 3 -- find() and comparison operators")
    # Step 1: Find by equality, and findOne
    equality_and_findone()
    # Step 2: Use the comparison operators
    comparison_operators()
    # Step 3: Write a range as one object
    a_range_is_one_object()
    # Step 4: Query a sub-document with dot notation
    dot_notation()
    # Step 5: See $ne match a missing field
    ne_matches_missing_fields()
    # Step 6: Count, and find the distinct values
    counting_and_distinct()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 03_find_compare.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find()                                  // everything
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ dept: "DS" })                    // equality
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find({ dept: "DS", age: 20 })           // implicit AND
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.findOne({ _id: 21 })                    // ONE document, or null
{
  _id: 21,
  name: 'Asha',
  dept: 'DS',
  marks: { maths: 88, stats: 91 },
  subjects: [ 'DS', 'Stats', 'Python' ],
  age: 20,
  active: true
}
collegeDB> db.students.find({ age: { $gt:  20 } })             // Ravi, Kiran, Bhanu
[
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ age: { $gte: 21 } })
[
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ age: { $lt:  21 } })             // Asha, Meena
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.find({ age: { $lte: 20 } })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.find({ age: { $ne:  20 } })
[
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ dept: { $in:  ["DS", "CS"] } })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find({ dept: { $nin: ["Stats"] } })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find({ age: { $gte: 20, $lte: 21 } })   // correct
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ age: { $gte: 20 }, age: { $lte: 21 } })   // WRONG, silently
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ "marks.maths": { $gte: 90 } })   // Meena
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.find({ "marks.maths": { $gt: 60, $lt: 90 } })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.insertOne({ _id: 26, name: "NoDept" })
{ acknowledged: true, insertedId: 26 }
collegeDB> db.students.find({ dept: { $ne: "DS" } })           // Stats students AND _id 26
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  },
  { _id: 26, name: 'NoDept' }
]
collegeDB> db.students.find({ dept: { $ne: "DS", $exists: true } })   // only real depts
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.countDocuments({ dept: "DS" })
3
collegeDB> db.students.distinct("dept")
[ 'DS', 'Stats' ]

In Python, through mongomock, 03_find_compare.py:

OUTPUT

Experiment 3 -- find() and comparison operators
  equality, implicit AND, findOne -> a document or None
  all seven comparison operators verified
  a range must be ONE object; two keys silently drops one condition
  dot notation reaches into sub-documents
  $ne matched the document with NO dept field at all --
       combine with $exists: true when that matters
  distinct on an array field flattens it: DS, Python, R, Stats

The range written as two keys returns more than the range: the first key is discarded, silently, and only $lte: 21 is applied.

RESULT

Every operator returns the students its comment names; $ne returns the document with no dept at all.

Experiment 4 — Logical operators

1. Question

Combine query conditions with the logical operators.

2. Aim

Combine conditions with $and, $or, $nor and $not.

3. Steps

In mongosh, 04_logical.js:

  1. Load the sample data.
  2. AND, implicit and explicit.
  3. OR.
  4. NOR.
  5. NOT, on an operator expression.
  6. Combine them.

In Python, through mongomock, 04_logical.py:

  1. AND, implicit and explicit.
  2. OR.
  3. NOR, by De Morgan.
  4. NOT, on an operator expression.
  5. Combine them.

THE POINT

$nor: [A, B] equals (NOT A) AND (NOT B) — De Morgan from Computer Fundamentals and Office Automation — and $not cannot take a plain value, only an operator expression. Both shown, and asserted by the Python half.

4. Programme

In mongosh, 04_logical.js:

// Experiment 4 -- Logical operators ($and, $or, $not, $nor) for complex
// queries.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 04_logical.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Implicit AND -- the usual form
// Step 2: AND, implicit and explicit
db.students.find({ dept: "DS", age: { $lt: 22 } })          // Asha, Ravi

// Explicit $and -- needed only for two conditions on the SAME field
db.students.find({ $and: [ { age: { $gte: 20 } }, { age: { $lte: 21 } } ] })

// Step 3: OR
db.students.find({ $or: [ { dept: "Stats" },
                          { "marks.maths": { $gt: 85 } } ] })

// $nor: NONE of the conditions. De Morgan: NOT(A OR B) = (NOT A) AND (NOT B)
// Step 4: NOR
db.students.find({ $nor: [ { dept: "DS" }, { age: 20 } ] })  // Bhanu

// $not inverts ONE OPERATOR EXPRESSION -- never a plain value
// Step 5: NOT, on an operator expression
db.students.find({ age: { $not: { $gt: 21 } } })            // NOT over 21
db.students.find({ age: { $not: 21 } })                     // ERROR

// Combining them
// Step 6: Combine them
db.students.find({
  dept: "DS",
  $or: [ { "marks.maths": { $gt: 80 } }, { "marks.stats": { $gt: 80 } } ]
})

// Nested
db.students.find({
  $and: [
    { $or: [ { dept: "DS" }, { dept: "Stats" } ] },
    { $or: [ { age: 20 }, { "marks.maths": { $gt: 70 } } ] }
  ]
})

In Python, through mongomock, 04_logical.py:

"""Experiment 4 — Logical operators."""
from fixtures import fresh_db, names


def implicit_and_explicit():
    db = fresh_db()

    implicit = names(db.students.find({"dept": "DS", "age": {"$lt": 22}}))
    explicit = names(db.students.find(
        {"$and": [{"dept": "DS"}, {"age": {"$lt": 22}}]}))
    assert implicit == explicit == ["Asha", "Ravi"], implicit

    # $and is REQUIRED for two conditions on the same field expressed
    # as separate clauses.
    both = names(db.students.find(
        {"$and": [{"age": {"$gte": 20}}, {"age": {"$lte": 21}}]}))
    assert both == ["Asha", "Bhanu", "Meena", "Ravi"], both

    print("  implicit AND == explicit $and; $and needed for same-field clauses")


def or_operator():
    db = fresh_db()
    got = names(db.students.find(
        {"$or": [{"dept": "Stats"}, {"marks.maths": {"$gt": 85}}]}))
    assert got == ["Asha", "Bhanu", "Meena"], got
    print(f"  $or (Stats OR maths>85) -> {got}")


def nor_is_de_morgan():
    db = fresh_db()

    nor = names(db.students.find({"$nor": [{"dept": "DS"}, {"age": 20}]}))
    assert nor == ["Bhanu"], nor

    # NOT(A OR B) == (NOT A) AND (NOT B) -- Course 1's De Morgan
    de_morgan = names(db.students.find(
        {"$and": [{"dept": {"$ne": "DS"}}, {"age": {"$ne": 20}}]}))
    assert nor == de_morgan, f"{nor} != {de_morgan}"

    print(f"  $nor [dept=DS, age=20] -> {nor}, identical to (NOT A) AND (NOT B)")


def not_needs_an_operator_expression():
    db = fresh_db()

    ok = names(db.students.find({"age": {"$not": {"$gt": 21}}}))
    assert ok == ["Asha", "Bhanu", "Meena", "Ravi"], ok

    # $not cannot take a plain value.
    try:
        list(db.students.find({"age": {"$not": 21}}))
        raise AssertionError("expected an error from $not with a plain value")
    except Exception as e:
        assert not isinstance(e, AssertionError), "should be a query error"

    print("  $not inverts an OPERATOR EXPRESSION; a plain value is an error")


def combining():
    db = fresh_db()

    got = names(db.students.find({
        "dept": "DS",
        "$or": [{"marks.maths": {"$gt": 80}}, {"marks.stats": {"$gt": 80}}]}))
    assert got == ["Asha"], got

    nested = names(db.students.find({
        "$and": [
            {"$or": [{"dept": "DS"}, {"dept": "Stats"}]},
            {"$or": [{"age": 20}, {"marks.maths": {"$gt": 70}}]},
        ]}))
    assert nested == ["Asha", "Kiran", "Meena"], nested

    print(f"  nested $and/$or -> {nested}")


def main():
    print("Experiment 4 -- Logical operators")
    # Step 1: AND, implicit and explicit
    implicit_and_explicit()
    # Step 2: OR
    or_operator()
    # Step 3: NOR, by De Morgan
    nor_is_de_morgan()
    # Step 4: NOT, on an operator expression
    not_needs_an_operator_expression()
    # Step 5: Combine them
    combining()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 04_logical.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find({ dept: "DS", age: { $lt: 22 } })          // Asha, Ravi
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ $and: [ { age: { $gte: 20 } }, { age: { $lte: 21 } } ] })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ $or: [ { dept: "Stats" },
...                           { "marks.maths": { $gt: 85 } } ] })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ $nor: [ { dept: "DS" }, { age: 20 } ] })  // Bhanu
[
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ age: { $not: { $gt: 21 } } })            // NOT over 21
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ age: { $not: 21 } })                     // ERROR
Uncaught
MongoServerError[BadValue]: $not argument must be a regex or an object
collegeDB> db.students.find({
...   dept: "DS",
...   $or: [ { "marks.maths": { $gt: 80 } }, { "marks.stats": { $gt: 80 } } ]
... })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.find({
...   $and: [
...     { $or: [ { dept: "DS" }, { dept: "Stats" } ] },
...     { $or: [ { age: 20 }, { "marks.maths": { $gt: 70 } } ] }
...   ]
... })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]

In Python, through mongomock, 04_logical.py:

OUTPUT

Experiment 4 -- Logical operators
  implicit AND == explicit $and; $and needed for same-field clauses
  $or (Stats OR maths>85) -> ['Asha', 'Bhanu', 'Meena']
  $nor [dept=DS, age=20] -> ['Bhanu'], identical to (NOT A) AND (NOT B)
  $not inverts an OPERATOR EXPRESSION; a plain value is an error
  nested $and/$or -> ['Asha', 'Kiran', 'Meena']

RESULT

$nor returns Bhanu alone, as (NOT DS) AND (NOT 20) does; $not with a plain value is an error.

Experiment 5 — Updating with $set, $unset, $inc, $rename

1. Question

Update documents with the update operators.

2. Aim

Change documents with $set, $unset, $inc and $rename, and see what replaceOne and upsert do.

3. Steps

In mongosh, 05_update.js:

  1. Load the sample data.
  2. $set, $inc, $unset and $rename.
  3. $mul, $max and $currentDate.
  4. Several operators at once.
  5. See updateOne change exactly one.
  6. See replaceOne drop every other field.
  7. Upsert.
  8. findOneAndUpdate.

In Python, through mongomock, 05_update.py:

  1. $set, $unset, $inc and $rename.
  2. Several operators at once.
  3. See updateOne change exactly one.
  4. See replaceOne drop every other field.
  5. Upsert.
  6. findOneAndUpdate.

THE POINT

replaceOne keeps only _id and discards every other field, while updateOne with $set preserves them. And updateOne changes exactly one document when three match — the commonest CRUD mistake, and silent.

4. Programme

In mongosh, 05_update.js:

// Experiment 5 -- Updating documents with $set, $unset, $inc, $rename.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 05_update.py, through
// mongomock. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: $set, $inc, $unset and $rename
db.students.updateOne({ _id: 21 }, { $set: { age: 21 } })
db.students.updateOne({ _id: 21 }, { $set: { "marks.python": 85 } })  // nested
db.students.updateMany({ dept: "DS" }, { $inc: { "marks.maths": 5 } })
db.students.updateOne({ _id: 21 }, { $inc: { age: -1 } })             // subtract
db.students.updateOne({ _id: 21 }, { $unset: { active: "" } })        // value ignored
db.students.updateMany({}, { $rename: { "dept": "department" } })
db.students.updateMany({}, { $rename: { "department": "dept" } })   // and back again
// [Corrected: the second $rename was missing. Without it every later query on
// dept matched nothing -- the updateOne below, commented ONE of three, and the
// updateMany, all three, both reported matchedCount: 0.]
// Step 3: $mul, $max and $currentDate
db.students.updateOne({ _id: 21 }, { $mul: { "marks.maths": 1.1 } })
db.students.updateOne({ _id: 21 }, { $max: { "marks.maths": 95 } })   // only if higher
db.students.updateOne({ _id: 21 }, { $currentDate: { updated: true } })

// Several operators in ONE update
// Step 4: Several operators at once
db.students.updateOne({ _id: 22 }, {
  $set:   { grade: "B" },
  $inc:   { age: 1 },
  $unset: { active: "" }
})

// Step 5: See updateOne change exactly one
db.students.updateOne({ dept: "DS" }, { $set: { flag: true } })   // ONE of three
db.students.updateMany({ dept: "DS" }, { $set: { flag: true } })  // all three

// Step 6: See replaceOne drop every other field
db.students.replaceOne({ _id: 21 }, { name: "Asha K" })
// the document is now { _id: 21, name: "Asha K" } -- everything else is GONE

// Step 7: Upsert
db.counters.updateOne(
  { _id: "visits" },
  { $inc: { count: 1 }, $setOnInsert: { created: new Date() } },
  { upsert: true }
)

// Step 8: findOneAndUpdate
db.students.findOneAndUpdate({ _id: 21 }, { $set: { age: 22 } },
                             { returnDocument: "after" })

In Python, through mongomock, 05_update.py:

"""Experiment 5 — Update operators."""
from fixtures import fresh_db, names


def set_unset_inc_rename():
    db = fresh_db()

    db.students.update_one({"_id": 21}, {"$set": {"age": 21}})
    assert db.students.find_one({"_id": 21})["age"] == 21

    db.students.update_one({"_id": 21}, {"$set": {"marks.python": 85}})
    assert db.students.find_one({"_id": 21})["marks"]["python"] == 85, \
        "$set creates a nested field"

    r = db.students.update_many({"dept": "DS"}, {"$inc": {"marks.maths": 5}})
    assert r.modified_count == 3
    assert db.students.find_one({"_id": 21})["marks"]["maths"] == 93

    db.students.update_one({"_id": 21}, {"$inc": {"age": -1}})
    assert db.students.find_one({"_id": 21})["age"] == 20, "a negative $inc subtracts"

    db.students.update_one({"_id": 21}, {"$unset": {"active": ""}})
    assert "active" not in db.students.find_one({"_id": 21}), \
        "$unset REMOVES the field; its value is ignored"

    db.students.update_many({}, {"$rename": {"dept": "department"}})
    doc = db.students.find_one({"_id": 22})
    assert "department" in doc and "dept" not in doc

    print("  $set (incl. nested), $inc (incl. negative), $unset, $rename")


def several_operators_at_once():
    db = fresh_db()
    db.students.update_one({"_id": 22}, {
        "$set":   {"grade": "B"},
        "$inc":   {"age": 1},
        "$unset": {"active": ""}})
    d = db.students.find_one({"_id": 22})
    assert d["grade"] == "B" and d["age"] == 22 and "active" not in d
    print("  several operators combine in one update document")


def update_one_changes_exactly_one():
    """The commonest CRUD mistake: silent, and it reports success."""
    db = fresh_db()
    assert db.students.count_documents({"dept": "DS"}) == 3

    r = db.students.update_one({"dept": "DS"}, {"$set": {"flag": True}})
    assert r.modified_count == 1, "ONE, even though three matched"
    assert db.students.count_documents({"flag": True}) == 1

    db2 = fresh_db()
    r2 = db2.students.update_many({"dept": "DS"}, {"$set": {"flag": True}})
    assert r2.modified_count == 3
    assert db2.students.count_documents({"flag": True}) == 3

    print("  updateOne -> modifiedCount 1 of 3 matches; updateMany -> 3")
    print("       the command SUCCEEDS either way -- nothing warns you")


def replace_one_destroys_everything():
    db = fresh_db()
    before = db.students.find_one({"_id": 21})
    assert set(before) >= {"name", "dept", "marks", "subjects", "age", "active"}

    db.students.replace_one({"_id": 21}, {"name": "Asha K"})
    after = db.students.find_one({"_id": 21})

    assert set(after) == {"_id", "name"}, f"only _id and name survive: {set(after)}"
    assert after["name"] == "Asha K"

    # updateOne with $set preserves the rest.
    db2 = fresh_db()
    db2.students.update_one({"_id": 21}, {"$set": {"name": "Asha K"}})
    kept = db2.students.find_one({"_id": 21})
    assert "marks" in kept and "subjects" in kept and kept["name"] == "Asha K"

    print("  replaceOne left only {_id, name}; updateOne+$set kept everything")


def upsert_is_atomic():
    db = fresh_db()

    # First call: inserts.
    r = db.counters.update_one({"_id": "visits"},
                               {"$inc": {"count": 1},
                                "$setOnInsert": {"created": "2026-08-26"}},
                               upsert=True)
    assert r.upserted_id == "visits"
    assert db.counters.find_one({"_id": "visits"})["count"] == 1

    # Later calls: update, and $setOnInsert does NOT re-apply.
    for _ in range(4):
        db.counters.update_one({"_id": "visits"},
                               {"$inc": {"count": 1},
                                "$setOnInsert": {"created": "LATER"}},
                               upsert=True)
    doc = db.counters.find_one({"_id": "visits"})
    assert doc["count"] == 5
    assert doc["created"] == "2026-08-26", "$setOnInsert applies ONLY on insert"

    print("  upsert: inserted then incremented to 5; $setOnInsert applied once")


def find_one_and_update():
    db = fresh_db()
    from pymongo import ReturnDocument

    after = db.students.find_one_and_update(
        {"_id": 21}, {"$set": {"age": 22}},
        return_document=ReturnDocument.AFTER)
    assert after["age"] == 22, "returnDocument: 'after' gives the NEW document"

    db2 = fresh_db()
    before = db2.students.find_one_and_update({"_id": 21}, {"$set": {"age": 22}})
    assert before["age"] == 20, "the default returns the document BEFORE the update"

    print("  findOneAndUpdate returns the OLD document by default, or the new one")


def main():
    print("Experiment 5 -- Updating documents")
    # Step 1: $set, $unset, $inc and $rename
    set_unset_inc_rename()
    # Step 2: Several operators at once
    several_operators_at_once()
    # Step 3: See updateOne change exactly one
    update_one_changes_exactly_one()
    # Step 4: See replaceOne drop every other field
    replace_one_destroys_everything()
    # Step 5: Upsert
    upsert_is_atomic()
    # Step 6: findOneAndUpdate
    find_one_and_update()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 05_update.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.updateOne({ _id: 21 }, { $set: { age: 21 } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $set: { "marks.python": 85 } })  // nested
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateMany({ dept: "DS" }, { $inc: { "marks.maths": 5 } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 3,
  modifiedCount: 3,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $inc: { age: -1 } })             // subtract
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $unset: { active: "" } })        // value ignored
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateMany({}, { $rename: { "dept": "department" } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 5,
  modifiedCount: 5,
  upsertedCount: 0
}
collegeDB> db.students.updateMany({}, { $rename: { "department": "dept" } })   // and back again
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 5,
  modifiedCount: 5,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $mul: { "marks.maths": 1.1 } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $max: { "marks.maths": 95 } })   // only if higher
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 0,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 }, { $currentDate: { updated: true } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 22 }, {
...   $set:   { grade: "B" },
...   $inc:   { age: 1 },
...   $unset: { active: "" }
... })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ dept: "DS" }, { $set: { flag: true } })   // ONE of three
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateMany({ dept: "DS" }, { $set: { flag: true } })  // all three
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 3,
  modifiedCount: 2,
  upsertedCount: 0
}
collegeDB> db.students.replaceOne({ _id: 21 }, { name: "Asha K" })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.counters.updateOne(
...   { _id: "visits" },
...   { $inc: { count: 1 }, $setOnInsert: { created: new Date() } },
...   { upsert: true }
... )
{
  acknowledged: true,
  insertedId: 'visits',
  matchedCount: 0,
  modifiedCount: 0,
  upsertedCount: 1
}
collegeDB> db.students.findOneAndUpdate({ _id: 21 }, { $set: { age: 22 } },
...                              { returnDocument: "after" })
{ _id: 21, name: 'Asha K', age: 22 }

In Python, through mongomock, 05_update.py:

OUTPUT

Experiment 5 -- Updating documents
  $set (incl. nested), $inc (incl. negative), $unset, $rename
  several operators combine in one update document
  updateOne -> modifiedCount 1 of 3 matches; updateMany -> 3
       the command SUCCEEDS either way -- nothing warns you
  replaceOne left only {_id, name}; updateOne+$set kept everything
  upsert: inserted then incremented to 5; $setOnInsert applied once
  findOneAndUpdate returns the OLD document by default, or the new one

Corrected: the $rename of dept to department was never undone, so every later query on dept matched nothing — the updateOne commented "ONE of three" and the updateMany "all three" both reported matchedCount: 0. A second $rename now puts it back. The Python half had passed, because each of its functions starts from fresh data.

RESULT

updateOne changed one of the three DS students and updateMany all three; replaceOne left Asha only a name; the upsert inserted the counter.

Experiment 6 — Deleting

1. Question

Delete documents, and a collection.

2. Aim

Delete one, many and all documents, and tell deleting them from dropping the collection.

3. Steps

In mongosh, 06_delete.js:

  1. Load the sample data.
  2. deleteOne and deleteMany.
  3. findOneAndDelete.
  4. deleteMany({}) against drop().

In Python, through mongomock, 06_delete.py:

  1. deleteOne and deleteMany.
  2. findOneAndDelete.
  3. Tell deleteMany({}) from drop().
  4. See an empty filter delete everything.

THE POINT

deleteMany({}) empties the collection but keeps it, its indexes and its validator; drop() removes all three.

There is no confirmation and no undo. In Database Management Systems, DELETE FROM t at least sat inside a transaction you could roll back.

4. Programme

In mongosh, 06_delete.js:

// Experiment 6 -- Deleting documents using deleteOne() and deleteMany().
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 06_delete.py, through
// mongomock. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: deleteOne and deleteMany
db.students.deleteOne({ _id: 25 })
db.students.deleteOne({ dept: "DS" })          // ONE of the three
db.students.deleteMany({ dept: "DS" })         // all of them
// Step 3: findOneAndDelete
db.students.findOneAndDelete({ _id: 23 })      // returns the deleted document
// [Corrected: this deleted _id 21, Asha -- already gone, as one of the three DS
// students deleted above -- so it returned null, not a document. Meena, 23, is
// the one student left.]

// deleteMany({}) removes EVERY document. No confirmation, no undo.
// Step 4: deleteMany({}) against drop()
load("00_sample_data.js")                      // five students again, to delete
db.students.deleteMany({})                     // the COLLECTION remains
db.getCollectionNames().sort()                 // students is still listed
db.students.drop()                             // the collection AND its indexes
db.getCollectionNames().sort()                 // and now it is not
// [Changed: the load and the two getCollectionNames() lines were added, so that
// deleteMany({}) has something to delete and the difference from drop() shows.
// The names are sorted because the server lists them in no fixed order.]

In Python, through mongomock, 06_delete.py:

"""Experiment 6 — Deleting documents."""
from fixtures import fresh_db, names, STUDENTS


def delete_one_and_many():
    db = fresh_db()
    assert db.students.count_documents({}) == 5

    r = db.students.delete_one({"_id": 25})
    assert r.deleted_count == 1 and db.students.count_documents({}) == 4

    r = db.students.delete_one({"dept": "DS"})
    assert r.deleted_count == 1, "ONE, even though three match"
    assert db.students.count_documents({"dept": "DS"}) == 2

    r = db.students.delete_many({"dept": "DS"})
    assert r.deleted_count == 2 and db.students.count_documents({"dept": "DS"}) == 0

    print("  deleteOne removes ONE of three matches; deleteMany removes all")


def find_one_and_delete():
    db = fresh_db()
    doc = db.students.find_one_and_delete({"_id": 21})
    assert doc["name"] == "Asha", "it RETURNS the deleted document"
    assert db.students.find_one({"_id": 21}) is None
    print("  findOneAndDelete returns the document it removed")


def delete_all_versus_drop():
    """deleteMany({}) empties the collection; drop() removes it entirely."""
    db = fresh_db()
    db.students.create_index("dept")
    before_indexes = len(list(db.students.list_indexes()))
    assert before_indexes >= 2, "_id plus the one we made"

    db.students.delete_many({})
    assert db.students.count_documents({}) == 0
    assert "students" in db.list_collection_names(), "the COLLECTION remains"
    assert len(list(db.students.list_indexes())) == before_indexes, \
        "and so do its INDEXES"

    db2 = fresh_db()
    db2.students.create_index("dept")
    db2.students.drop()
    assert "students" not in db2.list_collection_names(), "gone entirely"

    print(f"  deleteMany({{}}): 0 documents, collection and {before_indexes} indexes kept")
    print(f"  drop(): the collection, its documents and its indexes all removed")


def no_confirmation_no_undo():
    """A deliberate demonstration of how easy the mistake is."""
    db = fresh_db()
    intended = {"dept": "Physics"}          # matches nothing
    typo = {}                               # what a slip produces

    r_safe = db.students.delete_many(intended)
    assert r_safe.deleted_count == 0 and db.students.count_documents({}) == 5

    r_oops = db.students.delete_many(typo)
    assert r_oops.deleted_count == 5, "an empty filter matched EVERYTHING"
    assert db.students.count_documents({}) == 0

    print("  an empty filter deleted all 5 documents, reported success, and")
    print("       there is no transaction to roll back -- unlike Course 5")


def main():
    print("Experiment 6 -- Deleting documents")
    # Step 1: deleteOne and deleteMany
    delete_one_and_many()
    # Step 2: findOneAndDelete
    find_one_and_delete()
    # Step 3: Tell deleteMany({}) from drop()
    delete_all_versus_drop()
    # Step 4: See an empty filter delete everything
    no_confirmation_no_undo()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 06_delete.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.deleteOne({ _id: 25 })
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.students.deleteOne({ dept: "DS" })          // ONE of the three
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.students.deleteMany({ dept: "DS" })         // all of them
{ acknowledged: true, deletedCount: 2 }
collegeDB> db.students.findOneAndDelete({ _id: 23 })      // returns the deleted document
{
  _id: 23,
  name: 'Meena',
  dept: 'Stats',
  marks: { maths: 94, stats: 89 },
  subjects: [ 'Stats', 'R' ],
  age: 20,
  active: true
}
collegeDB> load("00_sample_data.js")                      // five students again, to delete
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.deleteMany({})                     // the COLLECTION remains
{ acknowledged: true, deletedCount: 5 }
collegeDB> db.getCollectionNames().sort()                 // students is still listed
[ 'courses', 'enrollments', 'students' ]
collegeDB> db.students.drop()                             // the collection AND its indexes
true
collegeDB> db.getCollectionNames().sort()                 // and now it is not
[ 'courses', 'enrollments' ]

In Python, through mongomock, 06_delete.py:

OUTPUT

Experiment 6 -- Deleting documents
  deleteOne removes ONE of three matches; deleteMany removes all
  findOneAndDelete returns the document it removed
  deleteMany({}): 0 documents, collection and 2 indexes kept
  drop(): the collection, its documents and its indexes all removed
  an empty filter deleted all 5 documents, reported success, and
       there is no transaction to roll back -- unlike Course 5

Corrected: findOneAndDelete was on _id 21, Asha, already deleted as one of the DS students, so it returned null; it is now on Meena, the one student left, and returns her. Changed: the sample data is loaded again before deleteMany({}), which otherwise had one document left to delete, and getCollectionNames() shows the collection before and after drop().

RESULT

deleteMany({}) deleted all five and left the collection listed; drop() removed it.

Experiment 7 — Projection

1. Question

Choose which fields a query returns, with a projection.

2. Aim

Include and exclude fields, nested ones and array elements.

3. Steps

In mongosh, 07_projection.js:

  1. Load the sample data.
  2. Include and exclude fields.
  3. Project a nested field.
  4. Project arrays.
  5. See that inclusion and exclusion cannot mix.

In Python, through mongomock, 07_projection.py:

  1. Include and exclude fields.
  2. See that the two cannot mix.
  3. Project nested fields.
  4. Project arrays.
  5. See projection reduce the work.

THE POINT

Mixing inclusion and exclusion raises an error, and _id is the one exception — it may be excluded alongside inclusions.

4. Programme

In mongosh, 07_projection.js:

// Experiment 7 -- Using projection to display selective fields.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 07_projection.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: Include and exclude fields
db.students.find({}, { name: 1, dept: 1 })            // these fields PLUS _id
db.students.find({}, { name: 1, _id: 0 })             // exclude _id
db.students.find({}, { marks: 0, subjects: 0 })       // everything EXCEPT these
// Step 3: Project a nested field
db.students.find({}, { "marks.maths": 1, _id: 0 })    // one nested field
// Step 4: Project arrays
db.students.find({}, { subjects: { $slice: 2 } })     // first 2 array elements
db.students.find({}, { subjects: { $slice: -1 } })    // the LAST element
db.students.find({ subjects: "DS" }, { "subjects.$": 1 })   // the MATCHING one

// You cannot MIX inclusion and exclusion...
// Step 5: See that inclusion and exclusion cannot mix
db.students.find({}, { name: 1, dept: 0 })            // ERROR
// ...except for _id, which is the one exception.
db.students.find({}, { name: 1, _id: 0 })             // fine

In Python, through mongomock, 07_projection.py:

"""Experiment 7 — Projection."""
from fixtures import fresh_db


def inclusion_and_exclusion():
    db = fresh_db()

    d = db.students.find_one({"_id": 21}, {"name": 1, "dept": 1})
    assert set(d) == {"_id", "name", "dept"}, "_id is included BY DEFAULT"

    d = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})
    assert set(d) == {"name"}

    d = db.students.find_one({"_id": 21}, {"marks": 0, "subjects": 0})
    assert "marks" not in d and "subjects" not in d
    assert {"name", "dept", "age", "active", "_id"} <= set(d)

    print("  inclusion adds _id automatically; exclusion keeps everything else")


def cannot_mix():
    db = fresh_db()
    try:
        list(db.students.find({}, {"name": 1, "dept": 0}))
        raise AssertionError("expected an error from mixing inclusion/exclusion")
    except Exception as e:
        assert not isinstance(e, AssertionError)

    # _id is the ONE exception.
    d = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})
    assert set(d) == {"name"}

    print("  mixing inclusion and exclusion raises; _id is the one exception")


def nested_projection():
    db = fresh_db()
    d = db.students.find_one({"_id": 21}, {"marks.maths": 1, "_id": 0})
    assert d == {"marks": {"maths": 88}}, d
    print("  'marks.maths': 1 keeps the sub-document with only that field")


def array_projection():
    db = fresh_db()

    d = db.students.find_one({"_id": 21}, {"subjects": {"$slice": 2}, "_id": 0,
                                           "name": 1})
    assert d["subjects"] == ["DS", "Stats"], "the first 2"

    d = db.students.find_one({"_id": 21}, {"subjects": {"$slice": -1}, "_id": 0,
                                           "name": 1})
    assert d["subjects"] == ["Python"], "a negative slice takes from the END"

    print("  $slice: 2 -> first two; $slice: -1 -> the last one")


def projection_reduces_work():
    """Not just cosmetic -- fewer bytes read and sent."""
    db = fresh_db()
    full = db.students.find_one({"_id": 21})
    thin = db.students.find_one({"_id": 21}, {"name": 1, "_id": 0})

    assert len(str(full)) > 3 * len(str(thin)), \
        "the projected document is a fraction of the size"

    print(f"  full document {len(str(full))} chars vs projected {len(str(thin))}")
    print(f"       and a projection fully covered by an index needs no document read")


def main():
    print("Experiment 7 -- Projection")
    # Step 1: Include and exclude fields
    inclusion_and_exclusion()
    # Step 2: See that the two cannot mix
    cannot_mix()
    # Step 3: Project nested fields
    nested_projection()
    # Step 4: Project arrays
    array_projection()
    # Step 5: See projection reduce the work
    projection_reduces_work()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 07_projection.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find({}, { name: 1, dept: 1 })            // these fields PLUS _id
[
  { _id: 21, name: 'Asha', dept: 'DS' },
  { _id: 22, name: 'Ravi', dept: 'DS' },
  { _id: 23, name: 'Meena', dept: 'Stats' },
  { _id: 24, name: 'Kiran', dept: 'DS' },
  { _id: 25, name: 'Bhanu', dept: 'Stats' }
]
collegeDB> db.students.find({}, { name: 1, _id: 0 })             // exclude _id
[
  { name: 'Asha' },
  { name: 'Ravi' },
  { name: 'Meena' },
  { name: 'Kiran' },
  { name: 'Bhanu' }
]
collegeDB> db.students.find({}, { marks: 0, subjects: 0 })       // everything EXCEPT these
[
  { _id: 21, name: 'Asha', dept: 'DS', age: 20, active: true },
  { _id: 22, name: 'Ravi', dept: 'DS', age: 21, active: true },
  { _id: 23, name: 'Meena', dept: 'Stats', age: 20, active: true },
  { _id: 24, name: 'Kiran', dept: 'DS', age: 22, active: false },
  { _id: 25, name: 'Bhanu', dept: 'Stats', age: 21, active: true }
]
collegeDB> db.students.find({}, { "marks.maths": 1, _id: 0 })    // one nested field
[
  { marks: { maths: 88 } },
  { marks: { maths: 65 } },
  { marks: { maths: 94 } },
  { marks: { maths: 71 } },
  { marks: { maths: 52 } }
]
collegeDB> db.students.find({}, { subjects: { $slice: 2 } })     // first 2 array elements
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({}, { subjects: { $slice: -1 } })    // the LAST element
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ subjects: "DS" }, { "subjects.$": 1 })   // the MATCHING one
[
  { _id: 21, subjects: [ 'DS' ] },
  { _id: 22, subjects: [ 'DS' ] },
  { _id: 24, subjects: [ 'DS' ] }
]
collegeDB> db.students.find({}, { name: 1, dept: 0 })            // ERROR
Uncaught
MongoServerError[Location31254]: Cannot do exclusion on field dept in inclusion projection
collegeDB> db.students.find({}, { name: 1, _id: 0 })             // fine
[
  { name: 'Asha' },
  { name: 'Ravi' },
  { name: 'Meena' },
  { name: 'Kiran' },
  { name: 'Bhanu' }
]

In Python, through mongomock, 07_projection.py:

OUTPUT

Experiment 7 -- Projection
  inclusion adds _id automatically; exclusion keeps everything else
  mixing inclusion and exclusion raises; _id is the one exception
  'marks.maths': 1 keeps the sub-document with only that field
  $slice: 2 -> first two; $slice: -1 -> the last one
  full document 144 chars vs projected 16
       and a projection fully covered by an index needs no document read

RESULT

Inclusion and exclusion cannot be mixed, except for _id; $slice and the positional $ cut arrays down.

Experiment 8 — Sorting, limiting, skipping

1. Question

Sort, limit and skip the results of a query.

2. Aim

Order results and page through them, the slow way and the fast way.

3. Steps

In mongosh, 08_sort_limit.js:

  1. Load the sample data.
  2. Sort.
  3. Limit and skip.
  4. See sort apply before limit, whatever the order written.
  5. Paginate by range, not by skip.
  6. Take the top document from the cursor.

In Python, through mongomock, 08_sort_limit.py:

  1. Sort.
  2. Limit and skip.
  3. See sort, skip and limit apply in a fixed order.
  4. Paginate by range, not by skip.
  5. Tell a cursor from a document.

THE POINT

The server applies sort, then skip, then limit, regardless of the chaining order — so .limit(3).sort(...) sorts everything and then takes three.

The script also demonstrates range pagination ({ _id: { $gt: last } }) as the fix for skip's linear cost.

4. Programme

In mongosh, 08_sort_limit.js:

// Experiment 8 -- Sorting documents, limiting output, skipping records.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 08_sort_limit.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: Sort
db.students.find().sort({ "marks.maths": -1 })              // -1 descending
db.students.find().sort({ dept: 1, "marks.maths": -1 })     // multi-key
// Step 3: Limit and skip
db.students.find().sort({ "marks.maths": -1 }).limit(3)     // top 3
db.students.find().skip(2).limit(2)                         // "page 2"

// The server ALWAYS applies sort, then skip, then limit -- whatever order you
// chain them in. So .limit(3).sort(...) sorts EVERYTHING and then takes three.
// Step 4: See sort apply before limit, whatever the order written
db.students.find().limit(3).sort({ "marks.maths": -1 })

// Step 5: Paginate by range, not by skip
// skip(100000) makes the server WALK AND DISCARD 100,000 documents.
// Range pagination uses the index to jump straight there:
db.students.find().sort({ _id: 1 }).limit(2)                       // page 1
const lastSeenId = 22                                              // the last _id on page 1
db.students.find({ _id: { $gt: lastSeenId } }).sort({ _id: 1 }).limit(2)   // page 2
// [Corrected: lastSeenId was never set, so this line failed with
// "ReferenceError: lastSeenId is not defined".]

// Step 6: Take the top document from the cursor
db.students.find().sort({ "marks.maths": -1 }).limit(1).next().name   // topper

In Python, through mongomock, 08_sort_limit.py:

"""Experiment 8 — Sorting, limiting and skipping."""
from fixtures import fresh_db


def sorting():
    db = fresh_db()

    desc = [d["name"] for d in db.students.find().sort("marks.maths", -1)]
    assert desc == ["Meena", "Asha", "Kiran", "Ravi", "Bhanu"], desc

    asc = [d["name"] for d in db.students.find().sort("marks.maths", 1)]
    assert asc == list(reversed(desc))

    multi = [(d["dept"], d["marks"]["maths"])
             for d in db.students.find().sort([("dept", 1), ("marks.maths", -1)])]
    assert multi == [("DS", 88), ("DS", 71), ("DS", 65),
                     ("Stats", 94), ("Stats", 52)], multi

    print(f"  sort by maths desc -> {desc}")
    print(f"  multi-key sort (dept asc, maths desc) groups then orders within")


def limit_and_skip():
    db = fresh_db()

    top3 = [d["name"] for d in db.students.find().sort("marks.maths", -1).limit(3)]
    assert top3 == ["Meena", "Asha", "Kiran"]

    page2 = [d["name"] for d in db.students.find().sort("_id", 1).skip(2).limit(2)]
    assert page2 == ["Meena", "Kiran"], page2

    print(f"  top 3 by maths: {top3}; page 2 by _id: {page2}")


def order_is_fixed():
    """sort, then skip, then limit -- whatever order you chain them."""
    db = fresh_db()

    a = [d["name"] for d in db.students.find().sort("marks.maths", -1).limit(3)]
    b = [d["name"] for d in db.students.find().limit(3).sort("marks.maths", -1)]
    assert a == b == ["Meena", "Asha", "Kiran"], (a, b)

    print("  .limit(3).sort(...) == .sort(...).limit(3) -- the server decides")
    print("       it sorts EVERYTHING and then takes three, not the reverse")


def range_pagination_beats_skip():
    db = fresh_db()

    # skip-based
    p1 = list(db.students.find().sort("_id", 1).limit(2))
    p2 = list(db.students.find().sort("_id", 1).skip(2).limit(2))
    p3 = list(db.students.find().sort("_id", 1).skip(4).limit(2))

    # range-based: remember the last _id seen
    r1 = list(db.students.find().sort("_id", 1).limit(2))
    r2 = list(db.students.find({"_id": {"$gt": r1[-1]["_id"]}})
                         .sort("_id", 1).limit(2))
    r3 = list(db.students.find({"_id": {"$gt": r2[-1]["_id"]}})
                         .sort("_id", 1).limit(2))

    for skip_page, range_page in [(p1, r1), (p2, r2), (p3, r3)]:
        assert [d["_id"] for d in skip_page] == [d["_id"] for d in range_page]

    print("  range pagination gives IDENTICAL pages, and every page costs the")
    print("       same -- skip(100000) would walk and discard 100,000 documents")


def cursor_versus_document():
    db = fresh_db()

    cursor = db.students.find({"_id": 21})
    assert not hasattr(cursor, "get"), "find() gives a CURSOR"
    assert list(cursor)[0]["name"] == "Asha"

    doc = db.students.find_one({"_id": 21})
    assert doc["name"] == "Asha", "findOne gives the DOCUMENT"

    print("  find() -> a cursor (iterate it); findOne() -> a document or None")


def main():
    print("Experiment 8 -- Sorting, limiting, skipping")
    # Step 1: Sort
    sorting()
    # Step 2: Limit and skip
    limit_and_skip()
    # Step 3: See sort, skip and limit apply in a fixed order
    order_is_fixed()
    # Step 4: Paginate by range, not by skip
    range_pagination_beats_skip()
    # Step 5: Tell a cursor from a document
    cursor_versus_document()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 08_sort_limit.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.find().sort({ "marks.maths": -1 })              // -1 descending
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find().sort({ dept: 1, "marks.maths": -1 })     // multi-key
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find().sort({ "marks.maths": -1 }).limit(3)     // top 3
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find().skip(2).limit(2)                         // "page 2"
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find().limit(3).sort({ "marks.maths": -1 })
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find().sort({ _id: 1 }).limit(2)                       // page 1
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  }
]
collegeDB> const lastSeenId = 22                                              // the last _id on page 1
collegeDB> db.students.find({ _id: { $gt: lastSeenId } }).sort({ _id: 1 }).limit(2)   // page 2
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find().sort({ "marks.maths": -1 }).limit(1).next().name   // topper
Meena

In Python, through mongomock, 08_sort_limit.py:

OUTPUT

Experiment 8 -- Sorting, limiting, skipping
  sort by maths desc -> ['Meena', 'Asha', 'Kiran', 'Ravi', 'Bhanu']
  multi-key sort (dept asc, maths desc) groups then orders within
  top 3 by maths: ['Meena', 'Asha', 'Kiran']; page 2 by _id: ['Meena', 'Kiran']
  .limit(3).sort(...) == .sort(...).limit(3) -- the server decides
       it sorts EVERYTHING and then takes three, not the reverse
  range pagination gives IDENTICAL pages, and every page costs the
       same -- skip(100000) would walk and discard 100,000 documents
  find() -> a cursor (iterate it); findOne() -> a document or None

Corrected: lastSeenId was used but never set, so the range query failed with "ReferenceError: lastSeenId is not defined". It is now set to 22, the last _id on page 1.

RESULT

limit(3).sort(...) gives the top three, not three sorted; page 2 by range is Meena and Kiran.

Experiment 9 — An embedded data model

1. Question

Design and query an embedded data model.

2. Aim

Store a student's address and enrolments inside the student, and query them.

3. Steps

In mongosh, 09_embedded.js:

  1. Insert two students, with their address and enrolments embedded.
  2. Read everything in one query.
  3. Query nested fields and arrays.
  4. Use $elemMatch for two conditions on one element.
  5. Update one array element, and push another.
  6. Total the credits.

In Python, through mongomock, 09_embedded.py:

  1. Read everything in one query.
  2. Query nested fields and arrays.
  3. Use $elemMatch for two conditions on one element.
  4. Update one array element.
  5. Push, and aggregate.
  6. State the model's limitation.

THE POINT

One read returns everything — no join anywhere — and the $elemMatch requirement for two conditions on the enrolment array.

4. Programme

In mongosh, 09_embedded.js:

// Experiment 9 -- Designing an Embedded Data Model for a student-course
// enrollment system.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 09_embedded.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)

// Step 1: Insert two students, with their address and enrolments embedded
use collegeDB
db.embedded.drop()

db.embedded.insertMany([
  { _id: 21, name: "Asha Kumari", dept: "DS",
    address: { city: "Vijayawada", state: "AP", pin: "520010" },   // 1-to-1
    enrollments: [                                                 // 1-to-few
      { course: "DSC301", title: "Data Science with R", grade: "A", credits: 4 },
      { course: "STA302", title: "Statistical Foundations", grade: "B", credits: 3 }
    ] },
  { _id: 22, name: "Ravi Teja", dept: "DS",
    address: { city: "Guntur", state: "AP", pin: "522002" },
    enrollments: [
      { course: "DSC301", title: "Data Science with R", grade: "C", credits: 4 }
    ] }
])

// ONE read gets the student, their address and every enrolment. No join.
// Step 2: Read everything in one query
db.embedded.findOne({ _id: 21 })

// Step 3: Query nested fields and arrays
db.embedded.find({ "address.city": "Vijayawada" })
db.embedded.find({ "enrollments.grade": "A" })

// TWO conditions on an array of sub-documents NEED $elemMatch, or different
// elements may satisfy different conditions.
// Step 4: Use $elemMatch for two conditions on one element
db.embedded.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })
db.embedded.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" })  // WRONG

// Updating one element: the POSITIONAL operator $
// Step 5: Update one array element, and push another
db.embedded.updateOne({ _id: 21, "enrollments.course": "STA302" },
                      { $set: { "enrollments.$.grade": "A" } })

db.embedded.updateOne({ _id: 22 },
  { $push: { enrollments: { course: "WEB303", title: "Web Technologies",
                            grade: "B", credits: 3 } } })

// Total credits, per student -- needs $unwind
// Step 6: Total the credits
db.embedded.aggregate([
  { $unwind: "$enrollments" },
  { $group: { _id: "$name", credits: { $sum: "$enrollments.credits" } } },
  { $sort: { _id: 1 } }
])
// [Changed: the $sort was added. $group returns its groups in no fixed order, and
// they came back in a different order from one run to the next.]

In Python, through mongomock, 09_embedded.py:

"""Experiment 9 — An embedded data model."""
import mongomock

DOCS = [
    {"_id": 21, "name": "Asha Kumari", "dept": "DS",
     "address": {"city": "Vijayawada", "state": "AP", "pin": "520010"},
     "enrollments": [
         {"course": "DSC301", "title": "Data Science with R", "grade": "A", "credits": 4},
         {"course": "STA302", "title": "Statistical Foundations", "grade": "B", "credits": 3}]},
    {"_id": 22, "name": "Ravi Teja", "dept": "DS",
     "address": {"city": "Guntur", "state": "AP", "pin": "522002"},
     "enrollments": [
         {"course": "DSC301", "title": "Data Science with R", "grade": "C", "credits": 4}]},
]


def db():
    d = mongomock.MongoClient().collegeDB
    d.embedded.insert_many([dict(x) for x in DOCS])
    return d


def one_read_gets_everything():
    d = db()
    doc = d.embedded.find_one({"_id": 21})
    assert doc["name"] == "Asha Kumari"
    assert doc["address"]["city"] == "Vijayawada"
    assert len(doc["enrollments"]) == 2
    assert doc["enrollments"][0]["title"] == "Data Science with R"

    print("  ONE read returned the student, the address and both enrolments")
    print("       -- no join anywhere, which is the point of embedding")


def querying_nested_and_arrays():
    d = db()
    assert [x["name"] for x in d.embedded.find({"address.city": "Vijayawada"})] \
        == ["Asha Kumari"]
    assert sorted(x["name"] for x in d.embedded.find({"enrollments.grade": "A"})) \
        == ["Asha Kumari"]
    assert sorted(x["name"] for x in d.embedded.find({"enrollments.course": "DSC301"})) \
        == ["Asha Kumari", "Ravi Teja"], "ANY element matches"
    print("  dot notation queries the embedded address and the enrolment array")


def elemmatch_is_required():
    """Two conditions on an array of sub-documents: different elements can
    satisfy different conditions unless you use $elemMatch."""
    d = db()
    # Ravi took DSC301 (grade C) and nothing with grade A -- so he should NOT
    # match "DSC301 with grade A".
    wrong = sorted(x["name"] for x in d.embedded.find(
        {"enrollments.course": "DSC301", "enrollments.grade": "A"}))
    right = sorted(x["name"] for x in d.embedded.find(
        {"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))

    assert right == ["Asha Kumari"], right
    assert wrong == ["Asha Kumari"], wrong   # Ravi has no grade A at all

    # Now make the trap visible: give Ravi an A in a DIFFERENT course.
    d.embedded.update_one({"_id": 22}, {"$push": {"enrollments": {
        "course": "WEB303", "title": "Web Technologies", "grade": "A", "credits": 3}}})

    wrong2 = sorted(x["name"] for x in d.embedded.find(
        {"enrollments.course": "DSC301", "enrollments.grade": "A"}))
    right2 = sorted(x["name"] for x in d.embedded.find(
        {"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))

    assert wrong2 == ["Asha Kumari", "Ravi Teja"], \
        "WRONG: Ravi matched using DSC301 from one element and grade A from another"
    assert right2 == ["Asha Kumari"], "RIGHT: one element must satisfy both"

    print("  after giving Ravi an A in a DIFFERENT course:")
    print(f"       without $elemMatch -> {wrong2}   (Ravi is a FALSE match)")
    print(f"       with    $elemMatch -> {right2}")


def positional_update():
    d = db()
    d.embedded.update_one({"_id": 21, "enrollments.course": "STA302"},
                          {"$set": {"enrollments.$.grade": "A"}})
    doc = d.embedded.find_one({"_id": 21})
    grades = {e["course"]: e["grade"] for e in doc["enrollments"]}
    assert grades == {"DSC301": "A", "STA302": "A"}, grades
    print("  the positional $ updated the element the QUERY matched")


def pushing_and_aggregating():
    d = db()
    d.embedded.update_one({"_id": 22}, {"$push": {"enrollments": {
        "course": "WEB303", "title": "Web Technologies", "grade": "B", "credits": 3}}})
    assert len(d.embedded.find_one({"_id": 22})["enrollments"]) == 2

    credits = {r["_id"]: r["credits"] for r in d.embedded.aggregate([
        {"$unwind": "$enrollments"},
        {"$group": {"_id": "$name", "credits": {"$sum": "$enrollments.credits"}}}])}
    assert credits == {"Asha Kumari": 7, "Ravi Teja": 7}, credits

    print(f"  total credits per student (needs $unwind): {credits}")


def the_limitation():
    """Embedding is right here because enrolments are BOUNDED. State the case
    where it would be wrong."""
    d = db()
    doc = d.embedded.find_one({"_id": 21})
    assert len(doc["enrollments"]) <= 10, "a degree has a bounded number of courses"

    print("  embedding is right here because a student's enrolments are BOUNDED")
    print("       -- attendance records or log entries would NOT be, and would")
    print("       eventually breach the 16 MB document limit")


def main():
    print("Experiment 9 -- An embedded data model")
    # Step 1: Read everything in one query
    one_read_gets_everything()
    # Step 2: Query nested fields and arrays
    querying_nested_and_arrays()
    # Step 3: Use $elemMatch for two conditions on one element
    elemmatch_is_required()
    # Step 4: Update one array element
    positional_update()
    # Step 5: Push, and aggregate
    pushing_and_aggregating()
    # Step 6: State the model's limitation
    the_limitation()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 09_embedded.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> db.embedded.drop()
true
collegeDB> db.embedded.insertMany([
...   { _id: 21, name: "Asha Kumari", dept: "DS",
...     address: { city: "Vijayawada", state: "AP", pin: "520010" },   // 1-to-1
...     enrollments: [                                                 // 1-to-few
...       { course: "DSC301", title: "Data Science with R", grade: "A", credits: 4 },
...       { course: "STA302", title: "Statistical Foundations", grade: "B", credits: 3 }
...     ] },
...   { _id: 22, name: "Ravi Teja", dept: "DS",
...     address: { city: "Guntur", state: "AP", pin: "522002" },
...     enrollments: [
...       { course: "DSC301", title: "Data Science with R", grade: "C", credits: 4 }
...     ] }
... ])
{ acknowledged: true, insertedIds: { '0': 21, '1': 22 } }
collegeDB> db.embedded.findOne({ _id: 21 })
{
  _id: 21,
  name: 'Asha Kumari',
  dept: 'DS',
  address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
  enrollments: [
    {
      course: 'DSC301',
      title: 'Data Science with R',
      grade: 'A',
      credits: 4
    },
    {
      course: 'STA302',
      title: 'Statistical Foundations',
      grade: 'B',
      credits: 3
    }
  ]
}
collegeDB> db.embedded.find({ "address.city": "Vijayawada" })
[
  {
    _id: 21,
    name: 'Asha Kumari',
    dept: 'DS',
    address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
    enrollments: [
      {
        course: 'DSC301',
        title: 'Data Science with R',
        grade: 'A',
        credits: 4
      },
      {
        course: 'STA302',
        title: 'Statistical Foundations',
        grade: 'B',
        credits: 3
      }
    ]
  }
]
collegeDB> db.embedded.find({ "enrollments.grade": "A" })
[
  {
    _id: 21,
    name: 'Asha Kumari',
    dept: 'DS',
    address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
    enrollments: [
      {
        course: 'DSC301',
        title: 'Data Science with R',
        grade: 'A',
        credits: 4
      },
      {
        course: 'STA302',
        title: 'Statistical Foundations',
        grade: 'B',
        credits: 3
      }
    ]
  }
]
collegeDB> db.embedded.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })
[
  {
    _id: 21,
    name: 'Asha Kumari',
    dept: 'DS',
    address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
    enrollments: [
      {
        course: 'DSC301',
        title: 'Data Science with R',
        grade: 'A',
        credits: 4
      },
      {
        course: 'STA302',
        title: 'Statistical Foundations',
        grade: 'B',
        credits: 3
      }
    ]
  }
]
collegeDB> db.embedded.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" })  // WRONG
[
  {
    _id: 21,
    name: 'Asha Kumari',
    dept: 'DS',
    address: { city: 'Vijayawada', state: 'AP', pin: '520010' },
    enrollments: [
      {
        course: 'DSC301',
        title: 'Data Science with R',
        grade: 'A',
        credits: 4
      },
      {
        course: 'STA302',
        title: 'Statistical Foundations',
        grade: 'B',
        credits: 3
      }
    ]
  }
]
collegeDB> db.embedded.updateOne({ _id: 21, "enrollments.course": "STA302" },
...                       { $set: { "enrollments.$.grade": "A" } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.embedded.updateOne({ _id: 22 },
...   { $push: { enrollments: { course: "WEB303", title: "Web Technologies",
...                             grade: "B", credits: 3 } } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.embedded.aggregate([
...   { $unwind: "$enrollments" },
...   { $group: { _id: "$name", credits: { $sum: "$enrollments.credits" } } },
...   { $sort: { _id: 1 } }
... ])
[
  { _id: 'Asha Kumari', credits: 7 },
  { _id: 'Ravi Teja', credits: 7 }
]

In Python, through mongomock, 09_embedded.py:

OUTPUT

Experiment 9 -- An embedded data model
  ONE read returned the student, the address and both enrolments
       -- no join anywhere, which is the point of embedding
  dot notation queries the embedded address and the enrolment array
  after giving Ravi an A in a DIFFERENT course:
       without $elemMatch -> ['Asha Kumari', 'Ravi Teja']   (Ravi is a FALSE match)
       with    $elemMatch -> ['Asha Kumari']
  the positional $ updated the element the QUERY matched
  total credits per student (needs $unwind): {'Asha Kumari': 7, 'Ravi Teja': 7}
  embedding is right here because a student's enrolments are BOUNDED
       -- attendance records or log entries would NOT be, and would
       eventually breach the 16 MB document limit

Changed: the credits total now ends with a $sort. $group returns its groups in no fixed order, and on the real server Asha and Ravi came back in a different order from one run to the next.

RESULT

One findOne returns the student with everything; $elemMatch finds Asha's A in DSC301 where the two separate conditions match wrongly.

Experiment 10 — A normalized model with references

1. Question

Design and query a normalised model, with documents that reference each other.

2. Aim

Join referenced collections with $lookup, and see what references do not guarantee.

3. Steps

In mongosh, 10_referenced.js:

  1. Insert a student, the courses and the enrolments, as references.
  2. Two reads, or one $lookup.
  3. See $lookup keep an unmatched row, as a left outer join.
  4. See references not enforced.
  5. Index the foreign fields.

In Python, through mongomock, 10_referenced.py:

  1. Two reads, or one $lookup.
  2. See $lookup always give an array.
  3. See $lookup is a left outer join.
  4. See references are not enforced.
  5. Index the foreign field.

THE POINT

$lookup produces an array even for a one-to-one match, which is why $unwind follows it; and it is a left outer join — an unmatched document gets an empty array, not nothing.

4. Programme

In mongosh, 10_referenced.js:

// Experiment 10 -- Designing a Normalized Data Model using document
// references.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 10_referenced.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)

// Step 1: Insert a student, the courses and the enrolments, as references
use collegeDB

db.students.insertOne({ _id: 21, name: "Asha Kumari", dept: "DS" })
db.courses.insertMany([
  { _id: "DSC301", title: "Data Science with R", credits: 4, instructor: "Dr. Rao" },
  { _id: "STA302", title: "Statistical Foundations", credits: 3, instructor: "Dr. Devi" }
])
db.enrollments.insertMany([
  { student_id: 21, course_id: "DSC301", grade: "A" },
  { student_id: 21, course_id: "STA302", grade: "B" }
])

// Two reads, application-side
// Step 2: Two reads, or one $lookup
const s = db.students.findOne({ _id: 21 })
const e = db.enrollments.find({ student_id: 21 }).toArray()

// Or one $lookup. Note: 'as' is ALWAYS an array, even for a 1-to-1 match,
// which is why $unwind almost always follows.
db.enrollments.aggregate([
  { $lookup: { from: "courses", localField: "course_id",
               foreignField: "_id", as: "course" } },
  { $unwind: "$course" },
  { $project: { _id: 0, course: "$course.title", grade: 1 } }
])

// $lookup is a LEFT OUTER JOIN -- an unmatched document gets an EMPTY ARRAY
// Step 3: See $lookup keep an unmatched row, as a left outer join
db.enrollments.insertOne({ student_id: 21, course_id: "GONE", grade: "F" })
db.enrollments.aggregate([
  { $lookup: { from: "courses", localField: "course_id",
               foreignField: "_id", as: "course" } }
])   // the GONE row has course: []

// NOTHING stops a reference pointing at a document that does not exist.
// In Course 5 a foreign key would. Here the application must check.
// Step 4: See references not enforced
db.courses.deleteOne({ _id: "DSC301" })     // the enrolments still reference it

// Index the foreignField, or every input document causes a collection scan
// Step 5: Index the foreign fields
db.enrollments.createIndex({ course_id: 1 })
db.enrollments.createIndex({ student_id: 1 })

In Python, through mongomock, 10_referenced.py:

"""Experiment 10 — A normalized model using document references."""
from fixtures import fresh_db


def two_reads_or_one_lookup():
    db = fresh_db()

    # Application-side: two reads.
    student = db.students.find_one({"_id": 21})
    enrolments = list(db.enrollments.find({"student_id": 21}))
    assert student["name"] == "Asha"
    assert len(enrolments) == 2

    # Or one $lookup.
    joined = list(db.enrollments.aggregate([
        {"$match": {"student_id": 21}},
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": "$course"},
        {"$project": {"_id": 0, "course": "$course.title", "grade": 1}}]))

    assert sorted(j["course"] for j in joined) == \
        ["Data Science with R", "Statistical Foundations"], joined

    print(f"  two reads, or one $lookup -> {[j['course'] for j in joined]}")


def lookup_always_produces_an_array():
    """Which is why $unwind almost always follows it."""
    db = fresh_db()

    raw = list(db.enrollments.aggregate([
        {"$match": {"student_id": 21, "course_id": "DSC301"}},
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}}]))

    assert isinstance(raw[0]["course"], list), "an ARRAY even for a 1-to-1 match"
    assert len(raw[0]["course"]) == 1

    unwound = list(db.enrollments.aggregate([
        {"$match": {"student_id": 21, "course_id": "DSC301"}},
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": "$course"}]))
    assert isinstance(unwound[0]["course"], dict), "$unwind makes it a sub-document"
    assert unwound[0]["course"]["title"] == "Data Science with R"

    print("  $lookup gave course: [ {...} ]; $unwind made it course: { ... }")
    print("       without $unwind, '$course.title' would be an ARRAY of titles")


def lookup_is_a_left_outer_join():
    db = fresh_db()
    db.enrollments.insert_one({"student_id": 21, "course_id": "GONE", "grade": "F"})

    rows = list(db.enrollments.aggregate([
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}}]))

    orphan = [r for r in rows if r["course_id"] == "GONE"][0]
    assert orphan["course"] == [], "an unmatched document gets an EMPTY ARRAY"
    assert len(rows) == 6, "the orphan is KEPT -- left outer join"

    # $unwind would then DROP it, unless you preserve empties.
    dropped = list(db.enrollments.aggregate([
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": "$course"}]))
    assert len(dropped) == 5, "$unwind silently dropped the orphan"

    kept = list(db.enrollments.aggregate([
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": {"path": "$course", "preserveNullAndEmptyArrays": True}}]))
    assert len(kept) == 6

    print("  $lookup kept the orphan with course: []; $unwind then DROPPED it")
    print("       preserveNullAndEmptyArrays: true keeps it -- 6 rows, not 5")


def references_are_not_enforced():
    """The row that matters most in the RDBMS comparison."""
    db = fresh_db()
    assert db.enrollments.count_documents({"course_id": "DSC301"}) == 2

    db.courses.delete_one({"_id": "DSC301"})
    assert db.courses.find_one({"_id": "DSC301"}) is None
    assert db.enrollments.count_documents({"course_id": "DSC301"}) == 2, \
        "the enrolments STILL reference a course that no longer exists"

    # An integrity check the application must run for itself.
    valid = set(db.courses.distinct("_id"))
    orphans = [e for e in db.enrollments.find()
               if e["course_id"] not in valid]
    assert len(orphans) == 2

    print("  deleting the course left 2 DANGLING references, with no error")
    print("       Course 5's foreign key would have refused; here the")
    print(f"       application must check -- {len(orphans)} orphans found")


def index_the_foreign_field():
    db = fresh_db()
    db.enrollments.create_index("course_id")
    db.enrollments.create_index("student_id")

    names = {i["name"] for i in db.enrollments.list_indexes()}
    assert "course_id_1" in names and "student_id_1" in names

    print("  indexed both foreignFields -- without them, $lookup scans the")
    print("       whole other collection ONCE PER INPUT DOCUMENT")


def main():
    print("Experiment 10 -- A normalized model with references")
    # Step 1: Two reads, or one $lookup
    two_reads_or_one_lookup()
    # Step 2: See $lookup always give an array
    lookup_always_produces_an_array()
    # Step 3: See $lookup is a left outer join
    lookup_is_a_left_outer_join()
    # Step 4: See references are not enforced
    references_are_not_enforced()
    # Step 5: Index the foreign field
    index_the_foreign_field()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 10_referenced.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> db.students.insertOne({ _id: 21, name: "Asha Kumari", dept: "DS" })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.courses.insertMany([
...   { _id: "DSC301", title: "Data Science with R", credits: 4, instructor: "Dr. Rao" },
...   { _id: "STA302", title: "Statistical Foundations", credits: 3, instructor: "Dr. Devi" }
... ])
{ acknowledged: true, insertedIds: { '0': 'DSC301', '1': 'STA302' } }
collegeDB> db.enrollments.insertMany([
...   { student_id: 21, course_id: "DSC301", grade: "A" },
...   { student_id: 21, course_id: "STA302", grade: "B" }
... ])
{
  acknowledged: true,
  insertedIds: {
    '0': ObjectId('6ac214d315618cb446277d3b'),
    '1': ObjectId('6ac214d315618cb446277d3c')
  }
}
collegeDB> const s = db.students.findOne({ _id: 21 })
collegeDB> const e = db.enrollments.find({ student_id: 21 }).toArray()
collegeDB> db.enrollments.aggregate([
...   { $lookup: { from: "courses", localField: "course_id",
...                foreignField: "_id", as: "course" } },
...   { $unwind: "$course" },
...   { $project: { _id: 0, course: "$course.title", grade: 1 } }
... ])
[
  { grade: 'A', course: 'Data Science with R' },
  { grade: 'B', course: 'Statistical Foundations' }
]
collegeDB> db.enrollments.insertOne({ student_id: 21, course_id: "GONE", grade: "F" })
{
  acknowledged: true,
  insertedId: ObjectId('6ac214d315618cb446277d3d')
}
collegeDB> db.enrollments.aggregate([
...   { $lookup: { from: "courses", localField: "course_id",
...                foreignField: "_id", as: "course" } }
... ])   // the GONE row has course: []
[
  {
    _id: ObjectId('6ac214d315618cb446277d3b'),
    student_id: 21,
    course_id: 'DSC301',
    grade: 'A',
    course: [
      {
        _id: 'DSC301',
        title: 'Data Science with R',
        credits: 4,
        instructor: 'Dr. Rao'
      }
    ]
  },
  {
    _id: ObjectId('6ac214d315618cb446277d3c'),
    student_id: 21,
    course_id: 'STA302',
    grade: 'B',
    course: [
      {
        _id: 'STA302',
        title: 'Statistical Foundations',
        credits: 3,
        instructor: 'Dr. Devi'
      }
    ]
  },
  {
    _id: ObjectId('6ac214d315618cb446277d3d'),
    student_id: 21,
    course_id: 'GONE',
    grade: 'F',
    course: []
  }
]
collegeDB> db.courses.deleteOne({ _id: "DSC301" })     // the enrolments still reference it
{ acknowledged: true, deletedCount: 1 }
collegeDB> db.enrollments.createIndex({ course_id: 1 })
course_id_1
collegeDB> db.enrollments.createIndex({ student_id: 1 })
student_id_1

In Python, through mongomock, 10_referenced.py:

OUTPUT

Experiment 10 -- A normalized model with references
  two reads, or one $lookup -> ['Data Science with R', 'Statistical Foundations']
  $lookup gave course: [ {...} ]; $unwind made it course: { ... }
       without $unwind, '$course.title' would be an ARRAY of titles
  $lookup kept the orphan with course: []; $unwind then DROPPED it
       preserveNullAndEmptyArrays: true keeps it -- 6 rows, not 5
  deleting the course left 2 DANGLING references, with no error
       Course 5's foreign key would have refused; here the
       application must check -- 2 orphans found
  indexed both foreignFields -- without them, $lookup scans the
       whole other collection ONCE PER INPUT DOCUMENT

RESULT

$lookup returns an array, and keeps the GONE enrolment with an empty one; deleting a course leaves its enrolments pointing at nothing.

Experiment 11 — One-to-one, one-to-many, many-to-many

1. Question

Model one-to-one, one-to-many and many-to-many relationships.

2. Aim

Model each kind of relationship, and query it from both ends.

3. Steps

In mongosh, 11_relationships.js:

  1. One-to-one: embed.
  2. One-to-many: reference from the child.
  3. One-to-few: embed an array.
  4. Many-to-many: reference both ways.

In Python, through mongomock, 11_relationships.py:

  1. One-to-one: embed.
  2. One-to-few: embed an array.
  3. One-to-many: reference from the child.
  4. Many-to-many: reference both ways.
  5. Give the relationship's own data to a junction collection.

THE POINT

All three modelled and queried:

Relationship Model Query direction
One-to-one Embed the address Both, from the student
One-to-many Reference from the child Course → its enrolments
Many-to-many A junction collection Both directions

The junction collection is what carries the grade — an attribute of the relationship, belonging to neither entity.

4. Programme

In mongosh, 11_relationships.js:

// Experiment 11 -- Modeling relationships: One-to-One, One-to-Many,
// Many-to-Many in MongoDB.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 11_relationships.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)

use collegeDB

// Step 1: One-to-one: embed
// Small, bounded, always read together, never queried alone.
db.people.insertOne({
  _id: 21, name: "Asha Kumari",
  address: { city: "Vijayawada", state: "AP", pin: "520010" }
})
db.people.find({ "address.pin": "520010" })

// Step 2: One-to-many: reference from the child
// A course has many enrolments. The array must NOT live on the course, because
// it is unbounded -- that is the 16 MB trap.
db.courses.insertOne({ _id: "DSC301", title: "Data Science with R" })
db.enrollments.insertMany([
  { course_id: "DSC301", student_id: 21, grade: "A" },
  { course_id: "DSC301", student_id: 22, grade: "C" }
])
db.enrollments.find({ course_id: "DSC301" })          // the "many" side

// Step 3: One-to-few: embed an array
db.people.updateOne({ _id: 21 },
  { $set: { phones: ["9876543210", "9876543211"] } })

// Step 4: Many-to-many: reference both ways
// Option A: an array of ids on one side
db.students.updateOne({ _id: 21 },
  { $set: { name: "Asha Kumari", course_ids: ["DSC301", "STA302"] } },
  { upsert: true })
// [Corrected: this was an updateOne without upsert, on a student 21 who exists
// only if Experiment 10 has just been run in the same database. On its own it
// matched nothing, and the find below returned no one.]
db.students.find({ course_ids: "DSC301" })            // who takes DSC301?

// Option C: a junction collection -- REQUIRED when the relationship itself
// has attributes. The grade belongs to neither the student nor the course.
db.enrollments.find({ student_id: 21 })               // this student's courses
db.enrollments.find({ course_id: "DSC301" })          // this course's students

In Python, through mongomock, 11_relationships.py:

"""Experiment 11 — One-to-one, one-to-many and many-to-many."""
import mongomock
from fixtures import fresh_db


def one_to_one_embed():
    db = mongomock.MongoClient().collegeDB
    db.people.insert_one({"_id": 21, "name": "Asha Kumari",
                          "address": {"city": "Vijayawada", "state": "AP",
                                      "pin": "520010"}})

    doc = db.people.find_one({"_id": 21})
    assert doc["address"]["city"] == "Vijayawada", "one read, no join"
    assert [d["name"] for d in db.people.find({"address.pin": "520010"})] \
        == ["Asha Kumari"], "and it is still queryable"

    print("  1-to-1: EMBED. One read; the sub-document is still queryable by dot")


def one_to_few_embed_array():
    db = mongomock.MongoClient().collegeDB
    db.people.insert_one({"_id": 21, "name": "Asha",
                          "phones": ["9876543210", "9876543211"]})

    assert len(db.people.find_one({"_id": 21})["phones"]) == 2
    assert [d["_id"] for d in db.people.find({"phones": "9876543210"})] == [21], \
        "any element matches"

    print("  1-to-few: embed as an ARRAY -- bounded, so no 16 MB risk")


def one_to_many_reference_from_child():
    """The array must NOT live on the parent: it is unbounded."""
    db = fresh_db()

    # The CHILD holds the parent's id.
    for e in db.enrollments.find({"course_id": "DSC301"}):
        assert "course_id" in e

    got = sorted(e["student_id"] for e in db.enrollments.find({"course_id": "DSC301"}))
    assert got == [21, 22], got

    # Adding a thousand enrolments does not grow the course document at all.
    before = len(str(db.courses.find_one({"_id": "DSC301"})))
    db.enrollments.insert_many([{"course_id": "DSC301", "student_id": 1000 + i,
                                 "grade": "B"} for i in range(1000)])
    after = len(str(db.courses.find_one({"_id": "DSC301"})))
    assert before == after, "the COURSE document is unchanged -- that is the point"
    assert db.enrollments.count_documents({"course_id": "DSC301"}) == 1002

    print(f"  1-to-many: reference from the CHILD. 1000 more enrolments left the")
    print(f"       course document at {after} chars -- an embedded array would")
    print(f"       have grown it, and eventually breached 16 MB")


def many_to_many_both_ways():
    db = fresh_db()

    # Option A: an array of ids on the student
    db.students.update_one({"_id": 21},
                           {"$set": {"course_ids": ["DSC301", "STA302"]}})
    assert [d["name"] for d in db.students.find({"course_ids": "DSC301"})] == ["Asha"]

    # Option C: the junction collection -- queryable from BOTH directions
    hers = sorted(e["course_id"] for e in db.enrollments.find({"student_id": 21}))
    assert hers == ["DSC301", "STA302"], hers

    theirs = sorted(e["student_id"] for e in db.enrollments.find({"course_id": "DSC301"}))
    assert theirs == [21, 22], theirs

    print(f"  M-to-M: student 21 takes {hers}; DSC301 has students {theirs}")


def the_junction_carries_the_relationship_attributes():
    """Why option C wins when the relationship has its own data."""
    db = fresh_db()

    e = db.enrollments.find_one({"student_id": 21, "course_id": "DSC301"})
    assert e["grade"] == "A"

    # The grade belongs to NEITHER entity:
    assert "grade" not in db.students.find_one({"_id": 21})
    assert "grade" not in db.courses.find_one({"_id": "DSC301"})

    # An array of ids could not hold it.
    db.students.update_one({"_id": 21}, {"$set": {"course_ids": ["DSC301"]}})
    s = db.students.find_one({"_id": 21})
    assert s["course_ids"] == ["DSC301"], "just an id -- nowhere to put the grade"

    print("  the GRADE lives on the enrolment, not on the student or the course")
    print("       -- exactly the reasoning that produces a junction TABLE in")
    print("       Course 5, and it survives the translation unchanged")


def main():
    print("Experiment 11 -- Modelling relationships")
    # Step 1: One-to-one: embed
    one_to_one_embed()
    # Step 2: One-to-few: embed an array
    one_to_few_embed_array()
    # Step 3: One-to-many: reference from the child
    one_to_many_reference_from_child()
    # Step 4: Many-to-many: reference both ways
    many_to_many_both_ways()
    # Step 5: Give the relationship's own data to a junction collection
    the_junction_carries_the_relationship_attributes()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 11_relationships.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> db.people.insertOne({
...   _id: 21, name: "Asha Kumari",
...   address: { city: "Vijayawada", state: "AP", pin: "520010" }
... })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.people.find({ "address.pin": "520010" })
[
  {
    _id: 21,
    name: 'Asha Kumari',
    address: { city: 'Vijayawada', state: 'AP', pin: '520010' }
  }
]
collegeDB> db.courses.insertOne({ _id: "DSC301", title: "Data Science with R" })
{ acknowledged: true, insertedId: 'DSC301' }
collegeDB> db.enrollments.insertMany([
...   { course_id: "DSC301", student_id: 21, grade: "A" },
...   { course_id: "DSC301", student_id: 22, grade: "C" }
... ])
{
  acknowledged: true,
  insertedIds: {
    '0': ObjectId('6ac214d8097ed1d9dbdc315a'),
    '1': ObjectId('6ac214d8097ed1d9dbdc315b')
  }
}
collegeDB> db.enrollments.find({ course_id: "DSC301" })          // the "many" side
[
  {
    _id: ObjectId('6ac214d8097ed1d9dbdc315a'),
    course_id: 'DSC301',
    student_id: 21,
    grade: 'A'
  },
  {
    _id: ObjectId('6ac214d8097ed1d9dbdc315b'),
    course_id: 'DSC301',
    student_id: 22,
    grade: 'C'
  }
]
collegeDB> db.people.updateOne({ _id: 21 },
...   { $set: { phones: ["9876543210", "9876543211"] } })
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 1,
  modifiedCount: 1,
  upsertedCount: 0
}
collegeDB> db.students.updateOne({ _id: 21 },
...   { $set: { name: "Asha Kumari", course_ids: ["DSC301", "STA302"] } },
...   { upsert: true })
{
  acknowledged: true,
  insertedId: 21,
  matchedCount: 0,
  modifiedCount: 0,
  upsertedCount: 1
}
collegeDB> db.students.find({ course_ids: "DSC301" })            // who takes DSC301?
[
  { _id: 21, course_ids: [ 'DSC301', 'STA302' ], name: 'Asha Kumari' }
]
collegeDB> db.enrollments.find({ student_id: 21 })               // this student's courses
[
  {
    _id: ObjectId('6ac214d8097ed1d9dbdc315a'),
    course_id: 'DSC301',
    student_id: 21,
    grade: 'A'
  }
]
collegeDB> db.enrollments.find({ course_id: "DSC301" })          // this course's students
[
  {
    _id: ObjectId('6ac214d8097ed1d9dbdc315a'),
    course_id: 'DSC301',
    student_id: 21,
    grade: 'A'
  },
  {
    _id: ObjectId('6ac214d8097ed1d9dbdc315b'),
    course_id: 'DSC301',
    student_id: 22,
    grade: 'C'
  }
]

In Python, through mongomock, 11_relationships.py:

OUTPUT

Experiment 11 -- Modelling relationships
  1-to-1: EMBED. One read; the sub-document is still queryable by dot
  1-to-few: embed as an ARRAY -- bounded, so no 16 MB risk
  1-to-many: reference from the CHILD. 1000 more enrolments left the
       course document at 88 chars -- an embedded array would
       have grown it, and eventually breached 16 MB
  M-to-M: student 21 takes ['DSC301', 'STA302']; DSC301 has students [21, 22]
  the GRADE lives on the enrolment, not on the student or the course
       -- exactly the reasoning that produces a junction TABLE in
       Course 5, and it survives the translation unchanged

Corrected: the many-to-many's updateOne was on student 21, who exists only if Experiment 10 has just been run in the same database; on its own it matched nothing, and "who takes DSC301?" returned no one. It is now an upsert, which works either way.

RESULT

Each relationship is modelled and queried both ways; the junction collection holds the grade.

Experiment 12 — Schema validation with JSON Schema

1. Question

Validate documents against a JSON Schema.

2. Aim

Enforce a schema on a collection, and add one to a collection that already holds data.

3. Steps

In mongosh, 12_validation.js:

  1. Load the sample data.
  2. Write the schema.
  3. Create a collection that enforces it.
  4. Insert a conforming document, and four that are not.
  5. Add validation to data already there, in stages.

In Python, through mongomock, 12_validation.py:

  1. Pass a conforming document.
  2. Catch each violation.
  3. Find the documents that do not conform.
  4. Tighten validation in stages.

THE POINT

Attach the schema in the safest mode first — moderate + warn — find the offenders with { $nor: [ { $jsonSchema: … } ] }, fix them, then tighten to strict + error.

mongomock does not enforce $jsonSchema, so the Python half implements the same rules in code and asserts that conforming documents pass and each kind of violation is caught — among them a missing required field, a wrong type, a value outside the enum. The real server enforces them, and its error names the rule that failed.

4. Programme

In mongosh, 12_validation.js:

// Experiment 12 -- Implementing schema validation using JSON Schema.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 12_validation.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")
db.validated.drop()

// Step 2: Write the schema
const schema = {
  bsonType: "object",
  required: ["roll", "name", "dept"],
  properties: {
    roll: { bsonType: "int", minimum: 1,
            description: "required integer, at least 1" },
    name: { bsonType: "string", minLength: 3, maxLength: 80 },
    dept: { enum: ["DS", "Stats", "CS"],
            description: "must be DS, Stats or CS" },
    marks: {
      bsonType: "object",
      properties: {
        maths: { bsonType: "int", minimum: 0, maximum: 100 },
        stats: { bsonType: "int", minimum: 0, maximum: 100 }
      }
    },
    email: { bsonType: "string", pattern: "^[^\\s@]+@[^\\s@]+\\.[^\\s@]{2,}$" }
  }
}
// Step 3: Create a collection that enforces it
db.createCollection("validated", {
  validator: { $jsonSchema: schema },
  validationLevel: "strict",
  validationAction: "error"
})

// Step 4: Insert a conforming document, and four that are not
db.validated.insertOne({ roll: NumberInt(21), name: "Asha", dept: "DS",
                         marks: { maths: NumberInt(88) } })     // OK

db.validated.insertOne({ name: "NoRoll", dept: "DS" })          // missing required
db.validated.insertOne({ roll: NumberInt(22), name: "Ab", dept: "DS" })  // too short
db.validated.insertOne({ roll: NumberInt(23), name: "Ravi", dept: "Physics" })  // enum
db.validated.insertOne({ roll: NumberInt(24), name: "Meena", dept: "DS",
                         marks: { maths: NumberInt(150) } })    // out of range

// Step 5: Add validation to data already there, in stages
// 1. Attach it in the SAFEST mode: log, do not block.
db.runCommand({ collMod: "students",
                validator: { $jsonSchema: schema },
                validationLevel: "moderate",
                validationAction: "warn" })

// 2. FIND the offenders -- $nor inverts a $jsonSchema match
db.students.find({ $nor: [ { $jsonSchema: schema } ] })
// [Corrected: these two read { $jsonSchema: { /* as above */ } }, an empty
// schema, which every document passes, so no offender could be found. The
// schema is now a const, defined once and used three times.]

// 3. Fix them, confirm the count is zero, then:
db.runCommand({ collMod: "students",
                validationLevel: "strict", validationAction: "error" })

In Python, through mongomock, 12_validation.py:

"""Experiment 12 — Schema validation with JSON Schema.

*** mongomock does NOT enforce $jsonSchema. ***

Rather than pretend it does, this script implements the SAME rules in code and
asserts that a conforming document passes and each kind of violation is caught.
The mongosh half (12_validation.js) is what you run on a real server.
"""
import re
import mongomock

SCHEMA = {
    "bsonType": "object",
    "required": ["roll", "name", "dept"],
    "properties": {
        "roll":  {"bsonType": "int", "minimum": 1},
        "name":  {"bsonType": "string", "minLength": 3, "maxLength": 80},
        "dept":  {"enum": ["DS", "Stats", "CS"]},
        "marks": {"bsonType": "object", "properties": {
            "maths": {"bsonType": "int", "minimum": 0, "maximum": 100},
            "stats": {"bsonType": "int", "minimum": 0, "maximum": 100}}},
        "email": {"bsonType": "string",
                  "pattern": r"^[^\s@]+@[^\s@]+\.[^\s@]{2,}$"},
    },
}

BSON_TYPES = {"int": int, "string": str, "object": dict, "double": float,
              "bool": bool, "array": list}


def violations(doc, schema=SCHEMA, path=""):
    """Return the list of ways `doc` fails `schema`. Empty means it conforms."""
    out = []
    for field in schema.get("required", []):
        if field not in doc:
            out.append(f"{path}{field}: required field missing")

    for field, rule in schema.get("properties", {}).items():
        if field not in doc:
            continue
        value = doc[field]
        where = f"{path}{field}"

        want = rule.get("bsonType")
        if want and not isinstance(value, BSON_TYPES[want]):
            out.append(f"{where}: expected {want}, got {type(value).__name__}")
            continue
        if "enum" in rule and value not in rule["enum"]:
            out.append(f"{where}: {value!r} not in {rule['enum']}")
        if "minimum" in rule and value < rule["minimum"]:
            out.append(f"{where}: {value} below minimum {rule['minimum']}")
        if "maximum" in rule and value > rule["maximum"]:
            out.append(f"{where}: {value} above maximum {rule['maximum']}")
        if "minLength" in rule and len(value) < rule["minLength"]:
            out.append(f"{where}: shorter than {rule['minLength']}")
        if "maxLength" in rule and len(value) > rule["maxLength"]:
            out.append(f"{where}: longer than {rule['maxLength']}")
        if "pattern" in rule and not re.match(rule["pattern"], value):
            out.append(f"{where}: does not match {rule['pattern']}")
        if rule.get("bsonType") == "object" and "properties" in rule:
            out.extend(violations(value, rule, path=f"{where}."))
    return out


def conforming_document_passes():
    good = {"roll": 21, "name": "Asha", "dept": "DS", "marks": {"maths": 88},
            "email": "asha@nri.ac.in"}
    assert violations(good) == [], violations(good)
    print("  a conforming document produces no violations")


def each_violation_is_caught():
    cases = {
        "missing required": ({"name": "NoRoll", "dept": "DS"}, "required"),
        "name too short":   ({"roll": 22, "name": "Ab", "dept": "DS"}, "shorter"),
        "dept not in enum": ({"roll": 23, "name": "Ravi", "dept": "Physics"}, "not in"),
        "marks over 100":   ({"roll": 24, "name": "Meena", "dept": "DS",
                              "marks": {"maths": 150}}, "above maximum"),
        "roll below 1":     ({"roll": 0, "name": "Zero", "dept": "DS"}, "below minimum"),
        "roll wrong type":  ({"roll": "21", "name": "Str", "dept": "DS"}, "expected int"),
        "bad email":        ({"roll": 25, "name": "Bhanu", "dept": "DS",
                              "email": "not-an-email"}, "does not match"),
    }
    for label, (doc, expected) in cases.items():
        found = violations(doc)
        assert found, f"{label}: expected a violation, got none"
        assert any(expected in v for v in found), f"{label}: {found}"
        print(f"    {label:18s} -> {found[0]}")

    print("  every rule in the schema is enforced")


def find_the_offenders():
    """The migration step: find what does NOT conform, before tightening."""
    db = mongomock.MongoClient().collegeDB
    db.messy.insert_many([
        {"roll": 21, "name": "Asha", "dept": "DS"},          # ok
        {"roll": 22, "name": "Ab", "dept": "DS"},            # name too short
        {"name": "NoRoll", "dept": "Stats"},                 # missing roll
        {"roll": 24, "name": "Kiran", "dept": "Physics"},    # bad dept
    ])

    offenders = [(d.get("name"), violations(d)) for d in db.messy.find()
                 if violations(d)]
    assert len(offenders) == 3, offenders
    assert all(v for _, v in offenders)

    conforming = [d for d in db.messy.find() if not violations(d)]
    assert len(conforming) == 1 and conforming[0]["name"] == "Asha"

    print(f"  3 of 4 documents violate the schema:")
    for name, vs in offenders:
        print(f"    {str(name):10s} {vs[0]}")
    print("  on a real server: $nor: [ { $jsonSchema: ... } ] finds exactly these")


def the_migration_path():
    """Turning strict validation on over dirty data breaks the application."""
    levels = {
        "strict":   "applies to EVERY insert and update",
        "moderate": "applies to inserts, and to updates of documents that ALREADY conform",
        "off":      "applies to nothing",
    }
    actions = {"error": "REJECT the write", "warn": "LOG it and accept"}

    assert set(levels) == {"strict", "moderate", "off"}
    assert set(actions) == {"error", "warn"}

    print("  validationLevel:")
    for k, v in levels.items():
        print(f"    {k:9s} {v}")
    print("  validationAction:")
    for k, v in actions.items():
        print(f"    {k:9s} {v}")
    print("  migration: moderate+warn -> find offenders -> fix -> strict+error")
    print("       going straight to strict+error breaks every update to a")
    print("       non-conforming document, INCLUDING the one that would fix it")


def main():
    print("Experiment 12 -- Schema validation")
    print("  NOTE: mongomock does not enforce $jsonSchema, so the same rules")
    print("        are implemented in code here and asserted.")
    # Step 1: Pass a conforming document
    conforming_document_passes()
    # Step 2: Catch each violation
    each_violation_is_caught()
    # Step 3: Find the documents that do not conform
    find_the_offenders()
    # Step 4: Tighten validation in stages
    the_migration_path()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 12_validation.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.validated.drop()
true
collegeDB> const schema = {
...   bsonType: "object",
...   required: ["roll", "name", "dept"],
...   properties: {
...     roll: { bsonType: "int", minimum: 1,
...             description: "required integer, at least 1" },
...     name: { bsonType: "string", minLength: 3, maxLength: 80 },
...     dept: { enum: ["DS", "Stats", "CS"],
...             description: "must be DS, Stats or CS" },
...     marks: {
...       bsonType: "object",
...       properties: {
...         maths: { bsonType: "int", minimum: 0, maximum: 100 },
...         stats: { bsonType: "int", minimum: 0, maximum: 100 }
...       }
...     },
...     email: { bsonType: "string", pattern: "^[^\\s@]+@[^\\s@]+\\.[^\\s@]{2,}$" }
...   }
... }
collegeDB> db.createCollection("validated", {
...   validator: { $jsonSchema: schema },
...   validationLevel: "strict",
...   validationAction: "error"
... })
{ ok: 1 }
collegeDB> db.validated.insertOne({ roll: NumberInt(21), name: "Asha", dept: "DS",
...                          marks: { maths: NumberInt(88) } })     // OK
{
  acknowledged: true,
  insertedId: ObjectId('6ac214dd5e9ae24277f97aaa')
}
collegeDB> db.validated.insertOne({ name: "NoRoll", dept: "DS" })          // missing required
Uncaught:

MongoServerError: Document failed validation
Additional information: {
  failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aab'),
  details: {
    operatorName: '$jsonSchema',
    schemaRulesNotSatisfied: [
      {
        operatorName: 'required',
        specifiedAs: { required: [ 'roll', 'name', 'dept' ] },
        missingProperties: [ 'roll' ]
      }
    ]
  }
}
collegeDB> db.validated.insertOne({ roll: NumberInt(22), name: "Ab", dept: "DS" })  // too short
Uncaught:

MongoServerError: Document failed validation
Additional information: {
  failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aac'),
  details: {
    operatorName: '$jsonSchema',
    schemaRulesNotSatisfied: [
      {
        operatorName: 'properties',
        propertiesNotSatisfied: [
          {
            propertyName: 'name',
            details: [
              {
                operatorName: 'minLength',
                specifiedAs: { minLength: 3 },
                reason: 'specified string length was not satisfied',
                consideredValue: 'Ab'
              }
            ]
          }
        ]
      }
    ]
  }
}
collegeDB> db.validated.insertOne({ roll: NumberInt(23), name: "Ravi", dept: "Physics" })  // enum
Uncaught:

MongoServerError: Document failed validation
Additional information: {
  failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aad'),
  details: {
    operatorName: '$jsonSchema',
    schemaRulesNotSatisfied: [
      {
        operatorName: 'properties',
        propertiesNotSatisfied: [
          {
            propertyName: 'dept',
            description: 'must be DS, Stats or CS',
            details: [
              {
                operatorName: 'enum',
                specifiedAs: { enum: [ 'DS', 'Stats', 'CS' ] },
                reason: 'value was not found in enum',
                consideredValue: 'Physics'
              }
            ]
          }
        ]
      }
    ]
  }
}
collegeDB> db.validated.insertOne({ roll: NumberInt(24), name: "Meena", dept: "DS",
...                          marks: { maths: NumberInt(150) } })    // out of range
Uncaught:

MongoServerError: Document failed validation
Additional information: {
  failingDocumentId: ObjectId('6ac214dd5e9ae24277f97aae'),
  details: {
    operatorName: '$jsonSchema',
    schemaRulesNotSatisfied: [
      {
        operatorName: 'properties',
        propertiesNotSatisfied: [
          {
            propertyName: 'marks',
            details: [
              {
                operatorName: 'properties',
                propertiesNotSatisfied: [
                  {
                    propertyName: 'maths',
                    details: [
                      {
                        operatorName: 'maximum',
                        specifiedAs: { maximum: 100 },
                        reason: 'comparison failed',
                        consideredValue: 150
                      }
                    ]
                  }
                ]
              }
            ]
          }
        ]
      }
    ]
  }
}
collegeDB> db.runCommand({ collMod: "students",
...                 validator: { $jsonSchema: schema },
...                 validationLevel: "moderate",
...                 validationAction: "warn" })
{ ok: 1 }
collegeDB> db.students.find({ $nor: [ { $jsonSchema: schema } ] })
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: [ 'Stats' ],
    age: 21,
    active: true
  }
]
collegeDB> db.runCommand({ collMod: "students",
...                 validationLevel: "strict", validationAction: "error" })
{ ok: 1 }

In Python, through mongomock, 12_validation.py:

OUTPUT

Experiment 12 -- Schema validation
  NOTE: mongomock does not enforce $jsonSchema, so the same rules
        are implemented in code here and asserted.
  a conforming document produces no violations
    missing required   -> roll: required field missing
    name too short     -> name: shorter than 3
    dept not in enum   -> dept: 'Physics' not in ['DS', 'Stats', 'CS']
    marks over 100     -> marks.maths: 150 above maximum 100
    roll below 1       -> roll: 0 below minimum 1
    roll wrong type    -> roll: expected int, got str
    bad email          -> email: does not match ^[^\s@]+@[^\s@]+\.[^\s@]{2,}$
  every rule in the schema is enforced
  3 of 4 documents violate the schema:
    Ab         name: shorter than 3
    NoRoll     roll: required field missing
    Kiran      dept: 'Physics' not in ['DS', 'Stats', 'CS']
  on a real server: $nor: [ { $jsonSchema: ... } ] finds exactly these
  validationLevel:
    strict    applies to EVERY insert and update
    moderate  applies to inserts, and to updates of documents that ALREADY conform
    off       applies to nothing
  validationAction:
    error     REJECT the write
    warn      LOG it and accept
  migration: moderate+warn -> find offenders -> fix -> strict+error
       going straight to strict+error breaks every update to a
       non-conforming document, INCLUDING the one that would fix it

Corrected: the second part read { $jsonSchema: { /* as above */ } }, an empty schema, which every document passes, so no offender could be found. The schema is now a const, written once and used three times.

RESULT

The conforming document goes in and each of the four violations is refused, by MongoDB itself; all five sample students fail the schema, having no roll.

Experiment 13 — Single-field and compound indexes

1. Question

Create single-field and compound indexes, and measure their effect.

2. Aim

Create indexes, read explain(), and apply the prefix and ESR rules.

3. Steps

In mongosh, 13_indexes.js:

  1. Load the sample data.
  2. Create and list indexes.
  3. Measure with explain().
  4. Apply the prefix rule.
  5. Apply the ESR rule.
  6. Cover a query.
  7. See two missing fields collide in a unique index.

In Python, through mongomock, 13_indexes.py:

  1. Create and list indexes.
  2. See a unique index enforced.
  3. See two missing fields collide.
  4. Apply the prefix rule.
  5. Apply the ESR rule.
  6. Cover a query.
  7. Count what indexes cost.

THE POINT

Run explain("executionStats") and read totalDocsExamined / nReturned — that ratio, not the wall-clock time, is what tells you whether the index is right. mongomock records indexes but has no planner, so the Python half asserts that the indexes are created and listed correctly and that a unique index rejects a duplicate. The prefix rule and ESR are demonstrated as a table of which queries each index serves; the server shows the plans.

4. Programme

In mongosh, 13_indexes.js:

// Experiment 13 -- Creating and testing single-field and compound indexes.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 13_indexes.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: Create and list indexes
db.students.createIndex({ dept: 1 })
db.students.createIndex({ dept: 1, "marks.maths": -1 })
db.students.createIndex({ email: 1 }, { unique: true })   // FAILS here: see below
db.students.createIndex({ dept: 1 }, { name: "dept_idx" })   // FAILS: see below
// Both fail, and both failures are worth knowing. None of the five students
// has an email, so all five index as email: null -- five duplicates of null.
// And { dept: 1 } is already indexed, as dept_1: the same keys under a second
// name are refused. [Note added: these two lines were written expecting both
// to succeed. Run, they fail, as above.]
db.students.getIndexes()
db.students.totalIndexSize()

// Step 3: Measure with explain()
db.students.find({ dept: "DS" }).explain("executionStats")
// stage:              COLLSCAN (bad)  vs  IXSCAN (good)
// totalDocsExamined / nReturned:  1 is ideal, 1000 means the index is wrong

// Step 4: Apply the prefix rule
db.students.createIndex({ dept: 1, year: 1, cgpa: 1 })
db.students.find({ dept: "DS" })                        // uses it
db.students.find({ dept: "DS", year: 4 })               // uses it
db.students.find({ year: 4 })                           // does NOT -- COLLSCAN
db.students.createIndex({ year: 1 })                    // so this is needed too

// Step 5: Apply the ESR rule
// Query: dept = "DS", maths > 70, sorted by age
db.students.createIndex({ dept: 1, age: 1, "marks.maths": 1 })
//                        ^equality ^sort   ^range
// A range predicate leaves everything AFTER it unordered, so a sort field
// placed after a range field cannot use the index.

// Step 6: Cover a query
db.students.createIndex({ dept: 1, name: 1 })
db.students.find({ dept: "DS" }, { _id: 0, dept: 1, name: 1 }).explain("executionStats")
// totalDocsExamined: 0
// [Corrected: the .explain(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find() without it, then rejected ".explain(...)" as an invalid command.]

// Step 7: See two missing fields collide in a unique index
db.people.drop()
db.people.createIndex({ email: 1 }, { unique: true })
db.people.insertOne({ name: "A" })         // ok -- email missing, indexed as null
db.people.insertOne({ name: "B" })         // DUPLICATE KEY ERROR -- a second null
db.people.dropIndex("email_1")
db.people.createIndex({ email: 1 },
  { unique: true, partialFilterExpression: { email: { $exists: true } } })
db.people.insertOne({ name: "B" })         // ok now: missing emails are not indexed
// [Corrected: this used db.students, whose five students already have no email,
// so the unique index could not be built at all, and the insert of B, commented
// DUPLICATE KEY ERROR, succeeded. A new collection, as in 13_indexes.py,
// shows what the comment says.]

db.students.dropIndex("dept_1")

In Python, through mongomock, 13_indexes.py:

"""Experiment 13 — Single-field and compound indexes.

mongomock records indexes but does not report IXSCAN, so this script asserts
that indexes are CREATED and ENFORCED correctly, and demonstrates the prefix
rule and ESR as tables. Run explain("executionStats") on a real server to see
the plan.
"""
import mongomock
from pymongo.errors import DuplicateKeyError
from fixtures import fresh_db


def creating_and_listing():
    db = fresh_db()
    db.students.create_index("dept")
    db.students.create_index([("dept", 1), ("marks.maths", -1)])
    db.students.create_index("age", name="age_idx")

    names = {i["name"] for i in db.students.list_indexes()}
    assert "_id_" in names, "_id is indexed AUTOMATICALLY"
    assert "dept_1" in names
    assert "age_idx" in names, "a custom name"
    assert any("marks.maths" in n for n in names), names

    db.students.drop_index("dept_1")
    assert "dept_1" not in {i["name"] for i in db.students.list_indexes()}

    print(f"  created and listed {len(names)} indexes, including the automatic _id")


def unique_is_enforced():
    db = fresh_db()
    db.students.create_index("name", unique=True)

    try:
        db.students.insert_one({"_id": 99, "name": "Asha"})
        raise AssertionError("expected a DuplicateKeyError")
    except DuplicateKeyError:
        pass

    db.students.insert_one({"_id": 99, "name": "Unique Name"})
    assert db.students.count_documents({}) == 6

    print("  a unique index rejected a duplicate name and accepted a new one")


def unique_and_missing_fields():
    """A missing field indexes as null, and two nulls collide."""
    db = mongomock.MongoClient().collegeDB
    db.people.create_index("email", unique=True)

    db.people.insert_one({"_id": 1, "name": "A"})          # no email -> null
    try:
        db.people.insert_one({"_id": 2, "name": "B"})      # also null
        raise AssertionError("expected a DuplicateKeyError from two nulls")
    except DuplicateKeyError:
        pass

    print("  a unique index allowed ONE document with no email, then rejected")
    print("       the next -- a missing field indexes as null, and nulls collide")
    print("       fix: partialFilterExpression: { email: { $exists: true } }")


def the_prefix_rule():
    """An index on {a,b,c} serves left-hand PREFIXES only."""
    index = ["dept", "year", "cgpa"]
    cases = [
        (["dept"],                 True,  "a prefix"),
        (["dept", "year"],         True,  "a prefix"),
        (["dept", "year", "cgpa"], True,  "the whole index"),
        (["year"],                 False, "not a prefix -- COLLSCAN"),
        (["cgpa"],                 False, "not a prefix -- COLLSCAN"),
        (["year", "cgpa"],         False, "not a prefix -- COLLSCAN"),
    ]

    def is_prefix(fields):
        return index[:len(fields)] == fields

    print(f"  index {{{', '.join(index)}}}:")
    for fields, expected, why in cases:
        assert is_prefix(fields) == expected, (fields, expected)
        mark = "uses it " if expected else "does NOT"
        print(f"    query on {str(fields):28s} {mark}  ({why})")

    print("       the phone book, sorted by (surname, forename): finding every")
    print("       Kumari is fast; finding every Asha means reading the whole book")


def the_esr_rule():
    """Equality, Sort, Range."""
    query = {"equality": "dept", "sort": "age", "range": "marks.maths"}
    correct = ["dept", "age", "marks.maths"]
    wrong = ["dept", "marks.maths", "age"]

    assert correct.index("age") < correct.index("marks.maths"), \
        "the SORT field must come before the RANGE field"
    assert wrong.index("marks.maths") < wrong.index("age"), \
        "this ordering puts the range first, and the sort cannot use the index"

    print(f"  ESR: query is dept = 'DS', maths > 70, sorted by age")
    print(f"    correct: {{{', '.join(correct)}}}   E, S, R")
    print(f"    wrong:   {{{', '.join(wrong)}}}   the range leaves everything")
    print(f"             after it unordered, so the sort falls back to memory")


def covered_query_fields():
    """Every field in the filter AND the projection must be in the index."""
    index = {"dept", "name"}
    filt = {"dept"}
    proj_bad = {"dept", "name", "_id"}      # _id is returned BY DEFAULT
    proj_good = {"dept", "name"}            # with _id: 0

    assert not (filt | proj_bad) <= index, "_id is not in the index -- NOT covered"
    assert (filt | proj_good) <= index, "with _id: 0 it IS covered"

    print("  covered query needs filter + projection inside the index")
    print("       find({dept}, {dept:1, name:1})        NOT covered -- _id sneaks in")
    print("       find({dept}, {dept:1, name:1, _id:0}) COVERED, totalDocsExamined 0")


def what_indexes_cost():
    db = fresh_db()
    for field in ["dept", "age", "active", "name"]:
        db.students.create_index(field)
    n = len(list(db.students.list_indexes()))
    assert n == 5, f"4 plus _id, got {n}"

    print(f"  {n} indexes means every insert, update and delete maintains {n}")
    print(f"       B-trees. Index what you QUERY, not everything -- the limit")
    print(f"       is 64 per collection, and reaching it means something is wrong")


def main():
    print("Experiment 13 -- Indexes")
    # Step 1: Create and list indexes
    creating_and_listing()
    # Step 2: See a unique index enforced
    unique_is_enforced()
    # Step 3: See two missing fields collide
    unique_and_missing_fields()
    # Step 4: Apply the prefix rule
    the_prefix_rule()
    # Step 5: Apply the ESR rule
    the_esr_rule()
    # Step 6: Cover a query
    covered_query_fields()
    # Step 7: Count what indexes cost
    what_indexes_cost()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 13_indexes.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.createIndex({ dept: 1 })
dept_1
collegeDB> db.students.createIndex({ dept: 1, "marks.maths": -1 })
dept_1_marks.maths_-1
collegeDB> db.students.createIndex({ email: 1 }, { unique: true })   // FAILS here: see below
Uncaught
MongoServerError[DuplicateKey]: Index build failed: 594da0ca-9e60-4781-be71-2b6fdab31bff: Collection collegeDB.students ( c4e0fefe-1c54-4b76-9c72-80d472ae564b ) :: caused by :: E11000 duplicate key error collection: collegeDB.students index: email_1 dup key: { email: null }
collegeDB> db.students.createIndex({ dept: 1 }, { name: "dept_idx" })   // FAILS: see below
Uncaught
MongoServerError[IndexOptionsConflict]: Index already exists with a different name: dept_1
collegeDB> db.students.getIndexes()
[
  { v: 2, key: { _id: 1 }, name: '_id_' },
  { v: 2, key: { dept: 1 }, name: 'dept_1' },
  {
    v: 2,
    key: { dept: 1, 'marks.maths': -1 },
    name: 'dept_1_marks.maths_-1'
  }
]
collegeDB> db.students.totalIndexSize()
45056
collegeDB> db.students.find({ dept: "DS" }).explain("executionStats")
{
  explainVersion: '1',
  queryPlanner: {
    namespace: 'collegeDB.students',
    parsedQuery: { dept: { '$eq': 'DS' } },
    indexFilterSet: false,
    queryHash: '8BDC9605',
    planCacheShapeHash: '8BDC9605',
    planCacheKey: 'F7C996A0',
    optimizationTimeMillis: 0,
    maxIndexedOrSolutionsReached: false,
    maxIndexedAndSolutionsReached: false,
    maxScansToExplodeReached: false,
    prunedSimilarIndexes: false,
    winningPlan: {
      isCached: false,
      stage: 'FETCH',
      nss: 'collegeDB.students',
      inputStage: {
        stage: 'IXSCAN',
        nss: 'collegeDB.students',
        keyPattern: { dept: 1 },
        indexName: 'dept_1',
        isMultiKey: false,
        multiKeyPaths: { dept: [] },
        isUnique: false,
        isSparse: false,
        isPartial: false,
        indexVersion: 2,
        direction: 'forward',
        indexBounds: { dept: [ '["DS", "DS"]' ] }
      }
    },
    rejectedPlans: [
      {
        isCached: false,
        stage: 'FETCH',
        nss: 'collegeDB.students',
        inputStage: {
          stage: 'IXSCAN',
          nss: 'collegeDB.students',
          keyPattern: { dept: 1, 'marks.maths': -1 },
          indexName: 'dept_1_marks.maths_-1',
          isMultiKey: false,
          multiKeyPaths: { dept: [], 'marks.maths': [] },
          isUnique: false,
          isSparse: false,
          isPartial: false,
          indexVersion: 2,
          direction: 'forward',
          indexBounds: {
            dept: [ '["DS", "DS"]' ],
            'marks.maths': [ '[MaxKey, MinKey]' ]
          }
        }
      }
    ]
  },
  executionStats: {
    executionSuccess: true,
    nReturned: 3,
    executionTimeMillis: 0,
    totalKeysExamined: 3,
    totalDocsExamined: 3,
    executionStages: {
      isCached: false,
      stage: 'FETCH',
      nReturned: 3,
      executionTimeMillisEstimate: 0,
      works: 5,
      advanced: 3,
      needTime: 0,
      needYield: 0,
      saveState: 1,
      restoreState: 1,
      isEOF: 1,
      nss: 'collegeDB.students',
      docsExamined: 3,
      alreadyHasObj: 0,
      inputStage: {
        stage: 'IXSCAN',
        nReturned: 3,
        executionTimeMillisEstimate: 0,
        works: 4,
        advanced: 3,
        needTime: 0,
        needYield: 0,
        saveState: 1,
        restoreState: 1,
        isEOF: 1,
        nss: 'collegeDB.students',
        keyPattern: { dept: 1 },
        indexName: 'dept_1',
        isMultiKey: false,
        multiKeyPaths: { dept: [] },
        isUnique: false,
        isSparse: false,
        isPartial: false,
        indexVersion: 2,
        direction: 'forward',
        indexBounds: { dept: [ '["DS", "DS"]' ] },
        keysExamined: 3,
        seeks: 1,
        dupsTested: 0,
        dupsDropped: 0,
        peakTrackedMemBytes: 0
      }
    }
  },
  queryShapeHash: '54C4AC01306DF576294D4F78A437455A0072E3A810EF2AF15CCC76DEB4F16CEE',
  command: { find: 'students', filter: { dept: 'DS' }, '$db': 'collegeDB' },
  serverInfo: {
    host: 'vm',
    port: 27017,
    version: '8.3.7',
    gitVersion: 'nogitversion'
  },
  serverParameters: {
    internalQueryFacetBufferSizeBytes: 104857600,
    internalDocumentSourceGroupMaxMemoryBytes: 104857600,
    internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
    internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
    internalQueryFacetMaxOutputDocSizeBytes: 104857600,
    internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
    internalQueryProhibitBlockingMergeOnMongoS: 0,
    internalQueryMaxAddToSetBytes: 104857600,
    internalQueryFrameworkControl: 'trySbeRestricted',
    internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
  },
  ok: 1
}
collegeDB> db.students.createIndex({ dept: 1, year: 1, cgpa: 1 })
dept_1_year_1_cgpa_1
collegeDB> db.students.find({ dept: "DS" })                        // uses it
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find({ dept: "DS", year: 4 })               // uses it
collegeDB> db.students.find({ year: 4 })                           // does NOT -- COLLSCAN
collegeDB> db.students.createIndex({ year: 1 })                    // so this is needed too
year_1
collegeDB> db.students.createIndex({ dept: 1, age: 1, "marks.maths": 1 })
dept_1_age_1_marks.maths_1
collegeDB> db.students.createIndex({ dept: 1, name: 1 })
dept_1_name_1
collegeDB> db.students.find({ dept: "DS" }, { _id: 0, dept: 1, name: 1 }).explain("executionStats")
{
  explainVersion: '1',
  queryPlanner: {
    namespace: 'collegeDB.students',
    parsedQuery: { dept: { '$eq': 'DS' } },
    indexFilterSet: false,
    queryHash: '1F033464',
    planCacheShapeHash: '1F033464',
    planCacheKey: '4D5148B1',
    optimizationTimeMillis: 1,
    maxIndexedOrSolutionsReached: false,
    maxIndexedAndSolutionsReached: false,
    maxScansToExplodeReached: false,
    prunedSimilarIndexes: false,
    winningPlan: {
      isCached: false,
      stage: 'PROJECTION_COVERED',
      transformBy: { _id: 0, dept: 1, name: 1 },
      inputStage: {
        stage: 'IXSCAN',
        nss: 'collegeDB.students',
        keyPattern: { dept: 1, name: 1 },
        indexName: 'dept_1_name_1',
        isMultiKey: false,
        multiKeyPaths: { dept: [], name: [] },
        isUnique: false,
        isSparse: false,
        isPartial: false,
        indexVersion: 2,
        direction: 'forward',
        indexBounds: { dept: [ '["DS", "DS"]' ], name: [ '[MinKey, MaxKey]' ] }
      }
    },
    rejectedPlans: [
      {
        isCached: false,
        stage: 'PROJECTION_SIMPLE',
        transformBy: { _id: 0, dept: 1, name: 1 },
        inputStage: {
          stage: 'FETCH',
          nss: 'collegeDB.students',
          inputStage: {
            stage: 'IXSCAN',
            nss: 'collegeDB.students',
            keyPattern: { dept: 1, 'marks.maths': -1 },
            indexName: 'dept_1_marks.maths_-1',
            isMultiKey: false,
            multiKeyPaths: { dept: [], 'marks.maths': [] },
            isUnique: false,
            isSparse: false,
            isPartial: false,
            indexVersion: 2,
            direction: 'forward',
            indexBounds: {
              dept: [ '["DS", "DS"]' ],
              'marks.maths': [ '[MaxKey, MinKey]' ]
            }
          }
        }
      },
      {
        isCached: false,
        stage: 'PROJECTION_SIMPLE',
        transformBy: { _id: 0, dept: 1, name: 1 },
        inputStage: {
          stage: 'FETCH',
          nss: 'collegeDB.students',
          inputStage: {
            stage: 'IXSCAN',
            nss: 'collegeDB.students',
            keyPattern: { dept: 1, year: 1, cgpa: 1 },
            indexName: 'dept_1_year_1_cgpa_1',
            isMultiKey: false,
            multiKeyPaths: { dept: [], year: [], cgpa: [] },
            isUnique: false,
            isSparse: false,
            isPartial: false,
            indexVersion: 2,
            direction: 'forward',
            indexBounds: {
              dept: [ '["DS", "DS"]' ],
              year: [ '[MinKey, MaxKey]' ],
              cgpa: [ '[MinKey, MaxKey]' ]
            }
          }
        }
      },
      {
        isCached: false,
        stage: 'PROJECTION_SIMPLE',
        transformBy: { _id: 0, dept: 1, name: 1 },
        inputStage: {
          stage: 'FETCH',
          nss: 'collegeDB.students',
          inputStage: {
            stage: 'IXSCAN',
            nss: 'collegeDB.students',
            keyPattern: { dept: 1, age: 1, 'marks.maths': 1 },
            indexName: 'dept_1_age_1_marks.maths_1',
            isMultiKey: false,
            multiKeyPaths: { dept: [], age: [], 'marks.maths': [] },
            isUnique: false,
            isSparse: false,
            isPartial: false,
            indexVersion: 2,
            direction: 'forward',
            indexBounds: {
              dept: [ '["DS", "DS"]' ],
              age: [ '[MinKey, MaxKey]' ],
              'marks.maths': [ '[MinKey, MaxKey]' ]
            }
          }
        }
      },
      {
        isCached: false,
        stage: 'PROJECTION_SIMPLE',
        transformBy: { _id: 0, dept: 1, name: 1 },
        inputStage: {
          stage: 'FETCH',
          nss: 'collegeDB.students',
          inputStage: {
            stage: 'IXSCAN',
            nss: 'collegeDB.students',
            keyPattern: { dept: 1 },
            indexName: 'dept_1',
            isMultiKey: false,
            multiKeyPaths: { dept: [] },
            isUnique: false,
            isSparse: false,
            isPartial: false,
            indexVersion: 2,
            direction: 'forward',
            indexBounds: { dept: [ '["DS", "DS"]' ] }
          }
        }
      }
    ]
  },
  executionStats: {
    executionSuccess: true,
    nReturned: 3,
    executionTimeMillis: 2,
    totalKeysExamined: 3,
    totalDocsExamined: 0,
    executionStages: {
      isCached: false,
      stage: 'PROJECTION_COVERED',
      nReturned: 3,
      executionTimeMillisEstimate: 0,
      works: 5,
      advanced: 3,
      needTime: 0,
      needYield: 0,
      saveState: 1,
      restoreState: 1,
      isEOF: 1,
      transformBy: { _id: 0, dept: 1, name: 1 },
      inputStage: {
        stage: 'IXSCAN',
        nReturned: 3,
        executionTimeMillisEstimate: 0,
        works: 5,
        advanced: 3,
        needTime: 0,
        needYield: 0,
        saveState: 1,
        restoreState: 1,
        isEOF: 1,
        nss: 'collegeDB.students',
        keyPattern: { dept: 1, name: 1 },
        indexName: 'dept_1_name_1',
        isMultiKey: false,
        multiKeyPaths: { dept: [], name: [] },
        isUnique: false,
        isSparse: false,
        isPartial: false,
        indexVersion: 2,
        direction: 'forward',
        indexBounds: { dept: [ '["DS", "DS"]' ], name: [ '[MinKey, MaxKey]' ] },
        keysExamined: 3,
        seeks: 1,
        dupsTested: 0,
        dupsDropped: 0,
        peakTrackedMemBytes: 0
      }
    }
  },
  queryShapeHash: '3702523BB00E9C352D45FA2B03292BA7E3DDDC04DB9C14A06C726E7CD9CF4AC9',
  command: {
    find: 'students',
    filter: { dept: 'DS' },
    projection: { _id: 0, dept: 1, name: 1 },
    '$db': 'collegeDB'
  },
  serverInfo: {
    host: 'vm',
    port: 27017,
    version: '8.3.7',
    gitVersion: 'nogitversion'
  },
  serverParameters: {
    internalQueryFacetBufferSizeBytes: 104857600,
    internalDocumentSourceGroupMaxMemoryBytes: 104857600,
    internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
    internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
    internalQueryFacetMaxOutputDocSizeBytes: 104857600,
    internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
    internalQueryProhibitBlockingMergeOnMongoS: 0,
    internalQueryMaxAddToSetBytes: 104857600,
    internalQueryFrameworkControl: 'trySbeRestricted',
    internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
  },
  ok: 1
}
collegeDB> db.people.drop()
true
collegeDB> db.people.createIndex({ email: 1 }, { unique: true })
email_1
collegeDB> db.people.insertOne({ name: "A" })         // ok -- email missing, indexed as null
{
  acknowledged: true,
  insertedId: ObjectId('6ac214e2f47d929cc4821b4c')
}
collegeDB> db.people.insertOne({ name: "B" })         // DUPLICATE KEY ERROR -- a second null
Uncaught
MongoServerError: E11000 duplicate key error collection: collegeDB.people index: email_1 dup key: { email: null }
collegeDB> db.people.dropIndex("email_1")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.people.createIndex({ email: 1 },
...   { unique: true, partialFilterExpression: { email: { $exists: true } } })
email_1
collegeDB> db.people.insertOne({ name: "B" })         // ok now: missing emails are not indexed
{
  acknowledged: true,
  insertedId: ObjectId('6ac214e3f47d929cc4821b4e')
}
collegeDB> db.students.dropIndex("dept_1")
{ nIndexesWas: 7, ok: 1 }

In Python, through mongomock, 13_indexes.py:

OUTPUT

Experiment 13 -- Indexes
  created and listed 4 indexes, including the automatic _id
  a unique index rejected a duplicate name and accepted a new one
  a unique index allowed ONE document with no email, then rejected
       the next -- a missing field indexes as null, and nulls collide
       fix: partialFilterExpression: { email: { $exists: true } }
  index {dept, year, cgpa}:
    query on ['dept']                     uses it   (a prefix)
    query on ['dept', 'year']             uses it   (a prefix)
    query on ['dept', 'year', 'cgpa']     uses it   (the whole index)
    query on ['year']                     does NOT  (not a prefix -- COLLSCAN)
    query on ['cgpa']                     does NOT  (not a prefix -- COLLSCAN)
    query on ['year', 'cgpa']             does NOT  (not a prefix -- COLLSCAN)
       the phone book, sorted by (surname, forename): finding every
       Kumari is fast; finding every Asha means reading the whole book
  ESR: query is dept = 'DS', maths > 70, sorted by age
    correct: {dept, age, marks.maths}   E, S, R
    wrong:   {dept, marks.maths, age}   the range leaves everything
             after it unordered, so the sort falls back to memory
  covered query needs filter + projection inside the index
       find({dept}, {dept:1, name:1})        NOT covered -- _id sneaks in
       find({dept}, {dept:1, name:1, _id:0}) COVERED, totalDocsExamined 0
  5 indexes means every insert, update and delete maintains 5
       B-trees. Index what you QUERY, not everything -- the limit
       is 64 per collection, and reaching it means something is wrong

The covered query's plan is PROJECTION_COVERED with totalDocsExamined: 0, as its comment says.

Corrected or noted, from running it: the unique index on email and the dept_idx index fail at the top, and the file now says why — none of the five students has an email, so all five index as null, and { dept: 1 } already exists as dept_1. The demonstration that two missing fields collide used db.students, where the unique index could not be built at all, so the insert commented DUPLICATE KEY ERROR succeeded; it now uses a new collection, as the Python half does. And .explain(...) began its own line, which mongosh does not join to the line above.

RESULT

find({ dept: "DS" }) is an IXSCAN examining 3 documents for 3 returned; the covered query examines none.

Experiment 14 — Text search and multikey indexes

1. Question

Search text with a text index, and index an array with a multikey index.

2. Aim

Index arrays and text, and search by words with relevance.

3. Steps

In mongosh, 14_text_multikey.js:

  1. Load the sample data.
  2. Index an array, which makes it multikey.
  3. Meet the restrictions.
  4. Index an array of sub-documents, and fall into the trap.
  5. Create a text index.
  6. Search.
  7. Rank by relevance.

In Python, through mongomock, 14_text_multikey.py:

  1. See an index on an array become multikey.
  2. Count one entry per element.
  3. Use $all, $size and the positional operator.
  4. Store the length to query it.
  5. Fall into the $elemMatch trap.
  6. State the compound restriction.
  7. Note that mongomock has no $text.
  8. Work out what a real server returns.
  9. Search by regex instead.
  10. State the text index rules.

THE POINT

An index on an array field creates one entry per element, so { subjects: "DS" } matches any student whose array contains it. Only one text index is allowed per collection — it may span several fields, but you cannot have two.

mongomock does not implement $text, so the Python half asserts the multikey behaviour and works out by hand what a server's text search returns; the session shows the server's.

4. Programme

In mongosh, 14_text_multikey.js:

// Experiment 14 -- Text search and multikey indexes.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 14_text_multikey.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// =============================================================================
// Step 2: Index an array, which makes it multikey
// PART A -- MULTIKEY INDEXES (an index on an array field)
// =============================================================================
// There is no "createMultikeyIndex". You index the field, and MongoDB makes
// the index multikey BY ITSELF the moment it meets an array value.

db.students.createIndex({ subjects: 1 })

db.students.find({ subjects: "DS" })          // matches if the ARRAY CONTAINS it
db.students.find({ subjects: { $all: ["DS", "Python"] } })   // contains BOTH
db.students.find({ subjects: { $size: 3 } })  // exactly three -- NOT indexed
db.students.find({ "subjects.0": "DS" })      // DS is the FIRST element

// One index ENTRY per array element. A student with 3 subjects contributes 3
// entries pointing at the same document, which is why multikey indexes are
// larger than they look, and why an array of 1,000 elements is a bad idea.

// Step 3: Meet the restrictions
// 1. A compound index may contain AT MOST ONE array field.
db.students.createIndex({ subjects: 1, dept: 1 })      // OK -- one array
// db.students.createIndex({ subjects: 1, tags: 1 })   // ERROR if BOTH arrays
// 2. A multikey index cannot be a shard key.
// 3. $size is never served by an index -- it must scan. Store a length field
//    alongside the array if you need to query on it:
db.students.updateMany({}, [ { $set: { nSubjects: { $size: "$subjects" } } } ])
db.students.createIndex({ nSubjects: 1 })

// Step 4: Index an array of sub-documents, and fall into the trap
db.transcripts.drop()
db.transcripts.insertOne({ _id: 21, name: "Asha", enrollments: [
  { course: "DSC301", grade: "B" }, { course: "STA302", grade: "A" } ] })
db.transcripts.createIndex({ "enrollments.grade": 1 })    // also multikey

// The trap from experiment 9, restated: without $elemMatch the two conditions
// may be satisfied by DIFFERENT elements of the array.
db.transcripts.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" })   // Asha -- wrongly
db.transcripts.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })   // nobody
// [Corrected: these queried db.students, whose documents have no enrollments,
// so both found nothing and the trap did not show. Asha's B in DSC301 and A in
// STA302 are two different elements, and only $elemMatch tells them apart.]

// =============================================================================
// Step 5: Create a text index
// PART B -- TEXT INDEXES
// =============================================================================
db.articles.drop()
db.articles.insertMany([
  { _id: 1, title: "Introduction to MongoDB",
    body: "MongoDB is a document database that stores data in BSON." },
  { _id: 2, title: "Aggregation pipelines explained",
    body: "The aggregation framework processes documents through stages." },
  { _id: 3, title: "Indexing strategy in MongoDB",
    body: "An index is a B-tree. Aggregation queries benefit from indexes too." },
  { _id: 4, title: "Relational databases",
    body: "SQL databases use tables, rows and joins." }
])

// Weights make a hit in the title count ten times a hit in the body.
db.articles.createIndex({ title: "text", body: "text" },
                        { weights: { title: 10, body: 1 },
                          name: "article_text",
                          default_language: "english" })

// Step 6: Search
db.articles.find({ $text: { $search: "mongodb" } })
db.articles.find({ $text: { $search: "mongodb aggregation" } })   // OR, not AND
db.articles.find({ $text: { $search: "\"aggregation framework\"" } })  // PHRASE
db.articles.find({ $text: { $search: "mongodb -relational" } })    // EXCLUDE

// Step 7: Rank by relevance
db.articles.find({ $text: { $search: "mongodb aggregation" } },
                 { score: { $meta: "textScore" }, title: 1 }).sort({ score: { $meta: "textScore" } })
// [Corrected: the .sort(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find(), unsorted, without it, then rejected ".sort(...)" as an invalid command.]

// The sort is NOT optional. $text returns matches in no particular order; the
// score exists only if you project it, and only sorts if you sort by it.

// --- the rules, all examinable ----------------------------------------------
// 1. ONE text index per collection. It may span many fields -- even every
//    field, via { "$**": "text" } -- but you cannot have two.
db.articles.dropIndex("article_text")
db.articles.createIndex({ "$**": "text" })     // a WILDCARD text index
db.articles.dropIndex("$**_text")
db.articles.createIndex({ title: "text", body: "text" },
                        { weights: { title: 10, body: 1 }, name: "article_text" })
// 2. $text searches WORDS, not substrings. "mongo" does not match "MongoDB".
//    For substrings and prefixes you need a regex, or Atlas Search.
// 3. Search is case-insensitive and diacritic-insensitive by default.
// 4. Stemming and stop words follow default_language: searching "stores" also
//    matches "store" and "storing"; "the" and "is" are ignored entirely.
// 5. Only ONE $text expression per query, and it cannot appear inside $or with
//    a non-text clause.

db.articles.getIndexes()

In Python, through mongomock, 14_text_multikey.py:

"""Experiment 14 — Text search and multikey indexes.

Two halves, and they run differently.

MULTIKEY is fully executed: mongomock matches array fields the way a real
server does, so every assertion here is a real one.

TEXT SEARCH is not. mongomock accepts createIndex([("x", "text")]) but raises
NotImplementedError on $text -- that is asserted below rather than glossed
over, and the ranking and stemming rules are then set out as data. The runnable
substitute is a regex scan, which is also the honest answer to "what do I do
when I have no text index?".
"""
import mongomock
from fixtures import fresh_db, names


# =============================================================================
# PART A -- multikey indexes: fully executed
# =============================================================================

def an_index_on_an_array_is_multikey_automatically():
    """You never ask for a multikey index. MongoDB decides."""
    db = fresh_db()
    db.students.create_index("subjects")

    # The index exists and looks like any other single-field index.
    assert "subjects_1" in {i["name"] for i in db.students.list_indexes()}

    # Matching is CONTAINS, not equals: no student's subjects field IS "DS".
    assert names(db.students.find({"subjects": "DS"})) == ["Asha", "Kiran", "Ravi"]
    assert names(db.students.find({"subjects": "R"})) == ["Meena"]

    # And the whole array still matches as a whole, if you give it exactly.
    assert names(db.students.find({"subjects": ["DS", "Python"]})) == ["Ravi"]
    assert names(db.students.find({"subjects": ["Python", "DS"]})) == [], \
        "the whole-array form is ORDER SENSITIVE; the contains form is not"

    print("  { subjects: 'DS' } matched Asha, Kiran, Ravi -- CONTAINS, not equals")
    print("  { subjects: ['Python','DS'] } matched nobody -- whole-array match")
    print("       is order sensitive, and Ravi's array is ['DS','Python']")


def one_index_entry_per_element():
    """The cost model: a document with n array values costs n index entries."""
    db = fresh_db()
    entries = sum(len(d["subjects"]) for d in db.students.find())
    docs = db.students.count_documents({})
    assert docs == 5
    assert entries == 9, entries          # 3 + 2 + 2 + 1 + 1

    print(f"  {docs} documents -> {entries} index entries "
          f"({entries / docs:.1f} per document)")
    print("       an array of 1,000 elements means 1,000 entries for ONE")
    print("       document -- multikey indexes are bigger than they look")


def all_and_size_and_positional():
    db = fresh_db()

    # $all -- contains ALL of these (order irrelevant)
    assert names(db.students.find({"subjects": {"$all": ["DS", "Python"]}})) \
        == ["Asha", "Ravi"]

    # bare {a: x, ...} on an array is OR-ish across elements; $all is AND
    assert names(db.students.find({"subjects": {"$all": ["Stats", "R"]}})) == ["Meena"]

    # $size -- exact length. NEVER served by an index.
    assert names(db.students.find({"subjects": {"$size": 3}})) == ["Asha"]
    assert names(db.students.find({"subjects": {"$size": 1}})) == ["Bhanu", "Kiran"]

    # positional: the FIRST element specifically
    assert names(db.students.find({"subjects.0": "DS"})) == ["Asha", "Kiran", "Ravi"]
    assert names(db.students.find({"subjects.0": "Stats"})) == ["Bhanu", "Meena"]

    print("  $all ['DS','Python'] -> Asha, Ravi        (contains BOTH)")
    print("  $size 3              -> Asha              (never uses the index)")
    print("  subjects.0 'Stats'   -> Bhanu, Meena      (FIRST element only)")


def store_the_length_if_you_query_it():
    """The fix for $size's collection scan: a field you CAN index."""
    db = fresh_db()
    for d in db.students.find():
        db.students.update_one({"_id": d["_id"]},
                               {"$set": {"nSubjects": len(d["subjects"])}})
    db.students.create_index("nSubjects")

    assert names(db.students.find({"nSubjects": 3})) == ["Asha"]
    assert names(db.students.find({"nSubjects": {"$gte": 2}})) \
        == ["Asha", "Meena", "Ravi"]

    print("  nSubjects >= 2 -> Asha, Meena, Ravi -- and unlike $size this one is")
    print("       indexable, and supports RANGES, which $size cannot express")


def the_elemmatch_trap_again():
    """Two conditions on an array of sub-documents. The classic wrong answer."""
    db = fresh_db()
    db.students.update_one(
        {"_id": 21},
        {"$set": {"enrollments": [{"course": "DSC301", "grade": "B"},
                                  {"course": "STA302", "grade": "A"}]}})
    db.students.create_index("enrollments.grade")

    # WRONG: satisfied by two DIFFERENT elements -- Asha got B in DSC301.
    wrong = names(db.students.find({"enrollments.course": "DSC301",
                                   "enrollments.grade": "A"}))
    assert wrong == ["Asha"], wrong

    # RIGHT: both conditions on the SAME element.
    right = names(db.students.find(
        {"enrollments": {"$elemMatch": {"course": "DSC301", "grade": "A"}}}))
    assert right == [], right

    print("  Asha: DSC301->B, STA302->A")
    print("    without $elemMatch -> ['Asha']   WRONG (two different elements)")
    print("    with    $elemMatch -> []         RIGHT")


def the_compound_restriction():
    """At most ONE array field in a compound index. Stated, not executed."""
    rules = [
        ("{ subjects: 1 }",            "OK", "one array field"),
        ("{ subjects: 1, dept: 1 }",   "OK", "one array, one scalar"),
        ("{ dept: 1, subjects: 1 }",   "OK", "order does not change the rule"),
        ("{ subjects: 1, tags: 1 }",   "ERROR",
         "two array fields -- cannot compute the cross product"),
    ]
    assert sum(1 for _, v, _ in rules if v == "ERROR") == 1

    print("  compound indexes containing arrays:")
    for spec, verdict, why in rules:
        print(f"    {spec:28s} {verdict:6s} {why}")
    print("       the reason: indexing both would need EVERY pair, so a")
    print("       document with 10 and 10 would need 100 index entries")


# =============================================================================
# PART B -- text search: mongomock cannot run it, so say so and prove it
# =============================================================================

ARTICLES = [
    {"_id": 1, "title": "Introduction to MongoDB",
     "body": "MongoDB is a document database that stores data in BSON."},
    {"_id": 2, "title": "Aggregation pipelines explained",
     "body": "The aggregation framework processes documents through stages."},
    {"_id": 3, "title": "Indexing strategy in MongoDB",
     "body": "An index is a B-tree. Aggregation queries benefit from indexes too."},
    {"_id": 4, "title": "Relational databases",
     "body": "SQL databases use tables, rows and joins."},
]


def text_search_is_not_implemented_here():
    """Asserted, so this file can never quietly start claiming to test $text."""
    db = mongomock.MongoClient().collegeDB
    db.articles.insert_many([dict(a) for a in ARTICLES])
    db.articles.create_index([("title", "text"), ("body", "text")])

    try:
        list(db.articles.find({"$text": {"$search": "mongodb"}}))
        raise SystemExit("mongomock now implements $text -- rewrite this file "
                         "to assert results instead of documenting them")
    except NotImplementedError as exc:
        message = str(exc)
    assert "$text" in message, message

    print("  mongomock accepted the text INDEX and then raised")
    print(f"    NotImplementedError: {message}")
    print("       so nothing below is a test result -- it is documentation")


def what_a_real_server_would_return():
    """The expected results, worked out by hand from the four articles."""
    expected = [
        ('"mongodb"',                  [1, 3],       "the word, in title or body"),
        ('"mongodb aggregation"',      [1, 2, 3],    "OR of the two terms, NOT and"),
        ('"\\"aggregation framework\\""', [2],       "a PHRASE -- adjacent words"),
        ('"mongodb -relational"',      [1, 3],       "- excludes; 4 never matched anyway"),
        ('"mongo"',                    [],           "WORDS, not substrings"),
        ('"stores"',                   [1],          "stemming: matches 'stores'/'storing'"),
        ('"the"',                      [],           "a stop word -- ignored entirely"),
    ]
    print("  on a real server, $text: { $search: ... } would return:")
    for term, ids, why in expected:
        print(f"    {term:32s} -> {str(ids):10s} {why}")
    print("       'mongo' returning NOTHING is the one that catches people:")
    print("       a text index stores WORDS. For prefixes use a regex anchored")
    print("       with ^, or Atlas Search; a bare /mongo/ scans the collection")


def regex_is_the_runnable_substitute():
    """What you actually do without a text index -- and it IS executed."""
    db = mongomock.MongoClient().collegeDB
    db.articles.insert_many([dict(a) for a in ARTICLES])

    hits = sorted(d["_id"] for d in
                  db.articles.find({"title": {"$regex": "mongo", "$options": "i"}}))
    assert hits == [1, 3], hits

    # And here is where regex BEATS $text: substrings.
    sub = sorted(d["_id"] for d in
                 db.articles.find({"body": {"$regex": "aggregat", "$options": "i"}}))
    assert sub == [2, 3], sub

    # And where it loses: no stemming, no ranking.
    stem = sorted(d["_id"] for d in
                  db.articles.find({"body": {"$regex": "store$", "$options": "i"}}))
    assert stem == [], "regex has no stemming -- 'stores' does not match 'store$'"

    print("  regex /mongo/i on title  -> [1, 3]   (substring: $text would miss)")
    print("  regex /aggregat/i on body-> [2, 3]   (substring again)")
    print("  regex /store$/i on body  -> []       (no stemming, no ranking,")
    print("       and an unanchored regex cannot use an index -- COLLSCAN)")


def the_text_index_rules():
    rules = [
        ("How many per collection?", "ONE",
         "it may span many fields, but you cannot have two"),
        ("Every field?", '{ "$**": "text" }',
         "a wildcard text index -- convenient, and large"),
        ("Weights", "{ title: 10, body: 1 }",
         "a title hit scores ten times a body hit"),
        ("Getting the score", '{ $meta: "textScore" }',
         "must be PROJECTED to exist"),
        ("Ordering by it", '.sort({ score: { $meta: "textScore" } })',
         "$text does NOT sort by relevance on its own"),
        ("Case / accents", "insensitive by default", "both, unless configured"),
        ("Inside $or", "not with a non-text clause", "one $text per query"),
    ]
    assert len(rules) == 7
    print("  text index rules:")
    for q, a, why in rules:
        print(f"    {q:24s} {a:34s} {why}")


def main():
    print("Experiment 14 -- Text search and multikey indexes")
    print("  PART A -- multikey: EXECUTED and asserted")
    # Step 1: See an index on an array become multikey
    an_index_on_an_array_is_multikey_automatically()
    # Step 2: Count one entry per element
    one_index_entry_per_element()
    # Step 3: Use $all, $size and the positional operator
    all_and_size_and_positional()
    # Step 4: Store the length to query it
    store_the_length_if_you_query_it()
    # Step 5: Fall into the $elemMatch trap
    the_elemmatch_trap_again()
    # Step 6: State the compound restriction
    the_compound_restriction()
    print("  PART B -- text search: NOT executable here")
    # Step 7: Note that mongomock has no $text
    text_search_is_not_implemented_here()
    # Step 8: Work out what a real server returns
    what_a_real_server_would_return()
    # Step 9: Search by regex instead
    regex_is_the_runnable_substitute()
    # Step 10: State the text index rules
    the_text_index_rules()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 14_text_multikey.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.createIndex({ subjects: 1 })
subjects_1
collegeDB> db.students.find({ subjects: "DS" })          // matches if the ARRAY CONTAINS it
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.find({ subjects: { $all: ["DS", "Python"] } })   // contains BOTH
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  }
]
collegeDB> db.students.find({ subjects: { $size: 3 } })  // exactly three -- NOT indexed
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  }
]
collegeDB> db.students.find({ "subjects.0": "DS" })      // DS is the FIRST element
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: [ 'DS', 'Python' ],
    age: 21,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false
  }
]
collegeDB> db.students.createIndex({ subjects: 1, dept: 1 })      // OK -- one array
subjects_1_dept_1
collegeDB> db.students.updateMany({}, [ { $set: { nSubjects: { $size: "$subjects" } } } ])
{
  acknowledged: true,
  insertedId: null,
  matchedCount: 5,
  modifiedCount: 5,
  upsertedCount: 0
}
collegeDB> db.students.createIndex({ nSubjects: 1 })
nSubjects_1
collegeDB> db.transcripts.drop()
true
collegeDB> db.transcripts.insertOne({ _id: 21, name: "Asha", enrollments: [
...   { course: "DSC301", grade: "B" }, { course: "STA302", grade: "A" } ] })
{ acknowledged: true, insertedId: 21 }
collegeDB> db.transcripts.createIndex({ "enrollments.grade": 1 })    // also multikey
enrollments.grade_1
collegeDB> db.transcripts.find({ "enrollments.course": "DSC301", "enrollments.grade": "A" })   // Asha -- wrongly
[
  {
    _id: 21,
    name: 'Asha',
    enrollments: [ { course: 'DSC301', grade: 'B' }, { course: 'STA302', grade: 'A' } ]
  }
]
collegeDB> db.transcripts.find({ enrollments: { $elemMatch: { course: "DSC301", grade: "A" } } })   // nobody
collegeDB> db.articles.drop()
true
collegeDB> db.articles.insertMany([
...   { _id: 1, title: "Introduction to MongoDB",
...     body: "MongoDB is a document database that stores data in BSON." },
...   { _id: 2, title: "Aggregation pipelines explained",
...     body: "The aggregation framework processes documents through stages." },
...   { _id: 3, title: "Indexing strategy in MongoDB",
...     body: "An index is a B-tree. Aggregation queries benefit from indexes too." },
...   { _id: 4, title: "Relational databases",
...     body: "SQL databases use tables, rows and joins." }
... ])
{ acknowledged: true, insertedIds: { '0': 1, '1': 2, '2': 3, '3': 4 } }
collegeDB> db.articles.createIndex({ title: "text", body: "text" },
...                         { weights: { title: 10, body: 1 },
...                           name: "article_text",
...                           default_language: "english" })
article_text
collegeDB> db.articles.find({ $text: { $search: "mongodb" } })
[
  {
    _id: 1,
    title: 'Introduction to MongoDB',
    body: 'MongoDB is a document database that stores data in BSON.'
  },
  {
    _id: 3,
    title: 'Indexing strategy in MongoDB',
    body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
  }
]
collegeDB> db.articles.find({ $text: { $search: "mongodb aggregation" } })   // OR, not AND
[
  {
    _id: 2,
    title: 'Aggregation pipelines explained',
    body: 'The aggregation framework processes documents through stages.'
  },
  {
    _id: 3,
    title: 'Indexing strategy in MongoDB',
    body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
  },
  {
    _id: 1,
    title: 'Introduction to MongoDB',
    body: 'MongoDB is a document database that stores data in BSON.'
  }
]
collegeDB> db.articles.find({ $text: { $search: "\"aggregation framework\"" } })  // PHRASE
[
  {
    _id: 2,
    title: 'Aggregation pipelines explained',
    body: 'The aggregation framework processes documents through stages.'
  }
]
collegeDB> db.articles.find({ $text: { $search: "mongodb -relational" } })    // EXCLUDE
[
  {
    _id: 1,
    title: 'Introduction to MongoDB',
    body: 'MongoDB is a document database that stores data in BSON.'
  },
  {
    _id: 3,
    title: 'Indexing strategy in MongoDB',
    body: 'An index is a B-tree. Aggregation queries benefit from indexes too.'
  }
]
collegeDB> db.articles.find({ $text: { $search: "mongodb aggregation" } },
...                  { score: { $meta: "textScore" }, title: 1 }).sort({ score: { $meta: "textScore" } })
[
  { _id: 1, title: 'Introduction to MongoDB', score: 8.083333333333334 },
  {
    _id: 2,
    title: 'Aggregation pipelines explained',
    score: 7.266666666666666
  },
  {
    _id: 3,
    title: 'Indexing strategy in MongoDB',
    score: 7.238095238095237
  }
]
collegeDB> db.articles.dropIndex("article_text")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.articles.createIndex({ "$**": "text" })     // a WILDCARD text index
$**_text
collegeDB> db.articles.dropIndex("$**_text")
{ nIndexesWas: 2, ok: 1 }
collegeDB> db.articles.createIndex({ title: "text", body: "text" },
...                         { weights: { title: 10, body: 1 }, name: "article_text" })
article_text
collegeDB> db.articles.getIndexes()
[
  { v: 2, key: { _id: 1 }, name: '_id_' },
  {
    v: 2,
    key: { _fts: 'text', _ftsx: 1 },
    name: 'article_text',
    weights: { body: 1, title: 10 },
    default_language: 'english',
    language_override: 'language',
    textIndexVersion: 3
  }
]

In Python, through mongomock, 14_text_multikey.py:

OUTPUT

Experiment 14 -- Text search and multikey indexes
  PART A -- multikey: EXECUTED and asserted
  { subjects: 'DS' } matched Asha, Kiran, Ravi -- CONTAINS, not equals
  { subjects: ['Python','DS'] } matched nobody -- whole-array match
       is order sensitive, and Ravi's array is ['DS','Python']
  5 documents -> 9 index entries (1.8 per document)
       an array of 1,000 elements means 1,000 entries for ONE
       document -- multikey indexes are bigger than they look
  $all ['DS','Python'] -> Asha, Ravi        (contains BOTH)
  $size 3              -> Asha              (never uses the index)
  subjects.0 'Stats'   -> Bhanu, Meena      (FIRST element only)
  nSubjects >= 2 -> Asha, Meena, Ravi -- and unlike $size this one is
       indexable, and supports RANGES, which $size cannot express
  Asha: DSC301->B, STA302->A
    without $elemMatch -> ['Asha']   WRONG (two different elements)
    with    $elemMatch -> []         RIGHT
  compound indexes containing arrays:
    { subjects: 1 }              OK     one array field
    { subjects: 1, dept: 1 }     OK     one array, one scalar
    { dept: 1, subjects: 1 }     OK     order does not change the rule
    { subjects: 1, tags: 1 }     ERROR  two array fields -- cannot compute the cross product
       the reason: indexing both would need EVERY pair, so a
       document with 10 and 10 would need 100 index entries
  PART B -- text search: NOT executable here
  mongomock accepted the text INDEX and then raised
    NotImplementedError: The $text operator is not implemented in mongomock yet
       so nothing below is a test result -- it is documentation
  on a real server, $text: { $search: ... } would return:
    "mongodb"                        -> [1, 3]     the word, in title or body
    "mongodb aggregation"            -> [1, 2, 3]  OR of the two terms, NOT and
    "\"aggregation framework\""      -> [2]        a PHRASE -- adjacent words
    "mongodb -relational"            -> [1, 3]     - excludes; 4 never matched anyway
    "mongo"                          -> []         WORDS, not substrings
    "stores"                         -> [1]        stemming: matches 'stores'/'storing'
    "the"                            -> []         a stop word -- ignored entirely
       'mongo' returning NOTHING is the one that catches people:
       a text index stores WORDS. For prefixes use a regex anchored
       with ^, or Atlas Search; a bare /mongo/ scans the collection
  regex /mongo/i on title  -> [1, 3]   (substring: $text would miss)
  regex /aggregat/i on body-> [2, 3]   (substring again)
  regex /store$/i on body  -> []       (no stemming, no ranking,
       and an unanchored regex cannot use an index -- COLLSCAN)
  text index rules:
    How many per collection? ONE                                it may span many fields, but you cannot have two
    Every field?             { "$**": "text" }                  a wildcard text index -- convenient, and large
    Weights                  { title: 10, body: 1 }             a title hit scores ten times a body hit
    Getting the score        { $meta: "textScore" }             must be PROJECTED to exist
    Ordering by it           .sort({ score: { $meta: "textScore" } }) $text does NOT sort by relevance on its own
    Case / accents           insensitive by default             both, unless configured
    Inside $or               not with a non-text clause         one $text per query

Corrected: the $elemMatch trap queried db.students, whose documents have no enrolments, so both queries found nothing; it now has a document to find, in which Asha has a B in DSC301 and an A in STA302. And the ranking's .sort(...) began its own line, so the search ran unsorted and the shell rejected the sort.

RESULT

The array index is multikey; $text matches whole words and ranks by score; without $elemMatch, Asha is found for an A in DSC301 she does not have.

Experiment 15 — $match, $group, $project, $sort

1. Question

Aggregate documents with $match, $group, $project and $sort.

2. Aim

Build an aggregation pipeline, and see each stage's SQL counterpart.

3. Steps

In mongosh, 15_aggregation.js:

  1. Load the sample data.
  2. $match, $group, $sort: the whole pipeline.
  3. $group without a filter.
  4. The accumulators.
  5. $project: include, exclude, compute, rename.
  6. Put $match first.
  7. See what each stage does.
  8. Send the result to a collection.

In Python, through mongomock, 15_aggregation.py:

  1. $match then $group: WHERE and HAVING.
  2. See the filter change the answer.
  3. Group by a key, or by null for a total.
  4. Use the accumulators.
  5. Note the operators mongomock lacks.
  6. $project: compute and rename.
  7. $addFields: keep everything.
  8. $match first.

THE POINT

Shown twice, and the pair is the point. With the active: true filter, Kiran (DS, 71, inactive) is excluded, so DS is (88+65)/2 = 76.5 over 2 and Stats (94+52)/2 = 73 over 2. Without it — Unit 4's Problem 1(a) — DS is (88+65+71)/3 = 74.667 over 3. Same grouping, and the $match before it moves the DS average up.

$match before and after $group are WHERE and HAVING — the same stage in different positions. mongomock lacks $round and $stdDevPop, so the Python half computes those in Python and says so; the server has both.

4. Programme

In mongosh, 15_aggregation.js:

// Experiment 15 -- The aggregation pipeline: $match, $group, $project, $sort.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 15_aggregation.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// Step 2: $match, $group, $sort: the whole pipeline
//   SELECT   dept, ROUND(AVG(maths),2) AS avg, COUNT(*) AS n
//   FROM     students
//   WHERE    active = true          -- $match BEFORE $group
//   GROUP BY dept
//   HAVING   COUNT(*) > 1           -- $match AFTER  $group
//   ORDER BY avg DESC;
db.students.aggregate([
  { $match:   { active: true } },
  { $group:   { _id: "$dept", avg: { $avg: "$marks.maths" },
                n: { $sum: 1 } } },
  { $match:   { n: { $gt: 1 } } },
  { $sort:    { avg: -1 } },
  { $project: { _id: 0, dept: "$_id", avg: { $round: ["$avg", 2] }, n: 1 } }
])

// WHERE and HAVING are THE SAME STAGE in different positions. Say that in the
// viva; it is the sentence that shows you understand pipelines.

// Step 3: $group without a filter
db.students.aggregate([
  { $group: { _id: "$dept", avgMaths: { $avg: "$marks.maths" },
              n: { $sum: 1 } } },
  { $sort: { _id: 1 } }
])
// _id is MANDATORY in $group. It is the grouping key. And $group returns its
// groups in NO fixed order -- sort them whenever the order is shown.
// [Changed: the $sort was added; the groups came back in different orders.]

db.students.aggregate([ { $group: { _id: null, avg: { $avg: "$marks.maths" },
                                    n: { $sum: 1 } } } ])
// _id: null groups EVERYTHING into one bucket -- a grand total.

// Step 4: The accumulators
db.students.aggregate([
  { $group: {
      _id: "$dept",
      n:        { $sum: 1 },                    // COUNT(*)
      totMaths: { $sum: "$marks.maths" },       // SUM
      avgMaths: { $avg: "$marks.maths" },       // AVG
      best:     { $max: "$marks.maths" },       // MAX
      worst:    { $min: "$marks.maths" },       // MIN
      sd:       { $stdDevPop: "$marks.maths" }, // Course 4's population sd
      everyone: { $push: "$name" },             // ALL values, as an array
      distinct: { $addToSet: "$name" },         // DISTINCT values
      anyone:   { $first: "$name" }             // needs a $sort to be meaningful
  } },
  { $set: { distinct: { $sortArray: { input: "$distinct", sortBy: 1 } } } },
  { $sort: { _id: 1 } }
])
// $addToSet keeps NO order, so the $set sorts that array before it is shown.
// [Changed: the $set and the $sort were added. The distinct names came back
// in a different order from one run to the next.]
// $push and $addToSet have no SQL equivalent, and are the reason MongoDB does
// not need GROUP_CONCAT.

// Step 5: $project: include, exclude, compute, rename
db.students.aggregate([
  { $project: {
      _id: 0,
      name: 1,                                       // include
      total: { $add: ["$marks.maths", "$marks.stats"] },
      pct:   { $round: [ { $divide: [ { $add: ["$marks.maths", "$marks.stats"] },
                                      2 ] }, 1 ] },
      dept:  "$dept",                                // rename by re-assigning
      band:  { $switch: { branches: [
                 { case: { $gte: ["$marks.maths", 75] }, then: "Distinction" },
                 { case: { $gte: ["$marks.maths", 60] }, then: "First" },
                 { case: { $gte: ["$marks.maths", 40] }, then: "Pass" } ],
               default: "Fail" } }
  } }
])

// $addFields (alias: $set) keeps everything and adds -- usually what you meant.
db.students.aggregate([
  { $addFields: { total: { $add: ["$marks.maths", "$marks.stats"] } } },
  { $sort: { total: -1 } },
  { $limit: 3 }
])

// Step 6: Put $match first
db.students.aggregate([ { $group: { _id: "$dept", n: { $sum: 1 } } },
                        { $match: { _id: "DS" } } ])       // SLOW: groups all
db.students.aggregate([ { $match: { dept: "DS" } },
                        { $group: { _id: "$dept", n: { $sum: 1 } } } ])   // fast
// Only the second can use an index on dept. Once documents have flowed through
// $group they are NEW documents, and no index describes them.

// Step 7: See what each stage does
db.students.aggregate([ { $match: { active: true } },
                        { $group: { _id: "$dept", n: { $sum: 1 } } } ],
                      { explain: true })
// Or truncate the pipeline and run the prefix -- the fastest way to find the
// stage that emptied your result. Compass's Aggregations tab does this for you.

// Step 8: Send the result to a collection
db.students.aggregate([
  { $group: { _id: "$dept", avg: { $avg: "$marks.maths" } } },
  { $merge: { into: "dept_summary", on: "_id",
              whenMatched: "replace", whenNotMatched: "insert" } }
])
// $out replaces the whole target collection; $merge updates it incrementally.
// Both must be the LAST stage.

In Python, through mongomock, 15_aggregation.py:

"""Experiment 15 — $match, $group, $project, $sort.

Every figure printed here is computed by mongomock and asserted against the
hand-worked arithmetic in unit-4.md, so the notes and the pipeline check each
other. The one exception is $stdDevPop, which mongomock raises
NotImplementedError on; that is asserted as a limitation and the value is
computed in Python instead, clearly labelled.
"""
import statistics

import mongomock
from fixtures import fresh_db


def by(rows, key="_id"):
    """A pipeline result keyed for assertion. Aggregation order is not a promise."""
    return {r[key]: r for r in rows}


def where_and_having_are_the_same_stage():
    """The whole pipeline, and the sentence that earns the marks."""
    db = fresh_db()
    rows = list(db.students.aggregate([
        {"$match":   {"active": True}},
        {"$group":   {"_id": "$dept", "avg": {"$avg": "$marks.maths"},
                      "n": {"$sum": 1}}},
        {"$match":   {"n": {"$gt": 1}}},
        {"$sort":    {"avg": -1}},
        # $round would go here on a real server; mongomock does not have it,
        # and unsupported_operators() below asserts that. Both averages are
        # exact anyway, so nothing is hidden by leaving it out.
        {"$project": {"_id": 0, "dept": "$_id", "avg": 1, "n": 1}},
    ]))

    # active: true drops Kiran (DS, 71), so DS is Asha and Ravi only.
    assert rows == [{"dept": "DS", "avg": 76.5, "n": 2},
                    {"dept": "Stats", "avg": 73.0, "n": 2}], rows
    assert (88 + 65) / 2 == 76.5
    assert (94 + 52) / 2 == 73.0

    # $sort was honoured: descending by avg.
    assert [r["avg"] for r in rows] == sorted((r["avg"] for r in rows), reverse=True)

    print("  WHERE active=true, GROUP BY dept, HAVING n>1, ORDER BY avg DESC")
    print("    DS    (88+65)/2 = 76.5  n=2")
    print("    Stats (94+52)/2 = 73.0  n=2")
    print("       Kiran is DS with 71 but active:false, so the $match BEFORE")
    print("       $group excludes him and the DS average RISES from 74.67")


def the_filter_changes_the_answer():
    """Same grouping, no $match: unit-4.md Problem 1(a)'s figures."""
    db = fresh_db()
    rows = by(db.students.aggregate([
        {"$group": {"_id": "$dept", "avg": {"$avg": "$marks.maths"},
                    "n": {"$sum": 1}}}]))

    assert rows["DS"]["n"] == 3 and rows["Stats"]["n"] == 2
    assert round(rows["DS"]["avg"], 3) == 74.667, rows["DS"]["avg"]
    assert (88 + 65 + 71) / 3 == rows["DS"]["avg"]
    assert rows["Stats"]["avg"] == 73.0

    print("  without the $match: DS (88+65+71)/3 = 74.667 over 3, Stats 73 over 2")
    print("       these are unit-4.md Problem 1(a)'s numbers, and the pair of")
    print("       results above is why 'which $match, and where' is the question")


def group_id_is_the_key_and_null_is_the_grand_total():
    db = fresh_db()
    total = list(db.students.aggregate([
        {"$group": {"_id": None, "avg": {"$avg": "$marks.maths"},
                    "n": {"$sum": 1}}}]))
    assert total == [{"_id": None, "avg": 74.0, "n": 5}], total
    assert (88 + 65 + 94 + 71 + 52) / 5 == 74.0

    # _id is not optional -- omitting it is an error, not a grand total.
    # (A real server says "a group specification must include an _id";
    #  mongomock reaches for the key and raises KeyError. Either way: rejected.)
    try:
        list(db.students.aggregate([{"$group": {"n": {"$sum": 1}}}]))
        raise SystemExit("$group without _id should not be accepted")
    except KeyError as exc:
        assert str(exc) == "'_id'", exc

    print("  _id: null -> one bucket: avg 74.0 over 5 (the grand total)")
    print("  _id omitted -> ERROR. It is the grouping KEY, not an option")


def the_accumulators():
    db = fresh_db()
    rows = by(db.students.aggregate([{"$group": {
        "_id":      "$dept",
        "n":        {"$sum": 1},
        "totMaths": {"$sum": "$marks.maths"},
        "avgMaths": {"$avg": "$marks.maths"},
        "best":     {"$max": "$marks.maths"},
        "worst":    {"$min": "$marks.maths"},
        "everyone": {"$push": "$name"},
        "distinct": {"$addToSet": "$dept"},
    }}]))

    ds = rows["DS"]
    assert ds["n"] == 3
    assert ds["totMaths"] == 224 == 88 + 65 + 71
    assert round(ds["avgMaths"], 3) == 74.667
    assert (ds["best"], ds["worst"]) == (88, 65)
    assert sorted(ds["everyone"]) == ["Asha", "Kiran", "Ravi"]
    assert ds["distinct"] == ["DS"], "addToSet de-duplicates; push does not"

    st = rows["Stats"]
    assert (st["n"], st["totMaths"], st["avgMaths"]) == (2, 146, 73.0)
    assert (st["best"], st["worst"]) == (94, 52)

    print("  dept   n  sum  avg      max  min  $push")
    for d in ("DS", "Stats"):
        r = rows[d]
        print(f"  {d:6s} {r['n']}  {r['totMaths']:3d}  {r['avgMaths']:7.3f}  "
              f"{r['best']:3d}  {r['worst']:3d}  {r['everyone']}")
    print("       $push and $addToSet have NO SQL equivalent -- they are why")
    print("       MongoDB never needed GROUP_CONCAT")


def unsupported_operators():
    """Two operators this pipeline would use on a real server, and cannot here.

    Asserted rather than commented, so the day mongomock gains them this file
    fails and gets rewritten to test them instead of describing them.
    """
    db = fresh_db()

    try:
        list(db.students.aggregate([
            {"$group": {"_id": "$dept", "sd": {"$stdDevPop": "$marks.maths"}}}]))
        raise SystemExit("mongomock now implements $stdDevPop -- assert it here")
    except NotImplementedError as exc:
        assert "$stdDevPop" in str(exc), exc

    try:
        list(db.students.aggregate([
            {"$project": {"m": {"$round": ["$marks.maths", 1]}}}]))
        raise SystemExit("mongomock now implements $round -- use it above")
    except Exception as exc:
        assert "$round" in str(exc), exc

    print("  not available in mongomock (both work on a real server):")
    print("    $round      OperationFailure: Unrecognized expression '$round'")
    print("    $stdDevPop  NotImplementedError")
    print("  so $stdDevPop is computed in Python here instead:")
    for dept in ("DS", "Stats"):
        vals = [d["marks"]["maths"] for d in db.students.find({"dept": dept})]
        pop = statistics.pstdev(vals)
        samp = statistics.stdev(vals)
        print(f"    {dept:6s} {vals}  $stdDevPop {pop:7.4f}   $stdDevSamp {samp:7.4f}")
    print("       Course 4's distinction, unchanged: $stdDevPop divides by n,")
    print("       $stdDevSamp by n-1. MongoDB makes you choose, as R does")


def project_computes_and_renames():
    db = fresh_db()
    rows = by(db.students.aggregate([{"$project": {
        "_id": 0, "name": 1,
        "total": {"$add": ["$marks.maths", "$marks.stats"]},
        "band": {"$switch": {"branches": [
            {"case": {"$gte": ["$marks.maths", 75]}, "then": "Distinction"},
            {"case": {"$gte": ["$marks.maths", 60]}, "then": "First"},
            {"case": {"$gte": ["$marks.maths", 40]}, "then": "Pass"}],
            "default": "Fail"}},
    }}], ), key="name")

    assert rows["Asha"] == {"name": "Asha", "total": 179, "band": "Distinction"}
    assert rows["Meena"]["total"] == 183 and rows["Meena"]["band"] == "Distinction"
    assert rows["Ravi"]["band"] == "First" and rows["Kiran"]["band"] == "First"
    assert rows["Bhanu"]["band"] == "Pass"
    assert all("dept" not in r for r in rows.values()), \
        "$project is EXCLUSIVE: naming any field drops the rest"

    print("  name   total  band")
    for n in ("Meena", "Asha", "Kiran", "Ravi", "Bhanu"):
        print(f"  {n:6s} {rows[n]['total']:5d}  {rows[n]['band']}")
    print("       dept is GONE -- $project keeps only what you name. Use")
    print("       $addFields when you meant 'everything, plus this'")


def addfields_keeps_everything():
    db = fresh_db()
    rows = list(db.students.aggregate([
        {"$addFields": {"total": {"$add": ["$marks.maths", "$marks.stats"]}}},
        {"$sort": {"total": -1}},
        {"$limit": 3},
    ]))

    assert [(r["name"], r["total"]) for r in rows] == \
        [("Meena", 183), ("Asha", 179), ("Kiran", 137)], rows
    assert all("dept" in r and "subjects" in r for r in rows), \
        "$addFields adds; it never removes"

    print("  top 3 by total, via $addFields: Meena 183, Asha 179, Kiran 137")
    print("       dept and subjects survived -- that is the difference")


def match_first_or_pay_for_it():
    """The optimiser helps, but only when it can. State the rule, not the hope."""
    db = fresh_db()

    late = list(db.students.aggregate([
        {"$group": {"_id": "$dept", "n": {"$sum": 1}}},
        {"$match": {"_id": "DS"}}]))
    early = list(db.students.aggregate([
        {"$match": {"dept": "DS"}},
        {"$group": {"_id": "$dept", "n": {"$sum": 1}}}]))

    assert late == early == [{"_id": "DS", "n": 3}], (late, early)

    print("  both orders return [{_id: 'DS', n: 3}] -- SAME ANSWER, different cost")
    print("       $match first can use an index on dept and groups 3 documents;")
    print("       $match last groups all 5 and then filters GROUPS, which no")
    print("       index describes, because they are new documents")
    print("       On 5 documents this is invisible. On 5,000,000 it is the")
    print("       whole query -- and that is what explain() shows you")


def main():
    print("Experiment 15 -- $match, $group, $project, $sort")
    # Step 1: $match then $group: WHERE and HAVING
    where_and_having_are_the_same_stage()
    # Step 2: See the filter change the answer
    the_filter_changes_the_answer()
    # Step 3: Group by a key, or by null for a total
    group_id_is_the_key_and_null_is_the_grand_total()
    # Step 4: Use the accumulators
    the_accumulators()
    # Step 5: Note the operators mongomock lacks
    unsupported_operators()
    # Step 6: $project: compute and rename
    project_computes_and_renames()
    # Step 7: $addFields: keep everything
    addfields_keeps_everything()
    # Step 8: $match first
    match_first_or_pay_for_it()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 15_aggregation.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.aggregate([
...   { $match:   { active: true } },
...   { $group:   { _id: "$dept", avg: { $avg: "$marks.maths" },
...                 n: { $sum: 1 } } },
...   { $match:   { n: { $gt: 1 } } },
...   { $sort:    { avg: -1 } },
...   { $project: { _id: 0, dept: "$_id", avg: { $round: ["$avg", 2] }, n: 1 } }
... ])
[ { n: 2, dept: 'DS', avg: 76.5 }, { n: 2, dept: 'Stats', avg: 73 } ]
collegeDB> db.students.aggregate([
...   { $group: { _id: "$dept", avgMaths: { $avg: "$marks.maths" },
...               n: { $sum: 1 } } },
...   { $sort: { _id: 1 } }
... ])
[
  { _id: 'DS', avgMaths: 74.66666666666667, n: 3 },
  { _id: 'Stats', avgMaths: 73, n: 2 }
]
collegeDB> db.students.aggregate([ { $group: { _id: null, avg: { $avg: "$marks.maths" },
...                                     n: { $sum: 1 } } } ])
[ { _id: null, avg: 74, n: 5 } ]
collegeDB> db.students.aggregate([
...   { $group: {
...       _id: "$dept",
...       n:        { $sum: 1 },                    // COUNT(*)
...       totMaths: { $sum: "$marks.maths" },       // SUM
...       avgMaths: { $avg: "$marks.maths" },       // AVG
...       best:     { $max: "$marks.maths" },       // MAX
...       worst:    { $min: "$marks.maths" },       // MIN
...       sd:       { $stdDevPop: "$marks.maths" }, // Course 4's population sd
...       everyone: { $push: "$name" },             // ALL values, as an array
...       distinct: { $addToSet: "$name" },         // DISTINCT values
...       anyone:   { $first: "$name" }             // needs a $sort to be meaningful
...   } },
...   { $set: { distinct: { $sortArray: { input: "$distinct", sortBy: 1 } } } },
...   { $sort: { _id: 1 } }
... ])
[
  {
    _id: 'DS',
    n: 3,
    totMaths: 224,
    avgMaths: 74.66666666666667,
    best: 88,
    worst: 65,
    sd: 9.741092797468305,
    everyone: [ 'Asha', 'Ravi', 'Kiran' ],
    distinct: [ 'Asha', 'Kiran', 'Ravi' ],
    anyone: 'Asha'
  },
  {
    _id: 'Stats',
    n: 2,
    totMaths: 146,
    avgMaths: 73,
    best: 94,
    worst: 52,
    sd: 21,
    everyone: [ 'Meena', 'Bhanu' ],
    distinct: [ 'Bhanu', 'Meena' ],
    anyone: 'Meena'
  }
]
collegeDB> db.students.aggregate([
...   { $project: {
...       _id: 0,
...       name: 1,                                       // include
...       total: { $add: ["$marks.maths", "$marks.stats"] },
...       pct:   { $round: [ { $divide: [ { $add: ["$marks.maths", "$marks.stats"] },
...                                       2 ] }, 1 ] },
...       dept:  "$dept",                                // rename by re-assigning
...       band:  { $switch: { branches: [
...                  { case: { $gte: ["$marks.maths", 75] }, then: "Distinction" },
...                  { case: { $gte: ["$marks.maths", 60] }, then: "First" },
...                  { case: { $gte: ["$marks.maths", 40] }, then: "Pass" } ],
...                default: "Fail" } }
...   } }
... ])
[
  {
    name: 'Asha',
    total: 179,
    pct: 89.5,
    dept: 'DS',
    band: 'Distinction'
  },
  { name: 'Ravi', total: 123, pct: 61.5, dept: 'DS', band: 'First' },
  {
    name: 'Meena',
    total: 183,
    pct: 91.5,
    dept: 'Stats',
    band: 'Distinction'
  },
  { name: 'Kiran', total: 137, pct: 68.5, dept: 'DS', band: 'First' },
  { name: 'Bhanu', total: 99, pct: 49.5, dept: 'Stats', band: 'Pass' }
]
collegeDB> db.students.aggregate([
...   { $addFields: { total: { $add: ["$marks.maths", "$marks.stats"] } } },
...   { $sort: { total: -1 } },
...   { $limit: 3 }
... ])
[
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: [ 'Stats', 'R' ],
    age: 20,
    active: true,
    total: 183
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: [ 'DS', 'Stats', 'Python' ],
    age: 20,
    active: true,
    total: 179
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: [ 'DS' ],
    age: 22,
    active: false,
    total: 137
  }
]
collegeDB> db.students.aggregate([ { $group: { _id: "$dept", n: { $sum: 1 } } },
...                         { $match: { _id: "DS" } } ])       // SLOW: groups all
[ { _id: 'DS', n: 3 } ]
collegeDB> db.students.aggregate([ { $match: { dept: "DS" } },
...                         { $group: { _id: "$dept", n: { $sum: 1 } } } ])   // fast
[ { _id: 'DS', n: 3 } ]
collegeDB> db.students.aggregate([ { $match: { active: true } },
...                         { $group: { _id: "$dept", n: { $sum: 1 } } } ],
...                       { explain: true })
{
  explainVersion: '2',
  queryPlanner: {
    namespace: 'collegeDB.students',
    parsedQuery: { active: { '$eq': true } },
    indexFilterSet: false,
    queryHash: '38BCAAF9',
    planCacheShapeHash: '38BCAAF9',
    planCacheKey: 'FE63AA7F',
    optimizationTimeMillis: 0,
    optimizedPipeline: true,
    maxIndexedOrSolutionsReached: false,
    maxIndexedAndSolutionsReached: false,
    maxScansToExplodeReached: false,
    prunedSimilarIndexes: false,
    winningPlan: {
      isCached: false,
      queryPlan: {
        stage: 'GROUP',
        planNodeId: 3,
        inputStage: {
          stage: 'COLLSCAN',
          planNodeId: 1,
          filter: { active: { '$eq': true } },
          nss: 'collegeDB.students',
          direction: 'forward'
        }
      },
      slotBasedPlan: {
        slots: '$$RESULT=s9 env: {  }',
        stages: '[3] project [s9 = newBsonObj("_id", s7, "n", s8)] \n' +
          '[3] project [s8 = (convert ( s6, int32) ?: s6)] \n' +
          '[3] group [s7] [s6 = count()] spillSlots[s5] mergingExprs[sum(s5)] \n' +
          '[3] project [s7 = (s3 ?: null)] \n' +
          '[1] filter {traverseF(s4, lambda(l2.0) { ((move(l2.0) == true) ?: false) }, false)} \n' +
          '[1] scan generic [s1 = record, s2 = recordId] [s3 = dept, s4 = active] @"ac85a8ec-c420-4d7a-9134-2c217094af59" '
      }
    },
    rejectedPlans: []
  },
  executionStats: {
    executionSuccess: true,
    nReturned: 2,
    executionTimeMillis: 8,
    totalKeysExamined: 0,
    totalDocsExamined: 5,
    executionStages: {
      stage: 'project',
      planNodeId: 3,
      nReturned: 2,
      executionTimeMillisEstimate: 6,
      opens: 1,
      closes: 1,
      saveState: 0,
      restoreState: 0,
      isEOF: 1,
      projections: { '9': 'newBsonObj("_id", s7, "n", s8) ' },
      inputStage: {
        stage: 'project',
        planNodeId: 3,
        nReturned: 2,
        executionTimeMillisEstimate: 6,
        opens: 1,
        closes: 1,
        saveState: 0,
        restoreState: 0,
        isEOF: 1,
        projections: { '8': '(convert ( s6, int32) ?: s6) ' },
        inputStage: {
          stage: 'group',
          planNodeId: 3,
          nReturned: 2,
          executionTimeMillisEstimate: 6,
          opens: 1,
          closes: 1,
          saveState: 0,
          restoreState: 0,
          isEOF: 1,
          groupBySlots: [ Long('7') ],
          expressions: {
            '6': 'count() ',
            initExprs: { '6': null, mergingExprs: { '5': 'sum(s5) ' } }
          },
          usedDisk: true,
          spills: 2,
          spilledBytes: 36,
          spilledRecords: 2,
          spilledDataStorageSize: 4096,
          peakTrackedMemBytes: 50,
          inputStage: {
            stage: 'project',
            planNodeId: 3,
            nReturned: 4,
            executionTimeMillisEstimate: 0,
            opens: 1,
            closes: 1,
            saveState: 0,
            restoreState: 0,
            isEOF: 1,
            projections: { '7': '(s3 ?: null) ' },
            inputStage: {
              stage: 'filter',
              planNodeId: 1,
              nReturned: 4,
              executionTimeMillisEstimate: 0,
              opens: 1,
              closes: 1,
              saveState: 0,
              restoreState: 0,
              isEOF: 1,
              numTested: 5,
              filter: 'traverseF(s4, lambda(l2.0) { ((move(l2.0) == true) ?: false) }, false) ',
              inputStage: {
                stage: 'scan',
                planNodeId: 1,
                nReturned: 5,
                executionTimeMillisEstimate: 0,
                opens: 1,
                closes: 1,
                saveState: 0,
                restoreState: 0,
                isEOF: 1,
                numReads: 5,
                recordSlot: 1,
                recordIdSlot: 2,
                scanFieldNames: [ 'dept', 'active' ],
                scanFieldSlots: [ Long('3'), Long('4') ]
              }
            }
          }
        }
      }
    },
    allPlansExecution: []
  },
  queryShapeHash: 'A17918C996816EA2A7275E096C6198D8C70A925310205C2EF51E25DFDC7B2BD2',
  peakTrackedMemBytes: Long('50'),
  command: {
    aggregate: 'students',
    pipeline: [
      { '$match': { active: true } },
      { '$group': { _id: '$dept', n: { '$sum': 1 } } }
    ],
    cursor: {},
    '$db': 'collegeDB'
  },
  serverInfo: {
    host: 'vm',
    port: 27017,
    version: '8.3.7',
    gitVersion: 'nogitversion'
  },
  serverParameters: {
    internalQueryFacetBufferSizeBytes: 104857600,
    internalDocumentSourceGroupMaxMemoryBytes: 104857600,
    internalQueryMaxBlockingSortMemoryUsageBytes: 104857600,
    internalDocumentSourceSetWindowFieldsMaxMemoryBytes: 104857600,
    internalQueryFacetMaxOutputDocSizeBytes: 104857600,
    internalLookupStageIntermediateDocumentMaxSizeBytes: 104857600,
    internalQueryProhibitBlockingMergeOnMongoS: 0,
    internalQueryMaxAddToSetBytes: 104857600,
    internalQueryFrameworkControl: 'trySbeRestricted',
    internalQueryPlannerIgnoreIndexWithCollationForRegex: 1
  },
  ok: 1
}
collegeDB> db.students.aggregate([
...   { $group: { _id: "$dept", avg: { $avg: "$marks.maths" } } },
...   { $merge: { into: "dept_summary", on: "_id",
...               whenMatched: "replace", whenNotMatched: "insert" } }
... ])

In Python, through mongomock, 15_aggregation.py:

OUTPUT

Experiment 15 -- $match, $group, $project, $sort
  WHERE active=true, GROUP BY dept, HAVING n>1, ORDER BY avg DESC
    DS    (88+65)/2 = 76.5  n=2
    Stats (94+52)/2 = 73.0  n=2
       Kiran is DS with 71 but active:false, so the $match BEFORE
       $group excludes him and the DS average RISES from 74.67
  without the $match: DS (88+65+71)/3 = 74.667 over 3, Stats 73 over 2
       these are unit-4.md Problem 1(a)'s numbers, and the pair of
       results above is why 'which $match, and where' is the question
  _id: null -> one bucket: avg 74.0 over 5 (the grand total)
  _id omitted -> ERROR. It is the grouping KEY, not an option
  dept   n  sum  avg      max  min  $push
  DS     3  224   74.667   88   65  ['Asha', 'Ravi', 'Kiran']
  Stats  2  146   73.000   94   52  ['Meena', 'Bhanu']
       $push and $addToSet have NO SQL equivalent -- they are why
       MongoDB never needed GROUP_CONCAT
  not available in mongomock (both work on a real server):
    $round      OperationFailure: Unrecognized expression '$round'
    $stdDevPop  NotImplementedError
  so $stdDevPop is computed in Python here instead:
    DS     [88, 65, 71]  $stdDevPop  9.7411   $stdDevSamp 11.9304
    Stats  [94, 52]  $stdDevPop 21.0000   $stdDevSamp 29.6985
       Course 4's distinction, unchanged: $stdDevPop divides by n,
       $stdDevSamp by n-1. MongoDB makes you choose, as R does
  name   total  band
  Meena    183  Distinction
  Asha     179  Distinction
  Kiran    137  First
  Ravi     123  First
  Bhanu     99  Pass
       dept is GONE -- $project keeps only what you name. Use
       $addFields when you meant 'everything, plus this'
  top 3 by total, via $addFields: Meena 183, Asha 179, Kiran 137
       dept and subjects survived -- that is the difference
  both orders return [{_id: 'DS', n: 3}] -- SAME ANSWER, different cost
       $match first can use an index on dept and groups 3 documents;
       $match last groups all 5 and then filters GROUPS, which no
       index describes, because they are new documents
       On 5 documents this is invisible. On 5,000,000 it is the
       whole query -- and that is what explain() shows you

Changed: two $groups now end with a $sort, and the $addToSet list of names is sorted with $sortArray. $group and $addToSet keep no order, and the names came back in a different order from one run to the next.

RESULT

With the active filter DS averages 76.5 over 2 and Stats 73 over 2; without it, DS averages 74.67 over 3.

Experiment 16 — $lookup, $unwind, $bucket

1. Question

Aggregate across collections and arrays with $lookup, $unwind and $bucket.

2. Aim

Unwind arrays, join collections, and bucket values into ranges.

3. Steps

In mongosh, 16_advanced_agg.js:

  1. Load the sample data.
  2. $unwind, and count the subjects.
  3. See $unwind drop empty and missing arrays.
  4. $lookup, and count enrolments per course.
  5. Filter the joined side first, with $lookup's pipeline.
  6. $bucket and $facet.

In Python, through mongomock, 16_advanced_agg.py:

  1. $unwind: one document per element.
  2. Count what the arrays hold.
  3. See $unwind drop empty and missing arrays.
  4. includeArrayIndex, and fields that are not arrays.
  5. $lookup: an array, by a left outer join.
  6. Join both ways.
  7. Note that mongomock lacks $lookup's pipeline form.
  8. $bucket: closed below, open above.
  9. See why the top boundary is 101.
  10. Note that mongomock lacks $bucketAuto.
  11. $facet: several pipelines in one pass.

THE POINT

Subject counts DS 3, Stats 3, Python 2, R 1; and the bucket distribution — with the top boundary at 101, not 100, because buckets are [lower, upper) and 100 as the boundary would lose a perfect scorer to default.

$unwind silently drops empty and missing arrays, and preserveNullAndEmptyArrays: true keeps them. That one is worth seeing fail.

4. Programme

In mongosh, 16_advanced_agg.js:

// Experiment 16 -- $lookup, $unwind and $bucket.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 16_advanced_agg.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// Start mongosh in this folder: the line after `use collegeDB` loads the
// sample data, 00_sample_data.js.

// Step 1: Load the sample data
use collegeDB
load("00_sample_data.js")

// =============================================================================
// $unwind -- one output document per array element
// =============================================================================
// Step 2: $unwind, and count the subjects
db.students.aggregate([ { $unwind: "$subjects" } ])
// 5 students with 3+2+2+1+1 subjects -> 9 documents out.

// The only way to count array CONTENTS:
db.students.aggregate([
  { $unwind: "$subjects" },
  { $group:  { _id: "$subjects", n: { $sum: 1 },
               who: { $push: "$name" } } },
  { $sort:   { n: -1, _id: 1 } }
])
// DS 3, Stats 3, Python 2, R 1

// Step 3: See $unwind drop empty and missing arrays
db.students.insertOne({ _id: 26, name: "Latha", dept: "DS", subjects: [] })
db.students.insertOne({ _id: 27, name: "Mohan", dept: "DS" })    // no field
db.students.aggregate([ { $unwind: "$subjects" },
                        { $count: "rows" } ])       // Latha and Mohan are GONE

db.students.aggregate([
  { $unwind: { path: "$subjects", preserveNullAndEmptyArrays: true } },
  { $count: "rows" }
])                                                  // both come back, subjects unset

// includeArrayIndex gives you the position, which $unwind otherwise loses:
db.students.aggregate([
  { $unwind: { path: "$subjects", includeArrayIndex: "pos" } },
  { $match:  { pos: 0 } }                           // each student's FIRST subject
])

// $unwind on a NON-array behaves as if it were a one-element array -- it does
// NOT error. That is why a typo'd path silently returns nothing instead.

// =============================================================================
// $lookup -- the left outer join
// =============================================================================
db.enrollments.aggregate([
  { $lookup: { from: "students", localField: "student_id",
               foreignField: "_id", as: "student" } },
  { $lookup: { from: "courses",  localField: "course_id",
               foreignField: "_id", as: "course" } },
  { $unwind: "$student" },
  { $unwind: "$course" },
  { $project: { _id: 0, name: "$student.name",
                title: "$course.title", grade: 1 } }
])

// as: ALWAYS an array, even for a one-to-one match -- hence the $unwind.
// LEFT OUTER: an unmatched document keeps its row with an EMPTY array, which
// is exactly why the $unwind after it silently deletes the unmatched rows.
// If you want them, preserveNullAndEmptyArrays: true.

// Step 4: $lookup, and count enrolments per course
db.courses.aggregate([
  { $lookup: { from: "enrollments", localField: "_id",
               foreignField: "course_id", as: "enrolled" } },
  { $project: { _id: 0, title: 1,
                n: { $size: "$enrolled" } } },       // $size, no $unwind needed
  { $sort: { n: -1 } }
])
// $size on the joined array beats $unwind + $group when you only want a count.

// Step 5: Filter the joined side first, with $lookup's pipeline
db.courses.aggregate([
  { $lookup: {
      from: "enrollments",
      let:  { cid: "$_id" },
      pipeline: [
        { $match: { $expr: { $and: [ { $eq: ["$course_id", "$$cid"] },
                                     { $eq: ["$grade", "A"] } ] } } },
        { $project: { _id: 0, student_id: 1 } }
      ],
      as: "aGrades" } }
])
// $$cid is the OUTER variable; $course_id the inner field. Two dollars means
// "from let". This form is how you avoid dragging 10,000 rows in to keep 3.

// =============================================================================
// $bucket and $bucketAuto -- histograms
// =============================================================================
// Step 6: $bucket and $facet
db.students.aggregate([
  { $bucket: {
      groupBy: "$marks.maths",
      boundaries: [0, 40, 60, 75, 101],
      default: "Other",
      output: { count: { $sum: 1 }, names: { $push: "$name" } } } }
])
// Boundaries are [lower, upper) -- CLOSED below, OPEN above.
// The last one is 101, NOT 100: with 100 as the top boundary a student who
// scored exactly 100 falls outside every bucket and lands in "Other".
// Without a `default`, an out-of-range value is an ERROR, not a silent drop.

db.students.aggregate([
  { $bucketAuto: { groupBy: "$marks.maths", buckets: 3 } }
])
// $bucketAuto picks the boundaries to even out the COUNTS. Good for
// exploration; useless for a report, because the boundaries move when the
// data does and yesterday's chart is not comparable with today's.

// $facet runs several pipelines over the SAME input, in one pass:
db.students.aggregate([
  { $facet: {
      byDept:    [ { $group: { _id: "$dept", n: { $sum: 1 } } }, { $sort: { _id: 1 } } ],
      byBand:    [ { $bucket: { groupBy: "$marks.maths",
                                boundaries: [0, 40, 60, 75, 101],
                                default: "Other",
                                output: { n: { $sum: 1 } } } } ],
      topThree:  [ { $sort: { "marks.maths": -1 } }, { $limit: 3 },
                   { $project: { _id: 0, name: 1 } } ] } }
])
// [Changed: the $sort in byDept was added. $group returns its groups in no fixed
// order, and they came back in a different order from one run to the next.]

In Python, through mongomock, 16_advanced_agg.py:

"""Experiment 16 — $lookup, $unwind and $bucket.

Executed and asserted: $unwind (including the empty-array trap), $lookup in
both directions, $bucket, and $facet.

Not available in mongomock, and asserted as unavailable rather than described
as if tested: $bucketAuto, and $lookup's let/pipeline form.
"""
from pymongo.errors import OperationFailure

from fixtures import fresh_db


def by(rows, key="_id"):
    return {r[key]: r for r in rows}


# =============================================================================
# $unwind
# =============================================================================

def one_document_per_element():
    db = fresh_db()
    before = db.students.count_documents({})
    after = len(list(db.students.aggregate([{"$unwind": "$subjects"}])))

    assert before == 5
    assert after == 9 == 3 + 2 + 2 + 1 + 1, after

    print(f"  {before} students -> {after} documents (3+2+2+1+1 subjects)")


def counting_array_contents():
    """The whole reason $unwind exists. $group alone cannot do this."""
    db = fresh_db()
    rows = by(db.students.aggregate([
        {"$unwind": "$subjects"},
        {"$group": {"_id": "$subjects", "n": {"$sum": 1},
                    "who": {"$push": "$name"}}},
        {"$sort": {"n": -1, "_id": 1}}]))

    assert {k: v["n"] for k, v in rows.items()} == \
        {"DS": 3, "Stats": 3, "Python": 2, "R": 1}
    assert sorted(rows["DS"]["who"]) == ["Asha", "Kiran", "Ravi"]
    assert sorted(rows["Stats"]["who"]) == ["Asha", "Bhanu", "Meena"]
    assert rows["R"]["who"] == ["Meena"]
    assert sum(v["n"] for v in rows.values()) == 9, "every element counted once"

    print("  subject counts:")
    for s in ("DS", "Stats", "Python", "R"):
        print(f"    {s:7s} {rows[s]['n']}  {sorted(rows[s]['who'])}")


def unwind_silently_drops_empty_and_missing():
    """The one worth seeing fail. Two students vanish from a count."""
    db = fresh_db()
    db.students.insert_one({"_id": 26, "name": "Latha", "dept": "DS",
                            "subjects": []})
    db.students.insert_one({"_id": 27, "name": "Mohan", "dept": "DS"})

    assert db.students.count_documents({}) == 7

    dropped = list(db.students.aggregate([{"$unwind": "$subjects"}]))
    assert len(dropped) == 9, len(dropped)
    assert "Latha" not in {d["name"] for d in dropped}
    assert "Mohan" not in {d["name"] for d in dropped}

    kept = list(db.students.aggregate([
        {"$unwind": {"path": "$subjects",
                     "preserveNullAndEmptyArrays": True}}]))
    assert len(kept) == 11, len(kept)
    survivors = {d["name"] for d in kept}
    assert "Latha" in survivors and "Mohan" in survivors
    assert all("subjects" not in d for d in kept
               if d["name"] in ("Latha", "Mohan")), \
        "they come back with the field UNSET, not with an empty array"

    print("  7 students, two with no subjects (empty array / missing field):")
    print(f"    plain $unwind                        -> {len(dropped)} docs, both LOST")
    print(f"    preserveNullAndEmptyArrays: true     -> {len(kept)} docs, both kept")
    print("       'my count is short and I cannot see why' is nearly always this")


def includearrayindex_and_non_arrays():
    db = fresh_db()
    firsts = list(db.students.aggregate([
        {"$unwind": {"path": "$subjects", "includeArrayIndex": "pos"}},
        {"$match": {"pos": 0}},
        {"$project": {"_id": 0, "name": 1, "subjects": 1}}]))

    assert len(firsts) == 5, firsts
    assert {d["name"]: d["subjects"] for d in firsts} == \
        {"Asha": "DS", "Ravi": "DS", "Meena": "Stats",
         "Kiran": "DS", "Bhanu": "Stats"}

    # $unwind on a NON-array is not an error: it acts as a one-element array.
    scalar = list(db.students.aggregate([{"$unwind": "$name"}]))
    assert len(scalar) == 5, "unwinding a string yields the same 5 documents"

    # A typo'd path is therefore SILENT -- it just returns nothing.
    typo = list(db.students.aggregate([{"$unwind": "$subject"}]))
    assert typo == [], "no error, no rows -- the commonest silent failure"

    print("  includeArrayIndex 'pos', $match pos:0 -> each student's FIRST subject")
    print("  $unwind on a string   -> 5 documents (treated as one element)")
    print("  $unwind on '$subject' -> 0 documents (a TYPO, and it does not error)")


# =============================================================================
# $lookup
# =============================================================================

def lookup_returns_an_array_and_is_a_left_outer_join():
    db = fresh_db()
    db.enrollments.insert_one({"student_id": 21, "course_id": "GONE404",
                               "grade": "F"})

    joined = list(db.enrollments.aggregate([
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}}]))

    assert len(joined) == 6, "every enrolment survives -- LEFT outer"
    sizes = {d["course_id"]: len(d["course"]) for d in joined}
    assert sizes["GONE404"] == 0, "no match -> EMPTY ARRAY, not a missing row"
    assert all(v == 1 for k, v in sizes.items() if k != "GONE404"), sizes
    assert all(isinstance(d["course"], list) for d in joined), \
        "as: is ALWAYS an array, even one-to-one -- that is why $unwind follows"

    # And now the consequence: the $unwind after it deletes the orphan.
    unwound = list(db.enrollments.aggregate([
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": "$course"}]))
    assert len(unwound) == 5, "the orphan's empty array was dropped by $unwind"

    print("  6 enrolments, one pointing at a course that does not exist:")
    print("    after $lookup           -> 6 rows, the orphan has course: []")
    print("    after $lookup + $unwind -> 5 rows, the orphan is GONE")
    print("       $lookup is a LEFT outer join; the $unwind after it turns it")
    print("       into an inner one. Nothing warns you")


def joining_both_directions():
    db = fresh_db()

    # enrolments -> the student and the course behind each one
    rows = list(db.enrollments.aggregate([
        {"$lookup": {"from": "students", "localField": "student_id",
                     "foreignField": "_id", "as": "student"}},
        {"$lookup": {"from": "courses", "localField": "course_id",
                     "foreignField": "_id", "as": "course"}},
        {"$unwind": "$student"},
        {"$unwind": "$course"},
        {"$project": {"_id": 0, "name": "$student.name",
                      "title": "$course.title", "grade": 1}},
        {"$sort": {"name": 1, "title": 1}}]))

    assert len(rows) == 5
    assert rows[0] == {"name": "Asha", "title": "Data Science with R",
                       "grade": "A"}, rows[0]
    assert {r["name"] for r in rows} == {"Asha", "Ravi", "Meena", "Kiran"}

    # courses -> how many enrolled. $size beats $unwind when you want a count.
    counts = list(db.courses.aggregate([
        {"$lookup": {"from": "enrollments", "localField": "_id",
                     "foreignField": "course_id", "as": "enrolled"}},
        {"$project": {"_id": 0, "title": 1, "n": {"$size": "$enrolled"}}},
        {"$sort": {"n": -1, "title": 1}}]))

    assert counts == [{"title": "Data Science with R", "n": 2},
                      {"title": "Statistical Foundations", "n": 2},
                      {"title": "Web Technologies", "n": 1}], counts

    print("  enrolments -> student + course:")
    for r in rows:
        print(f"    {r['name']:6s} {r['title']:24s} {r['grade']}")
    print("  courses -> enrolment counts, via $size on the joined array:")
    for c in counts:
        print(f"    {c['title']:24s} {c['n']}")
    print("       $size needs no $unwind and no $group -- one stage, one pass")


def the_let_pipeline_form_is_not_implemented_here():
    """Asserted, so this file cannot start claiming to test what it describes."""
    db = fresh_db()
    try:
        list(db.courses.aggregate([{"$lookup": {
            "from": "enrollments",
            "let": {"cid": "$_id"},
            "pipeline": [{"$match": {"$expr": {"$and": [
                {"$eq": ["$course_id", "$$cid"]},
                {"$eq": ["$grade", "A"]}]}}}],
            "as": "aGrades"}}]))
        raise SystemExit("mongomock now implements let/pipeline -- assert it")
    except NotImplementedError as exc:
        assert "let" in str(exc), exc

    # The runnable equivalent: join everything, then filter. Same answer,
    # more work -- which is exactly the point the let form makes.
    rows = list(db.courses.aggregate([
        {"$lookup": {"from": "enrollments", "localField": "_id",
                     "foreignField": "course_id", "as": "e"}},
        {"$unwind": "$e"},
        {"$match": {"e.grade": "A"}},
        {"$group": {"_id": "$title", "n": {"$sum": 1}}},
        {"$sort": {"_id": 1}}]))
    assert rows == [{"_id": "Data Science with R", "n": 1},
                    {"_id": "Statistical Foundations", "n": 1}], rows

    print("  $lookup with let/pipeline: NotImplementedError in mongomock")
    print("  the join-then-filter equivalent DOES run, and gives:")
    for r in rows:
        print(f"    {r['_id']:24s} {r['n']} grade-A enrolment(s)")
    print("       same answer, and on real data far more expensive: it drags")
    print("       every enrolment in and then throws most of them away.")
    print("       $$cid is the OUTER let variable, $course_id the inner field")


# =============================================================================
# $bucket
# =============================================================================

def bucket_boundaries_are_closed_below_and_open_above():
    db = fresh_db()
    rows = by(db.students.aggregate([{"$bucket": {
        "groupBy": "$marks.maths",
        "boundaries": [0, 40, 60, 75, 101],
        "default": "Other",
        "output": {"count": {"$sum": 1}, "names": {"$push": "$name"}}}}]))

    assert {k: v["count"] for k, v in rows.items()} == {40: 1, 60: 2, 75: 2}
    assert rows[40]["names"] == ["Bhanu"]                       # 52
    assert sorted(rows[60]["names"]) == ["Kiran", "Ravi"]        # 71, 65
    assert sorted(rows[75]["names"]) == ["Asha", "Meena"]        # 88, 94
    assert 0 not in rows, "the 0-40 bucket is EMPTY and is simply not emitted"

    print("  bucket   count  names")
    for lo, hi in ((0, 40), (40, 60), (60, 75), (75, 101)):
        r = rows.get(lo)
        label = f"[{lo:3d},{hi:4d})"
        if r:
            print(f"  {label}  {r['count']:5d}  {sorted(r['names'])}")
        else:
            print(f"  {label}  {'--':>5s}  (empty buckets are NOT emitted)")


def why_the_top_boundary_is_101():
    """A perfect score is the test case, and 100 as the boundary loses it."""
    db = fresh_db()
    db.students.insert_one({"_id": 28, "name": "Perfect", "dept": "DS",
                            "marks": {"maths": 100, "stats": 100},
                            "subjects": ["DS"], "age": 20, "active": True})

    with_101 = by(db.students.aggregate([{"$bucket": {
        "groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 101],
        "default": "Other", "output": {"names": {"$push": "$name"}}}}]))
    assert "Perfect" in with_101[75]["names"], with_101

    with_100 = by(db.students.aggregate([{"$bucket": {
        "groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 100],
        "default": "Other", "output": {"names": {"$push": "$name"}}}}]))
    assert with_100["Other"]["names"] == ["Perfect"], with_100
    assert "Perfect" not in with_100[75]["names"]

    # And without a default, an out-of-range value is an ERROR.
    try:
        list(db.students.aggregate([{"$bucket": {
            "groupBy": "$marks.maths", "boundaries": [0, 40, 60, 75, 100],
            "output": {"n": {"$sum": 1}}}}]))
        raise SystemExit("$bucket should reject an out-of-range value with no default")
    except OperationFailure as exc:
        assert "no default was specified" in str(exc), exc

    print("  a student who scored exactly 100:")
    print("    boundaries [..., 75, 101] -> lands in [75,101)   CORRECT")
    print("    boundaries [..., 75, 100] -> lands in 'Other'    WRONG")
    print("    boundaries [..., 75, 100], no default -> ERROR")
    print("       buckets are [lower, upper): closed below, OPEN above. The top")
    print("       boundary must exceed the maximum, so it is max+1, not max")


def bucketauto_is_not_implemented_here():
    db = fresh_db()
    try:
        list(db.students.aggregate([
            {"$bucketAuto": {"groupBy": "$marks.maths", "buckets": 3}}]))
        raise SystemExit("mongomock now implements $bucketAuto -- assert it")
    except NotImplementedError as exc:
        assert "$bucketAuto" in str(exc), exc

    print("  $bucketAuto: NotImplementedError in mongomock")
    print("       it picks boundaries to even out the COUNTS, so you never")
    print("       state them. Good for a first look at unfamiliar data; wrong")
    print("       for a report, because the boundaries MOVE when the data does")
    print("       and last month's chart is no longer comparable with this one")


def facet_runs_several_pipelines_in_one_pass():
    db = fresh_db()
    out = list(db.students.aggregate([{"$facet": {
        "byDept":   [{"$group": {"_id": "$dept", "n": {"$sum": 1}}},
                     {"$sort": {"_id": 1}}],
        "topThree": [{"$sort": {"marks.maths": -1}}, {"$limit": 3},
                     {"$project": {"_id": 0, "name": 1}}],
    }}]))

    assert len(out) == 1, "$facet emits exactly ONE document"
    result = out[0]
    assert result["byDept"] == [{"_id": "DS", "n": 3},
                                {"_id": "Stats", "n": 2}], result["byDept"]
    assert [d["name"] for d in result["topThree"]] == ["Meena", "Asha", "Kiran"]

    print("  $facet -> ONE document holding both results:")
    print(f"    byDept   {result['byDept']}")
    print(f"    topThree {[d['name'] for d in result['topThree']]}")
    print("       one pass over the collection instead of two queries, which")
    print("       is how a dashboard gets all its panels in a single round trip")


def main():
    print("Experiment 16 -- $lookup, $unwind, $bucket")
    # Step 1: $unwind: one document per element
    one_document_per_element()
    # Step 2: Count what the arrays hold
    counting_array_contents()
    # Step 3: See $unwind drop empty and missing arrays
    unwind_silently_drops_empty_and_missing()
    # Step 4: includeArrayIndex, and fields that are not arrays
    includearrayindex_and_non_arrays()
    # Step 5: $lookup: an array, by a left outer join
    lookup_returns_an_array_and_is_a_left_outer_join()
    # Step 6: Join both ways
    joining_both_directions()
    # Step 7: Note that mongomock lacks $lookup's pipeline form
    the_let_pipeline_form_is_not_implemented_here()
    # Step 8: $bucket: closed below, open above
    bucket_boundaries_are_closed_below_and_open_above()
    # Step 9: See why the top boundary is 101
    why_the_top_boundary_is_101()
    # Step 10: Note that mongomock lacks $bucketAuto
    bucketauto_is_not_implemented_here()
    # Step 11: $facet: several pipelines in one pass
    facet_runs_several_pipelines_in_one_pass()


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 16_advanced_agg.js:

OUTPUT

test> use collegeDB
switched to db collegeDB
collegeDB> load("00_sample_data.js")
sample data loaded: 5 students, 3 courses, 5 enrollments
true
collegeDB> db.students.aggregate([ { $unwind: "$subjects" } ])
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: 'DS',
    age: 20,
    active: true
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: 'Stats',
    age: 20,
    active: true
  },
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: 'Python',
    age: 20,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: 'DS',
    age: 21,
    active: true
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: 'Python',
    age: 21,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: 'Stats',
    age: 20,
    active: true
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: 'R',
    age: 20,
    active: true
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: 'DS',
    age: 22,
    active: false
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: 'Stats',
    age: 21,
    active: true
  }
]
collegeDB> db.students.aggregate([
...   { $unwind: "$subjects" },
...   { $group:  { _id: "$subjects", n: { $sum: 1 },
...                who: { $push: "$name" } } },
...   { $sort:   { n: -1, _id: 1 } }
... ])
[
  { _id: 'DS', n: 3, who: [ 'Asha', 'Ravi', 'Kiran' ] },
  { _id: 'Stats', n: 3, who: [ 'Asha', 'Meena', 'Bhanu' ] },
  { _id: 'Python', n: 2, who: [ 'Asha', 'Ravi' ] },
  { _id: 'R', n: 1, who: [ 'Meena' ] }
]
collegeDB> db.students.insertOne({ _id: 26, name: "Latha", dept: "DS", subjects: [] })
{ acknowledged: true, insertedId: 26 }
collegeDB> db.students.insertOne({ _id: 27, name: "Mohan", dept: "DS" })    // no field
{ acknowledged: true, insertedId: 27 }
collegeDB> db.students.aggregate([ { $unwind: "$subjects" },
...                         { $count: "rows" } ])       // Latha and Mohan are GONE
[ { rows: 9 } ]
collegeDB> db.students.aggregate([
...   { $unwind: { path: "$subjects", preserveNullAndEmptyArrays: true } },
...   { $count: "rows" }
... ])                                                  // both come back, subjects unset
[ { rows: 11 } ]
collegeDB> db.students.aggregate([
...   { $unwind: { path: "$subjects", includeArrayIndex: "pos" } },
...   { $match:  { pos: 0 } }                           // each student's FIRST subject
... ])
[
  {
    _id: 21,
    name: 'Asha',
    dept: 'DS',
    marks: { maths: 88, stats: 91 },
    subjects: 'DS',
    age: 20,
    active: true,
    pos: Long('0')
  },
  {
    _id: 22,
    name: 'Ravi',
    dept: 'DS',
    marks: { maths: 65, stats: 58 },
    subjects: 'DS',
    age: 21,
    active: true,
    pos: Long('0')
  },
  {
    _id: 23,
    name: 'Meena',
    dept: 'Stats',
    marks: { maths: 94, stats: 89 },
    subjects: 'Stats',
    age: 20,
    active: true,
    pos: Long('0')
  },
  {
    _id: 24,
    name: 'Kiran',
    dept: 'DS',
    marks: { maths: 71, stats: 66 },
    subjects: 'DS',
    age: 22,
    active: false,
    pos: Long('0')
  },
  {
    _id: 25,
    name: 'Bhanu',
    dept: 'Stats',
    marks: { maths: 52, stats: 47 },
    subjects: 'Stats',
    age: 21,
    active: true,
    pos: Long('0')
  }
]
collegeDB> db.enrollments.aggregate([
...   { $lookup: { from: "students", localField: "student_id",
...                foreignField: "_id", as: "student" } },
...   { $lookup: { from: "courses",  localField: "course_id",
...                foreignField: "_id", as: "course" } },
...   { $unwind: "$student" },
...   { $unwind: "$course" },
...   { $project: { _id: 0, name: "$student.name",
...                 title: "$course.title", grade: 1 } }
... ])
[
  { grade: 'A', name: 'Asha', title: 'Data Science with R' },
  { grade: 'B', name: 'Asha', title: 'Statistical Foundations' },
  { grade: 'C', name: 'Ravi', title: 'Data Science with R' },
  { grade: 'A', name: 'Meena', title: 'Statistical Foundations' },
  { grade: 'B', name: 'Kiran', title: 'Web Technologies' }
]
collegeDB> db.courses.aggregate([
...   { $lookup: { from: "enrollments", localField: "_id",
...                foreignField: "course_id", as: "enrolled" } },
...   { $project: { _id: 0, title: 1,
...                 n: { $size: "$enrolled" } } },       // $size, no $unwind needed
...   { $sort: { n: -1 } }
... ])
[
  { title: 'Data Science with R', n: 2 },
  { title: 'Statistical Foundations', n: 2 },
  { title: 'Web Technologies', n: 1 }
]
collegeDB> db.courses.aggregate([
...   { $lookup: {
...       from: "enrollments",
...       let:  { cid: "$_id" },
...       pipeline: [
...         { $match: { $expr: { $and: [ { $eq: ["$course_id", "$$cid"] },
...                                      { $eq: ["$grade", "A"] } ] } } },
...         { $project: { _id: 0, student_id: 1 } }
...       ],
...       as: "aGrades" } }
... ])
[
  {
    _id: 'DSC301',
    title: 'Data Science with R',
    credits: 4,
    instructor: 'Dr. Rao',
    aGrades: [ { student_id: 21 } ]
  },
  {
    _id: 'STA302',
    title: 'Statistical Foundations',
    credits: 3,
    instructor: 'Dr. Devi',
    aGrades: [ { student_id: 23 } ]
  },
  {
    _id: 'WEB303',
    title: 'Web Technologies',
    credits: 3,
    instructor: 'Dr. Kumar',
    aGrades: []
  }
]
collegeDB> db.students.aggregate([
...   { $bucket: {
...       groupBy: "$marks.maths",
...       boundaries: [0, 40, 60, 75, 101],
...       default: "Other",
...       output: { count: { $sum: 1 }, names: { $push: "$name" } } } }
... ])
[
  { _id: 40, count: 1, names: [ 'Bhanu' ] },
  { _id: 60, count: 2, names: [ 'Ravi', 'Kiran' ] },
  { _id: 75, count: 2, names: [ 'Asha', 'Meena' ] },
  { _id: 'Other', count: 2, names: [ 'Latha', 'Mohan' ] }
]
collegeDB> db.students.aggregate([
...   { $bucketAuto: { groupBy: "$marks.maths", buckets: 3 } }
... ])
[
  { _id: { min: null, max: 52 }, count: 2 },
  { _id: { min: 52, max: 71 }, count: 2 },
  { _id: { min: 71, max: 94 }, count: 3 }
]
collegeDB> db.students.aggregate([
...   { $facet: {
...       byDept:    [ { $group: { _id: "$dept", n: { $sum: 1 } } }, { $sort: { _id: 1 } } ],
...       byBand:    [ { $bucket: { groupBy: "$marks.maths",
...                                 boundaries: [0, 40, 60, 75, 101],
...                                 default: "Other",
...                                 output: { n: { $sum: 1 } } } } ],
...       topThree:  [ { $sort: { "marks.maths": -1 } }, { $limit: 3 },
...                    { $project: { _id: 0, name: 1 } } ] } }
... ])
[
  {
    byDept: [ { _id: 'DS', n: 5 }, { _id: 'Stats', n: 2 } ],
    byBand: [
      { _id: 40, n: 1 },
      { _id: 60, n: 2 },
      { _id: 75, n: 2 },
      { _id: 'Other', n: 2 }
    ],
    topThree: [ { name: 'Meena' }, { name: 'Asha' }, { name: 'Kiran' } ]
  }
]

In Python, through mongomock, 16_advanced_agg.py:

OUTPUT

Experiment 16 -- $lookup, $unwind, $bucket
  5 students -> 9 documents (3+2+2+1+1 subjects)
  subject counts:
    DS      3  ['Asha', 'Kiran', 'Ravi']
    Stats   3  ['Asha', 'Bhanu', 'Meena']
    Python  2  ['Asha', 'Ravi']
    R       1  ['Meena']
  7 students, two with no subjects (empty array / missing field):
    plain $unwind                        -> 9 docs, both LOST
    preserveNullAndEmptyArrays: true     -> 11 docs, both kept
       'my count is short and I cannot see why' is nearly always this
  includeArrayIndex 'pos', $match pos:0 -> each student's FIRST subject
  $unwind on a string   -> 5 documents (treated as one element)
  $unwind on '$subject' -> 0 documents (a TYPO, and it does not error)
  6 enrolments, one pointing at a course that does not exist:
    after $lookup           -> 6 rows, the orphan has course: []
    after $lookup + $unwind -> 5 rows, the orphan is GONE
       $lookup is a LEFT outer join; the $unwind after it turns it
       into an inner one. Nothing warns you
  enrolments -> student + course:
    Asha   Data Science with R      A
    Asha   Statistical Foundations  B
    Kiran  Web Technologies         B
    Meena  Statistical Foundations  A
    Ravi   Data Science with R      C
  courses -> enrolment counts, via $size on the joined array:
    Data Science with R      2
    Statistical Foundations  2
    Web Technologies         1
       $size needs no $unwind and no $group -- one stage, one pass
  $lookup with let/pipeline: NotImplementedError in mongomock
  the join-then-filter equivalent DOES run, and gives:
    Data Science with R      1 grade-A enrolment(s)
    Statistical Foundations  1 grade-A enrolment(s)
       same answer, and on real data far more expensive: it drags
       every enrolment in and then throws most of them away.
       $$cid is the OUTER let variable, $course_id the inner field
  bucket   count  names
  [  0,  40)     --  (empty buckets are NOT emitted)
  [ 40,  60)      1  ['Bhanu']
  [ 60,  75)      2  ['Kiran', 'Ravi']
  [ 75, 101)      2  ['Asha', 'Meena']
  a student who scored exactly 100:
    boundaries [..., 75, 101] -> lands in [75,101)   CORRECT
    boundaries [..., 75, 100] -> lands in 'Other'    WRONG
    boundaries [..., 75, 100], no default -> ERROR
       buckets are [lower, upper): closed below, OPEN above. The top
       boundary must exceed the maximum, so it is max+1, not max
  $bucketAuto: NotImplementedError in mongomock
       it picks boundaries to even out the COUNTS, so you never
       state them. Good for a first look at unfamiliar data; wrong
       for a report, because the boundaries MOVE when the data does
       and last month's chart is no longer comparable with this one
  $facet -> ONE document holding both results:
    byDept   [{'n': 3, '_id': 'DS'}, {'n': 2, '_id': 'Stats'}]
    topThree ['Meena', 'Asha', 'Kiran']
       one pass over the collection instead of two queries, which
       is how a dashboard gets all its panels in a single round trip

Latha (an empty subjects array) and Mohan (none at all) have no maths mark either, so $bucket puts them in Other. Changed: the $facet's count by department now ends with a $sort, for the same reason as Experiment 15.

RESULT

Subject counts are DS 3, Stats 3, Python 2, R 1; $unwind drops Latha and Mohan; the buckets use 101 as the top boundary so a perfect score is kept.

Experiment 17 — Replication with a replica set

1. Question

Set up a replica set, and watch it replicate and fail over.

2. Aim

Initiate a three-member replica set, read from a secondary, force a failover, and configure the members.

3. Steps

  1. Get three servers.
  2. Initiate the set, and read its status.
  3. Write on the primary, and read on a secondary.
  4. Look at the oplog.
  5. Step the primary down.
  6. Set the write concern and the read concern.
  7. Make a member hidden and delayed, and add an arbiter.

WHAT TO DEMONSTRATE

What to demonstrate: rs.status() showing one PRIMARY and two SECONDARY; rs.stepDown() triggering an election; writes failing during it; and w: "majority" versus w: 1. The theory is Unit 5 §5.7 — and the question "why an odd number of members?" is asked every year.

_drive_17_replication.py starts three mongod processes on ports 27017–27019 and a fourth on 27020, as section 0's route without Docker describes, and types each part of the script into a shell on the member it names. It waits where you would wait — for the election after rs.initiate(), and for a new primary after rs.stepDown() — and asserts what the experiment shows. An election, the oplog and every time in rs.status() differ on every run, so this is one run's output, recorded.

To do this on your own machine, docker compose with three mongo services is the easiest route, or use Atlas — its free tier is a three-member replica set.

4. Programme

// Experiment 17 -- Replication: setting up and observing a replica set.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server:
// _drive_17_replication.py starts three mongod processes, as section 0
// describes, and runs each part of this file on the member it names. What each
// line printed is on the lab page, and
// tools/data-science/capture_lab_outputs.py runs it again. There is no .py
// half: mongomock is a library, not a server, and nothing about replication
// could stand in for three of them. (Until October 2026 mongod could not be
// installed where these labs are checked, and this file was desk-checked only.)

// =============================================================================
// Step 1: Get three servers
// 0. Getting three servers
// =============================================================================
// EASIEST -- MongoDB Atlas. The free tier IS a three-member replica set, which
// a local install is not. Connect with mongodb+srv:// and rs.status() works.
//
// LOCAL -- docker compose, three services on one network:
//
//   services:
//     mongo1: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27017:27017"] }
//     mongo2: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27018:27017"] }
//     mongo3: { image: mongo, command: --replSet rs0 --bind_ip_all, ports: ["27019:27017"] }
//
//   docker compose up -d
//   docker compose exec mongo1 mongosh
//
// WITHOUT DOCKER -- three mongod processes, three data directories, three ports:
//
//   mkdir -p /data/rs0-{1,2,3}
//   mongod --replSet rs0 --port 27017 --dbpath /data/rs0-1 --bind_ip localhost &
//   mongod --replSet rs0 --port 27018 --dbpath /data/rs0-2 --bind_ip localhost &
//   mongod --replSet rs0 --port 27019 --dbpath /data/rs0-3 --bind_ip localhost &

// =============================================================================
// Step 2: Initiate the set, and read its status
// 1. Initiating the set -- run ONCE, on ONE member
// =============================================================================
rs.initiate({
  _id: "rs0",
  members: [
    { _id: 0, host: "localhost:27017" },
    { _id: 1, host: "localhost:27018" },
    { _id: 2, host: "localhost:27019" }
  ]
})
// With docker compose, the hosts are the service names: mongo1:27017,
// mongo2:27017 and mongo3:27017.
// [Changed: the hosts were mongo1, mongo2 and mongo3, which exist only on the
// docker compose network. These are the three processes of the route without
// Docker, which is how this file is run.]

// The shell prompt changes: rs0 [direct: other] > ... then rs0 [primary] >
// Election takes a second or two. Until it finishes there is no primary and
// no writes are accepted.

rs.status()          // members[], each with stateStr, health, optimeDate
rs.conf()            // the configuration, with each member's votes and priority
rs.isMaster()        // or db.hello() -- who is primary right now?

// WHAT TO SHOW THE EXAMINER in rs.status():
//   * exactly ONE member with stateStr "PRIMARY"
//   * two with "SECONDARY"
//   * health: 1 on all three
//   * optimeDate close together -- that closeness IS replication lag

// =============================================================================
// Step 3: Write on the primary, and read on a secondary
// 2. Watching data replicate
// =============================================================================
// On the PRIMARY:
use collegeDB
db.students.insertOne({ _id: 21, name: "Asha", dept: "DS" })

// On a SECONDARY -- a second terminal, connected to port 27018:
//   mongosh --port 27018
use collegeDB
db.students.find()                      // Asha is there: the write has replicated
db.getMongo().getReadPref()             // mode 'primary' -- and yet it read a secondary

// A shell connected to ONE member reads from that member, whatever the read
// preference says. Connected to the SET --
//   mongosh "mongodb://localhost:27017,localhost:27018/?replicaSet=rs0"
// -- it sends every read to the primary, unless you say otherwise:
db.getMongo().setReadPref("secondaryPreferred")
// A secondary may be behind the primary, so reading from one is something
// you choose, by read preference, when slightly stale data is acceptable.
// [Corrected: this said the first find() fails, "not primary and
// secondaryOk=false", until setReadPref() is called. That was the old mongo
// shell. In mongosh 2, connected straight to a secondary, it reads, as above.
// It also had no use collegeDB: a new shell starts in test, where there is no
// Asha to find.]

// =============================================================================
// Step 4: Look at the oplog
// 3. The oplog -- how replication actually works
// =============================================================================
use local
db.oplog.rs.find().sort({ $natural: -1 }).limit(5)
db.oplog.rs.stats().maxSize                  // the CAP, in bytes
rs.printReplicationInfo()                    // oplog size and its time window

// The oplog is a CAPPED collection of idempotent operations. Secondaries tail
// it and replay it. Two consequences worth stating in the viva:
//   * idempotent, so replaying an entry twice is safe -- which is what makes
//     recovery after a crash possible at all
//   * capped, so if a secondary falls further behind than the oplog's time
//     window, it can no longer catch up and needs a FULL resync

// =============================================================================
// Step 5: Step the primary down
// 4. Failover -- the demonstration that earns the marks
// =============================================================================
rs.printSecondaryReplicationInfo()      // lag per secondary, in seconds

rs.stepDown(60)                         // primary steps down for 60 seconds
// Watch: an election starts, a secondary becomes PRIMARY, and for roughly
// 10-30 seconds there is NO primary and every write fails. Show that gap.

// Or pull the plug, which is more convincing:
//   docker compose stop mongo1
//   rs.status()      // mongo1 health 0, stateStr "(not reachable/healthy)"
//   docker compose start mongo1     // it rejoins as a SECONDARY, not primary

// =============================================================================
// Step 6: Set the write concern and the read concern
// 5. Write concern and read concern -- the durability dial
// =============================================================================
// On the NEW primary: after the step-down, connect to whichever member
// rs.status() now shows as PRIMARY.
use collegeDB
// [Corrected: this use collegeDB was missing. After section 3's use local, the
// inserts below went into the local database, which is never replicated.]
db.students.insertOne({ _id: 22, name: "Ravi" },
                      { writeConcern: { w: 1 } })
// Acknowledged by the PRIMARY only. Fast. Lost if the primary dies before the
// secondaries have it -- a "rollback".

db.students.insertOne({ _id: 23, name: "Meena" },
                      { writeConcern: { w: "majority", j: true, wtimeout: 5000 } })
// Acknowledged by a MAJORITY, and on disk (j: true). Slower, and survives the
// loss of any one member. ALWAYS set wtimeout, or a stalled member hangs you.

db.students.find().readConcern("majority")   // only data a majority holds
db.students.find().readConcern("local")      // the default: may be rolled back

// | w        | acknowledged by     | survives primary loss? |
// |----------|---------------------|------------------------|
// | 0        | nobody (fire and forget) | no                |
// | 1        | the primary         | NO                     |
// | majority | 2 of 3              | yes                    |

// =============================================================================
// Step 7: Make a member hidden and delayed, and add an arbiter
// 6. Priority, hidden members and arbiters
// =============================================================================
cfg = rs.conf()
cfg.members[2].priority = 0            // never becomes primary
cfg.members[2].hidden = true           // and clients never see it
cfg.members[2].secondaryDelaySecs = 3600   // an HOUR behind -- a live undo button
// [Corrected: this was slaveDelay, which MongoDB 5.0 renamed. MongoDB 8 refuses
// the old name: "BSON field 'MemberConfig.slaveDelay' is an unknown field".]
rs.reconfig(cfg)

// A delayed hidden member is the answer to "someone ran deleteMany({})".
// It has the data as it was an hour ago.

db.adminCommand({ setDefaultRWConcern: 1, defaultWriteConcern: { w: "majority" } })
rs.addArb("localhost:27020")           // an ARBITER: votes, stores no data
// An arbiter changes what "majority" means without holding any data, so
// MongoDB 5 and later refuse to add one until the default write concern has
// been set explicitly -- the line before it.
// [Corrected: the setDefaultRWConcern line was missing, and without it
// rs.addArb() fails: "Reconfig attempted to install a config that would change
// the implicit default write concern".]

// =============================================================================
// WHY AN ODD NUMBER OF MEMBERS? -- asked every year
// =============================================================================
// A primary must be elected by a STRICT MAJORITY of votes.
//
//   3 members: majority 2, tolerates 1 failure
//   4 members: majority 3, tolerates 1 failure   <-- no better than 3
//   5 members: majority 3, tolerates 2 failures
//
// The fourth member buys NO extra fault tolerance and adds a machine, network
// traffic and a chance of a tie. So: odd numbers.
//
// The majority rule also prevents SPLIT BRAIN. If the network partitions 3-2,
// only the side of 3 can elect a primary; the side of 2 has no majority and
// steps down to secondary. Two primaries accepting conflicting writes is
// impossible by construction, not by convention.
//
// This is Unit 5 §5.7, and it is CAP in practice: MongoDB chooses CONSISTENCY
// over availability, so the minority side refuses writes rather than diverge.

5. Execution and Results

OUTPUT

[mongosh connected to 127.0.0.1:27017, not yet in a set]
test> rs.initiate({
...   _id: "rs0",
...   members: [
...     { _id: 0, host: "localhost:27017" },
...     { _id: 1, host: "localhost:27018" },
...     { _id: 2, host: "localhost:27019" }
...   ]
... })
{
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104251, i: 1 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104251, i: 1 })
}

[waited for the election: PRIMARY,SECONDARY,SECONDARY]

[mongosh connected to 127.0.0.1:27019, the primary]
rs0 [direct: primary] test> rs.status()          // members[], each with stateStr, health, optimeDate
{
  set: 'rs0',
  date: ISODate('2026-10-04T08:57:46.891Z'),
  myState: 1,
  term: Long('1'),
  syncSourceHost: '',
  syncSourceId: -1,
  heartbeatIntervalMillis: Long('2000'),
  majorityVoteCount: 2,
  writeMajorityCount: 2,
  votingMembersCount: 3,
  writableVotingMembersCount: 3,
  optimes: {
    lastCommittedOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    lastCommittedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
    readConcernMajorityOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    appliedOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    durableOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    writtenOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
    lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
    lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z')
  },
  lastStableRecoveryTimestamp: Timestamp({ t: 1791104251, i: 1 }),
  electionCandidateMetrics: {
    lastElectionReason: 'electionTimeout',
    lastElectionDate: ISODate('2026-10-04T08:57:42.079Z'),
    electionTerm: Long('1'),
    lastCommittedOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
    lastSeenWrittenOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
    lastSeenOpTimeAtElection: { ts: Timestamp({ t: 1791104251, i: 1 }), t: Long('-1') },
    numVotesNeeded: 2,
    priorityAtElection: 1,
    electionTimeoutMillis: Long('10000'),
    numCatchUpOps: Long('0'),
    newTermStartDate: ISODate('2026-10-04T08:57:42.106Z'),
    wMajorityWriteAvailabilityDate: ISODate('2026-10-04T08:57:42.598Z')
  },
  members: [
    {
      _id: 0,
      name: 'localhost:27017',
      health: 1,
      state: 2,
      stateStr: 'SECONDARY',
      uptime: 15,
      optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeDurable: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
      optimeDurableDate: ISODate('2026-10-04T08:57:42.000Z'),
      optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
      lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastHeartbeat: ISODate('2026-10-04T08:57:46.098Z'),
      lastHeartbeatRecv: ISODate('2026-10-04T08:57:46.599Z'),
      pingMs: Long('0'),
      lastHeartbeatMessage: '',
      syncSourceHost: 'localhost:27019',
      syncSourceId: 2,
      infoMessage: '',
      configVersion: 1,
      configTerm: 1
    },
    {
      _id: 1,
      name: 'localhost:27018',
      health: 1,
      state: 2,
      stateStr: 'SECONDARY',
      uptime: 15,
      optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeDurable: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
      optimeDurableDate: ISODate('2026-10-04T08:57:42.000Z'),
      optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
      lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastHeartbeat: ISODate('2026-10-04T08:57:46.092Z'),
      lastHeartbeatRecv: ISODate('2026-10-04T08:57:45.094Z'),
      pingMs: Long('0'),
      lastHeartbeatMessage: '',
      syncSourceHost: 'localhost:27019',
      syncSourceId: 2,
      infoMessage: '',
      configVersion: 1,
      configTerm: 1
    },
    {
      _id: 2,
      name: 'localhost:27019',
      health: 1,
      state: 1,
      stateStr: 'PRIMARY',
      uptime: 18,
      optime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeDate: ISODate('2026-10-04T08:57:42.000Z'),
      optimeWritten: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
      optimeWrittenDate: ISODate('2026-10-04T08:57:42.000Z'),
      lastAppliedWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastDurableWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      lastWrittenWallTime: ISODate('2026-10-04T08:57:42.140Z'),
      syncSourceHost: '',
      syncSourceId: -1,
      infoMessage: 'Could not find member to sync from',
      electionTime: Timestamp({ t: 1791104262, i: 1 }),
      electionDate: ISODate('2026-10-04T08:57:42.000Z'),
      configVersion: 1,
      configTerm: 1,
      self: true,
      lastHeartbeatMessage: ''
    }
  ],
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104262, i: 17 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104262, i: 17 })
}
rs0 [direct: primary] test> rs.conf()            // the configuration, with each member's votes and priority
{
  _id: 'rs0',
  version: 1,
  term: 1,
  members: [
    {
      _id: 0,
      host: 'localhost:27017',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    },
    {
      _id: 1,
      host: 'localhost:27018',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    },
    {
      _id: 2,
      host: 'localhost:27019',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    }
  ],
  protocolVersion: Long('1'),
  writeConcernMajorityJournalDefault: true,
  settings: {
    chainingAllowed: true,
    heartbeatIntervalMillis: 2000,
    heartbeatTimeoutSecs: 10,
    electionTimeoutMillis: 10000,
    catchUpTimeoutMillis: -1,
    catchUpTakeoverDelayMillis: 30000,
    getLastErrorModes: {},
    getLastErrorDefaults: { w: 1, wtimeout: 0 },
    replicaSetId: ObjectId('6ac214fbd195cc8f74280fe8')
  }
}
rs0 [direct: primary] test> rs.isMaster()        // or db.hello() -- who is primary right now?
{
  topologyVersion: { processId: ObjectId('6ac214f81bed474ff801f3e6'), counter: Long('6') },
  hosts: [ 'localhost:27017', 'localhost:27018', 'localhost:27019' ],
  setName: 'rs0',
  setVersion: 1,
  ismaster: true,
  secondary: false,
  primary: 'localhost:27019',
  me: 'localhost:27019',
  electionId: ObjectId('7fffffff0000000000000001'),
  lastWrite: {
    opTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    lastWriteDate: ISODate('2026-10-04T08:57:42.000Z'),
    majorityOpTime: { ts: Timestamp({ t: 1791104262, i: 17 }), t: Long('1') },
    majorityWriteDate: ISODate('2026-10-04T08:57:42.000Z')
  },
  maxBsonObjectSize: 16777216,
  maxMessageSizeBytes: 48000000,
  maxWriteBatchSize: 100000,
  localTime: ISODate('2026-10-04T08:57:46.953Z'),
  logicalSessionTimeoutMinutes: 30,
  connectionId: 26,
  minWireVersion: 0,
  maxWireVersion: 28,
  readOnly: false,
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104262, i: 17 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104262, i: 17 }),
  isWritablePrimary: true
}
rs0 [direct: primary] test> use collegeDB
switched to db collegeDB
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 21, name: "Asha", dept: "DS" })
{ acknowledged: true, insertedId: 21 }

[mongosh connected to 127.0.0.1:27018, a secondary]
rs0 [direct: secondary] test> use collegeDB
switched to db collegeDB
rs0 [direct: secondary] collegeDB> db.students.find()                      // Asha is there: the write has replicated
[ { _id: 21, name: 'Asha', dept: 'DS' } ]
rs0 [direct: secondary] collegeDB> db.getMongo().getReadPref()             // mode 'primary' -- and yet it read a secondary
ReadPreference {
  mode: 'primary',
  tags: undefined,
  hedge: undefined,
  maxStalenessSeconds: undefined
}
rs0 [direct: secondary] collegeDB> db.getMongo().setReadPref("secondaryPreferred")

[mongosh connected to 127.0.0.1:27019, the primary]
rs0 [direct: primary] test> use local
switched to db local
rs0 [direct: primary] local> db.oplog.rs.find().sort({ $natural: -1 }).limit(5)
[
  {
    lsid: {
      id: UUID('5d38bf63-2894-4ec9-9e62-a15c41016796'),
      uid: Binary.createFromBase64('47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU=', 0)
    },
    txnNumber: Long('1'),
    op: 'i',
    ns: 'collegeDB.students',
    ui: UUID('bbcd39f2-8897-449a-9be1-e8bd273a7ff0'),
    o: { _id: 21, name: 'Asha', dept: 'DS' },
    o2: { _id: 21 },
    stmtId: 0,
    ts: Timestamp({ t: 1791104267, i: 2 }),
    t: Long('1'),
    v: Long('2'),
    wall: ISODate('2026-10-04T08:57:47.089Z'),
    prevOpTime: { ts: Timestamp({ t: 0, i: 0 }), t: Long('-1') }
  },
  {
    op: 'c',
    ns: 'collegeDB.$cmd',
    ui: UUID('bbcd39f2-8897-449a-9be1-e8bd273a7ff0'),
    o: {
      create: 'students',
      idIndex: { v: 2, key: { _id: 1 }, name: '_id_' }
    },
    o2: {
      catalogId: Long('21'),
      ident: 'a1566e98-ef1a-4d0a-adb2-031960a74fde',
      idIndexIdent: 'ff62dc33-5d23-426b-9110-5af4399f2413'
    },
    versionContext: { OFCV: '8.3' },
    ts: Timestamp({ t: 1791104267, i: 1 }),
    t: Long('1'),
    v: Long('2'),
    wall: ISODate('2026-10-04T08:57:47.089Z')
  },
  {
    op: 'c',
    ns: 'config.$cmd',
    ui: UUID('c69231f3-5a46-44c1-89e6-fb542ea67c9f'),
    o: {
      createIndexes: 'sampledQueriesDiff',
      v: 2,
      key: { expireAt: 1 },
      name: 'SampledQueriesDiffTTLIndex',
      expireAfterSeconds: 0
    },
    o2: { indexIdent: 'bf6df23a-ba1e-45c1-ba39-8e77f78c5ae1' },
    ts: Timestamp({ t: 1791104262, i: 17 }),
    t: Long('1'),
    v: Long('2'),
    wall: ISODate('2026-10-04T08:57:42.140Z')
  },
  {
    op: 'c',
    ns: 'config.$cmd',
    ui: UUID('c69231f3-5a46-44c1-89e6-fb542ea67c9f'),
    o: {
      create: 'sampledQueriesDiff',
      idIndex: { v: 2, key: { _id: 1 }, name: '_id_' }
    },
    o2: {
      catalogId: Long('20'),
      ident: 'f5ea0985-ff6c-4433-afa0-e359fa0d57a5',
      idIndexIdent: 'eed90427-f9ea-4074-a290-de4b7efc9bbd'
    },
    versionContext: { OFCV: '8.3' },
    ts: Timestamp({ t: 1791104262, i: 16 }),
    t: Long('1'),
    v: Long('2'),
    wall: ISODate('2026-10-04T08:57:42.140Z')
  },
  {
    op: 'c',
    ns: 'config.$cmd',
    ui: UUID('96827267-ad25-4d3c-9d1b-670dc16c0bf5'),
    o: {
      createIndexes: 'sampledQueries',
      v: 2,
      key: { expireAt: 1 },
      name: 'SampledQueriesTTLIndex',
      expireAfterSeconds: 0
    },
    o2: { indexIdent: '71e3a6c9-faf8-4928-9f86-babfe77fdbb3' },
    ts: Timestamp({ t: 1791104262, i: 15 }),
    t: Long('1'),
    v: Long('2'),
    wall: ISODate('2026-10-04T08:57:42.133Z')
  }
]
rs0 [direct: primary] local> db.oplog.rs.stats().maxSize                  // the CAP, in bytes
1038090240
rs0 [direct: primary] local> rs.printReplicationInfo()                    // oplog size and its time window
actual oplog size
'990 MB'
---
configured oplog size
'990 MB'
---
log length start to end
'16 secs (0 hrs)'
---
oplog first event time
'Sun Oct 04 2026 08:57:31 GMT+0000 (Coordinated Universal Time)'
---
oplog last event time
'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)'
---
now
'Sun Oct 04 2026 08:57:53 GMT+0000 (Coordinated Universal Time)'
rs0 [direct: primary] local> rs.printSecondaryReplicationInfo()      // lag per secondary, in seconds
source: localhost:27017
{
  syncedTo: 'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)',
  replLag: '0 secs (0 hrs) behind the primary '
}
---
source: localhost:27018
{
  syncedTo: 'Sun Oct 04 2026 08:57:47 GMT+0000 (Coordinated Universal Time)',
  replLag: '0 secs (0 hrs) behind the primary '
}
rs0 [direct: primary] local> rs.stepDown(60)                         // primary steps down for 60 seconds
{
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104267, i: 2 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104267, i: 2 })
}

[waited for the election after the step-down: the primary is now 127.0.0.1:27017]

[mongosh connected to 127.0.0.1:27017, the new primary]
rs0 [direct: primary] test> use collegeDB
switched to db collegeDB
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 22, name: "Ravi" },
...                       { writeConcern: { w: 1 } })
{ acknowledged: true, insertedId: 22 }
rs0 [direct: primary] collegeDB> db.students.insertOne({ _id: 23, name: "Meena" },
...                       { writeConcern: { w: "majority", j: true, wtimeout: 5000 } })
{ acknowledged: true, insertedId: 23 }
rs0 [direct: primary] collegeDB> db.students.find().readConcern("majority")   // only data a majority holds
[
  { _id: 21, name: 'Asha', dept: 'DS' },
  { _id: 22, name: 'Ravi' },
  { _id: 23, name: 'Meena' }
]
rs0 [direct: primary] collegeDB> db.students.find().readConcern("local")      // the default: may be rolled back
[
  { _id: 21, name: 'Asha', dept: 'DS' },
  { _id: 22, name: 'Ravi' },
  { _id: 23, name: 'Meena' }
]
rs0 [direct: primary] collegeDB> cfg = rs.conf()
{
  _id: 'rs0',
  version: 1,
  term: 2,
  members: [
    {
      _id: 0,
      host: 'localhost:27017',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    },
    {
      _id: 1,
      host: 'localhost:27018',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    },
    {
      _id: 2,
      host: 'localhost:27019',
      arbiterOnly: false,
      buildIndexes: true,
      hidden: false,
      priority: 1,
      tags: {},
      secondaryDelaySecs: Long('0'),
      votes: 1
    }
  ],
  protocolVersion: Long('1'),
  writeConcernMajorityJournalDefault: true,
  settings: {
    chainingAllowed: true,
    heartbeatIntervalMillis: 2000,
    heartbeatTimeoutSecs: 10,
    electionTimeoutMillis: 10000,
    catchUpTimeoutMillis: -1,
    catchUpTakeoverDelayMillis: 30000,
    getLastErrorModes: {},
    getLastErrorDefaults: { w: 1, wtimeout: 0 },
    replicaSetId: ObjectId('6ac214fbd195cc8f74280fe8')
  }
}
rs0 [direct: primary] collegeDB> cfg.members[2].priority = 0            // never becomes primary
0
rs0 [direct: primary] collegeDB> cfg.members[2].hidden = true           // and clients never see it
true
rs0 [direct: primary] collegeDB> cfg.members[2].secondaryDelaySecs = 3600   // an HOUR behind -- a live undo button
3600
rs0 [direct: primary] collegeDB> rs.reconfig(cfg)
{
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104275, i: 3 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104275, i: 3 })
}
rs0 [direct: primary] collegeDB> db.adminCommand({ setDefaultRWConcern: 1, defaultWriteConcern: { w: "majority" } })
{
  defaultReadConcern: { level: 'local' },
  defaultWriteConcern: { w: 'majority', wtimeout: 0 },
  updateOpTime: Timestamp({ t: 1791104275, i: 3 }),
  updateWallClockTime: ISODate('2026-10-04T08:57:55.610Z'),
  defaultWriteConcernSource: 'global',
  defaultReadConcernSource: 'implicit',
  localUpdateWallClockTime: ISODate('2026-10-04T08:57:55.621Z'),
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104275, i: 5 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104275, i: 5 })
}
rs0 [direct: primary] collegeDB> rs.addArb("localhost:27020")           // an ARBITER: votes, stores no data
{
  ok: 1,
  '$clusterTime': {
    clusterTime: Timestamp({ t: 1791104280, i: 1 }),
    signature: {
      hash: Binary.createFromBase64('AAAAAAAAAAAAAAAAAAAAAAAAAAA=', 0),
      keyId: Long('0')
    }
  },
  operationTime: Timestamp({ t: 1791104280, i: 1 })
}

[rs.status() at the end: 27017 PRIMARY, 27018 SECONDARY, 27019 SECONDARY, 27020 ARBITER]
checked: one PRIMARY and two SECONDARY members; Asha replicated to the secondary, which read her with the read preference 'primary'; after rs.stepDown() the primary moved from 27019 to 27017; the reconfiguration and the arbiter were accepted

Four corrections, each found by running it:

And the hosts: rs.initiate named mongo1–mongo3, which exist only on the Docker network; it now names the local route's three processes, with the Docker form in a comment.

RESULT

One PRIMARY and two SECONDARY; Asha replicated to the secondary; after rs.stepDown() a new primary was elected; the hidden delayed member and the arbiter were accepted.

Experiment 18 — GridFS

1. Question

Store and retrieve a large file with GridFS.

2. Aim

Put a file into GridFS, see how it is stored, and get it back.

3. Steps

  1. Put a file in GridFS with mongofiles.
  2. Look at what it stored, and count the chunks.
  3. Do it from a driver.
  4. Query by metadata.
  5. Delete it, with its chunks.

WHAT TO DEMONSTRATE

What to demonstrate: that a file appears as one document in fs.files and many in fs.chunks, and that chunks == ceil(bytes / 261120). The point to state: GridFS is for files over 16 MB, or where you need to read ranges — for smaller files, object storage is usually better.

_drive_18_gridfs.py makes a 10 MB file and runs section 1's mongofiles commands for real — put, list, and get to a copy, compared byte for byte — then types sections 2 to 4 into mongosh, and deletes the file with mongofiles delete, as section 5 recommends.

4. Programme

// Experiment 18 -- GridFS: storing and retrieving large files.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server:
// _drive_18_gridfs.py puts a 10 MB file into GridFS with mongofiles, as section
// 1 describes, then types the rest of this file into mongosh. What each line
// printed is on the lab page, and tools/data-science/capture_lab_outputs.py
// runs it again. There is no .py half: mongomock does not implement GridFS.
// (Until October 2026 mongod could not be installed where these labs are
// checked, and this file was desk-checked only.)

// =============================================================================
// WHY GridFS EXISTS
// =============================================================================
// A BSON document is capped at 16 MB. GridFS gets round that by splitting the
// file into chunks of 255 KB and storing each chunk as its own document.
//
//   fs.files   -- ONE metadata document per file
//   fs.chunks  -- MANY documents, each a 255 KB slice, with files_id and n
//
// The 16 MB cap is not arbitrary: it is what keeps a single document cheap to
// move around in memory and over the wire. GridFS does not remove the cap; it
// works within it.

// =============================================================================
// Step 1: Put a file in GridFS with mongofiles
// 1. From the command line: mongofiles
// =============================================================================
//   mongofiles -d collegeDB put lecture.mp4
//   mongofiles -d collegeDB list
//   mongofiles -d collegeDB get lecture.mp4
//   mongofiles -d collegeDB delete lecture.mp4
//   mongofiles -d collegeDB --local ./copy.mp4 get lecture.mp4
//
// mongofiles ships with the MongoDB Database Tools, a SEPARATE download from
// the server. That trips people up.

// =============================================================================
// Step 2: Look at what it stored, and count the chunks
// 2. Looking at what it stored
// =============================================================================
use collegeDB
db.fs.files.find().pretty()
// { _id, length, chunkSize: 261120, uploadDate, filename, metadata }

db.fs.chunks.find({}, { data: 0 }).sort({ n: 1 }).limit(3)
// { _id, files_id, n: 0 }, { ..., n: 1 }, ...   -- n is the ORDER

const file = db.fs.files.findOne({ filename: "lecture.mp4" })
db.fs.chunks.countDocuments({ files_id: file._id })     // 41 for a 10 MB file
// [Corrected: this read files_id: <the _id from fs.files>, a placeholder, which
// mongosh rejects as a syntax error. The line before it looks the _id up.]

// --- THE ARITHMETIC, which is what gets asked ---------------------------------
// chunkSize is 255 KB = 255 * 1024 = 261120 bytes.
//
//   chunks = ceil(length / 261120)
//
//   1 MB     =  1048576 B -> ceil(1048576/261120)  =     5 chunks
//   10 MB    = 10485760 B -> ceil(10485760/261120) =    41 chunks
//   100 MB   = 104857600  -> ceil(104857600/261120)=   402 chunks
//   1 GB     = 1073741824 -> ceil(1073741824/261120)=  4113 chunks
//
// The LAST chunk is short -- GridFS does not pad. So the stored size is the
// file's size plus a little metadata, not a multiple of 255 KB.

// --- the indexes GridFS creates for itself -----------------------------------
db.fs.chunks.getIndexes()      // { files_id: 1, n: 1 }, UNIQUE
db.fs.files.getIndexes()       // { filename: 1, uploadDate: 1 }
// The unique compound index on (files_id, n) is what guarantees the chunks
// reassemble in the right order and cannot be duplicated.

// =============================================================================
// Step 3: Do it from a driver
// 3. From the shell / a driver
// =============================================================================
// mongosh has no built-in put; you use a driver. In Node:
//
//   const bucket = new GridFSBucket(db, { bucketName: "lectures" });
//   fs.createReadStream("lecture.mp4").pipe(bucket.openUploadStream("lecture.mp4",
//       { metadata: { course: "DSC301", week: 3 } }));
//   bucket.openDownloadStreamByName("lecture.mp4").pipe(res);
//
// A custom bucketName gives lectures.files / lectures.chunks instead of fs.*.
//
// STREAMING IS THE POINT. openDownloadStream can start at any byte:
//   bucket.openDownloadStreamByName("lecture.mp4", { start: 5_000_000 })
// which is how a video seeks without downloading the whole file first.

// =============================================================================
// Step 4: Query by metadata
// 4. Querying by metadata -- what a filesystem cannot do
// =============================================================================
db.fs.files.find({ "metadata.course": "DSC301" })
db.fs.files.find({ length: { $gt: 50 * 1024 * 1024 } })
db.fs.files.aggregate([
  { $group: { _id: "$metadata.course", n: { $sum: 1 },
              totalBytes: { $sum: "$length" } } },
  { $sort: { _id: 1 } }
])
// [Changed: the $sort was added. $group returns its groups in no fixed order.]
db.fs.files.createIndex({ "metadata.course": 1, uploadDate: -1 })

// =============================================================================
// Step 5: Delete it, with its chunks
// 5. Deleting -- the one thing to be careful about
// =============================================================================
// WRONG: this orphans every chunk of the file.
//   db.fs.files.deleteOne({ filename: "lecture.mp4" })
//
// RIGHT: use the driver's bucket.delete(id) or mongofiles delete, which
// removes the metadata document AND its chunks. GridFS is two collections
// kept consistent by the DRIVER, not by the database -- there is no cascade.

// =============================================================================
// WHEN NOT TO USE GridFS -- worth a mark, and usually the right answer
// =============================================================================
// USE IT when:
//   * files exceed 16 MB
//   * you need RANGE reads (video seeking)
//   * you want the file and its metadata under one backup and one replica set
//   * atomic-ish updates matter more than throughput
//
// DO NOT use it when:
//   * files are small and numerous -- the chunk documents cost more than they
//     save, and a plain BinData field under 16 MB is simpler
//   * you are serving them over HTTP at volume -- S3 or a CDN is faster,
//     cheaper, and does not put the read load on your database
//
// The honest summary: GridFS is a good answer when the files must live with
// the data. It is rarely the best answer for a website's static assets.

5. Execution and Results

OUTPUT

made lecture.mp4: 10,485,760 bytes

$ mongofiles -d collegeDB put lecture.mp4
(nothing printed)
$ mongofiles -d collegeDB list
lecture.mp4 10485760
$ mongofiles -d collegeDB --local ./copy.mp4 get lecture.mp4
(nothing printed)
copy.mp4 is the same as lecture.mp4, byte for byte: True

test> use collegeDB
switched to db collegeDB
collegeDB> db.fs.files.find().pretty()
[
  {
    _id: ObjectId('6ac2154da083e669e2dec979'),
    length: Long('10485760'),
    chunkSize: 261120,
    uploadDate: ISODate('2026-10-04T08:58:53.727Z'),
    filename: 'lecture.mp4',
    metadata: {}
  }
]
collegeDB> db.fs.chunks.find({}, { data: 0 }).sort({ n: 1 }).limit(3)
[
  {
    _id: ObjectId('6ac2154da083e669e2dec97a'),
    files_id: ObjectId('6ac2154da083e669e2dec979'),
    n: 0
  },
  {
    _id: ObjectId('6ac2154da083e669e2dec97b'),
    files_id: ObjectId('6ac2154da083e669e2dec979'),
    n: 1
  },
  {
    _id: ObjectId('6ac2154da083e669e2dec97c'),
    files_id: ObjectId('6ac2154da083e669e2dec979'),
    n: 2
  }
]
collegeDB> const file = db.fs.files.findOne({ filename: "lecture.mp4" })
collegeDB> db.fs.chunks.countDocuments({ files_id: file._id })     // 41 for a 10 MB file
41
collegeDB> db.fs.chunks.getIndexes()      // { files_id: 1, n: 1 }, UNIQUE
[
  { v: 2, key: { _id: 1 }, name: '_id_' },
  {
    v: 2,
    key: { files_id: 1, n: 1 },
    name: 'files_id_1_n_1',
    unique: true
  }
]
collegeDB> db.fs.files.getIndexes()       // { filename: 1, uploadDate: 1 }
[
  { v: 2, key: { _id: 1 }, name: '_id_' },
  {
    v: 2,
    key: { filename: 1, uploadDate: 1 },
    name: 'filename_1_uploadDate_1'
  }
]
collegeDB> db.fs.files.find({ "metadata.course": "DSC301" })
collegeDB> db.fs.files.find({ length: { $gt: 50 * 1024 * 1024 } })
collegeDB> db.fs.files.aggregate([
...   { $group: { _id: "$metadata.course", n: { $sum: 1 },
...               totalBytes: { $sum: "$length" } } },
...   { $sort: { _id: 1 } }
... ])
[ { _id: null, n: 1, totalBytes: Long('10485760') } ]
collegeDB> db.fs.files.createIndex({ "metadata.course": 1, uploadDate: -1 })
metadata.course_1_uploadDate_-1

$ mongofiles -d collegeDB delete lecture.mp4
(nothing printed)
after mongofiles delete: 0 files, 0 chunks -- the metadata document AND its chunks

mongofiles cannot attach metadata, so section 4's queries on metadata.course find nothing, as they would after any mongofiles put; a driver upload, as in section 3, sets it.

Corrected: the chunk count read { files_id: <the _id from fs.files> }, a placeholder that mongosh rejects as a syntax error; the line before it now looks the _id up. Changed: the metadata $group ends with a $sort.

RESULT

The 10 MB file is one document in fs.files and 41 in fs.chunks, as ceil(10485760 / 261120) says; the copy back is identical; mongofiles delete removes both.

Experiment 19 — Transactions

1. Question

Make a multi-document transaction, committed and aborted.

2. Aim

Transfer money between two accounts in a transaction, and see an abort undo everything.

3. Steps

  1. Set up two accounts.
  2. Transfer, and commit.
  3. Overdraw, and abort.
  4. Retry on a transient error.

WHAT TO DEMONSTRATE

Transactions are unavailable on a standalone mongod — they depend on the oplog and majority commit, so a replica set is required. That fact is itself a five-mark answer. This experiment runs on a one-member replica set, which is enough.

What to demonstrate: a transfer that commits, and one that aborts midway leaving both balances unchanged. And the point from Unit 5 §5.9: a schema that needs transactions for its common operations is usually one that should have embedded.

4. Programme

// Experiment 19 -- Multi-document ACID transactions.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. There is no .py half: transactions need a replica set, which mongomock
// is not. (Until October 2026 mongod could not be installed where these labs
// are checked, and this file was desk-checked only.)

// =============================================================================
// THE PREREQUISITE, which is itself a five-mark answer
// =============================================================================
// A standalone mongod REFUSES to start a transaction. In mongosh 2, against
// MongoDB 8, the first write inside one fails with:
//
//   MongoServerError: This MongoDB deployment does not support retryable
//   writes. Please add retryWrites=false to your connection string.
//
// -- and with retryWrites=false in the connection string, still with that.
// [Corrected: this quoted "Transaction numbers are only allowed on a replica
// set member or mongos", which MongoDB 8 with mongosh 2 does not print. The
// message above is what a standalone gives, tried both ways.]
//
// WHY: a transaction's commit must be durable and visible atomically, and
// MongoDB implements that on the oplog with a majority write concern. A
// standalone has no oplog to write to and no majority to reach. So: set up
// experiment 17 first, or use Atlas, whose free tier is a replica set.
//
// Single-DOCUMENT operations have always been atomic, replica set or not.
// That is the point most students miss, and it is why most well-modelled
// MongoDB applications never need this feature at all.

// =============================================================================
// Step 1: Set up two accounts
// 1. Setting up something worth a transaction
// =============================================================================
use bankDB
db.accounts.drop()
db.accounts.insertMany([
  { _id: "A", holder: "Asha", balance: 5000 },
  { _id: "B", holder: "Ravi", balance: 3000 }
])
// A transfer touches TWO documents. Nothing about the document model makes
// that atomic, and here the two balances genuinely belong to two owners --
// so this is the case where embedding is not the answer.

// =============================================================================
// Step 2: Transfer, and commit
// 2. The core API -- a transfer that COMMITS
// =============================================================================
const session = db.getMongo().startSession()
const accounts = session.getDatabase("bankDB").accounts

session.startTransaction({
  readConcern:  { level: "snapshot" },
  writeConcern: { w: "majority" }
})

try {
  accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
  accounts.updateOne({ _id: "B" }, { $inc: { balance:  500 } })
  session.commitTransaction()
  print("committed: A 4500, B 3500")
} catch (e) {
  session.abortTransaction()
  print("aborted: " + e)
  throw e
} finally {
  session.endSession()
}

// IMPORTANT: reads and writes must go through session.getDatabase(...).
// db.accounts.updateOne(...) inside the block is NOT in the transaction --
// it commits immediately, and nothing warns you. That is the single commonest
// mistake with this API.

// =============================================================================
// Step 3: Overdraw, and abort
// 3. The demonstration that matters: an ABORT leaves BOTH unchanged
// =============================================================================
// Show the balances before, run this, show them after. Nothing moved.
const s2 = db.getMongo().startSession()
const acc2 = s2.getDatabase("bankDB").accounts
s2.startTransaction()
try {
  acc2.updateOne({ _id: "A" }, { $inc: { balance: -9999 } })   // overdraws
  const a = acc2.findOne({ _id: "A" })
  if (a.balance < 0) throw new Error("insufficient funds")
  acc2.updateOne({ _id: "B" }, { $inc: { balance: 9999 } })
  s2.commitTransaction()
} catch (e) {
  s2.abortTransaction()          // A's -9999 is UNDONE
  print("aborted, both balances unchanged: " + e.message)
} finally { s2.endSession() }
db.accounts.find()       // A 4500, B 3500: the first transfer, and nothing of the second
// [Added: the balances were only printed as text above. This reads them back.]

// Read A from ANOTHER shell while the transaction is open: you see the OLD
// balance. Uncommitted writes are invisible outside the session -- snapshot
// isolation, and the visible proof that this is a real transaction.

// =============================================================================
// Step 4: Retry on a transient error
// 4. Retrying -- required, not optional
// =============================================================================
// A transaction can fail with a TRANSIENT error (a write conflict, a failover
// mid-commit). The error carries a label, and the caller is expected to retry:
//
//   e.hasErrorLabel("TransientTransactionError")   -> retry the WHOLE thing
//   e.hasErrorLabel("UnknownTransactionCommitResult") -> retry the COMMIT only
//
// Drivers wrap this: session.withTransaction(fn) retries for you and is what
// you should actually use.
//
//   session.withTransaction(() => {
//     accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
//     accounts.updateOne({ _id: "B" }, { $inc: { balance:  500 } })
//   })
//
// The callback must be IDEMPOTENT, because it may run more than once.

// =============================================================================
// 5. The limits
// =============================================================================
//   * default 60-second time limit (transactionLifetimeLimitSeconds); a
//     transaction that exceeds it is aborted
//   * 16 MB of oplog entries per transaction
//   * they hold locks, so long transactions block other writers
//   * a sharded transaction is slower again -- it coordinates across shards
//   * DDL inside a transaction is restricted: no createIndex, no dropDatabase;
//     collection creation is allowed only from MongoDB 4.4 onward

// =============================================================================
// THE POINT, from Unit 5 §5.9
// =============================================================================
// Transactions arrived in MongoDB 4.0 and are genuinely ACID. But:
//
//   IF YOUR COMMON OPERATIONS NEED THEM, YOUR SCHEMA IS PROBABLY WRONG.
//
// A student and their address, an order and its lines, a post and its
// comments -- embed those and every update is a single-document write, atomic
// with no transaction at all. Transactions are for the genuine cross-entity
// case: a bank transfer, where the two balances belong to different people and
// no amount of remodelling puts them in one document.
//
// Compare with Course 5: in SQL, transactions are how you do ordinary work.
// Here they are the escape hatch for the case the document model does not
// cover -- and reaching for it often is the signal to reconsider the model,
// or to ask whether the data was relational all along.

5. Execution and Results

OUTPUT

rs0 [direct: primary] test> use bankDB
switched to db bankDB
rs0 [direct: primary] bankDB> db.accounts.drop()
true
rs0 [direct: primary] bankDB> db.accounts.insertMany([
...   { _id: "A", holder: "Asha", balance: 5000 },
...   { _id: "B", holder: "Ravi", balance: 3000 }
... ])
{ acknowledged: true, insertedIds: { '0': 'A', '1': 'B' } }
rs0 [direct: primary] bankDB> const session = db.getMongo().startSession()
rs0 [direct: primary] bankDB> const accounts = session.getDatabase("bankDB").accounts
rs0 [direct: primary] bankDB> session.startTransaction({
...   readConcern:  { level: "snapshot" },
...   writeConcern: { w: "majority" }
... })
rs0 [direct: primary] bankDB> try {
...   accounts.updateOne({ _id: "A" }, { $inc: { balance: -500 } })
...   accounts.updateOne({ _id: "B" }, { $inc: { balance:  500 } })
...   session.commitTransaction()
...   print("committed: A 4500, B 3500")
... } catch (e) {
...   session.abortTransaction()
...   print("aborted: " + e)
...   throw e
... } finally {
...   session.endSession()
... }
committed: A 4500, B 3500
rs0 [direct: primary] bankDB> const s2 = db.getMongo().startSession()
rs0 [direct: primary] bankDB> const acc2 = s2.getDatabase("bankDB").accounts
rs0 [direct: primary] bankDB> s2.startTransaction()
rs0 [direct: primary] bankDB> try {
...   acc2.updateOne({ _id: "A" }, { $inc: { balance: -9999 } })   // overdraws
...   const a = acc2.findOne({ _id: "A" })
...   if (a.balance < 0) throw new Error("insufficient funds")
...   acc2.updateOne({ _id: "B" }, { $inc: { balance: 9999 } })
...   s2.commitTransaction()
... } catch (e) {
...   s2.abortTransaction()          // A's -9999 is UNDONE
...   print("aborted, both balances unchanged: " + e.message)
... } finally { s2.endSession() }
aborted, both balances unchanged: insufficient funds
rs0 [direct: primary] bankDB> db.accounts.find()       // A 4500, B 3500: the first transfer, and nothing of the second
[
  { _id: 'A', holder: 'Asha', balance: 4500 },
  { _id: 'B', holder: 'Ravi', balance: 3500 }
]

Corrected: the script quoted a standalone server's refusal as "Transaction numbers are only allowed on a replica set member or mongos". MongoDB 8 with mongosh 2 says instead "This MongoDB deployment does not support retryable writes. Please add retryWrites=false to your connection string" — and says it again with retryWrites=false. Added: a find() at the end reads the balances back; both were only printed as text.

RESULT

The transfer committed, A 4500 and B 3500; the overdraft aborted, and both balances are unchanged, read back from the database.

Experiment 20 — Case study: a mini-application

1. Question

Build a mini-application: a library management system.

2. Aim

Implement the library schema, issue and return books, run the reports, and check the stock counts stay true.

3. Steps

In mongosh, 20_case_study.js:

  1. Seed the books and members.
  2. Create the indexes.
  3. Issue books, with a conditional decrement.
  4. Return them, and charge the fines.
  5. Run the five reports.
  6. Check the stock counts against the loans.

In Python, through mongomock, 20_case_study.py:

  1. Seed the books, members and loans.
  2. Issue and return a book.
  3. Refuse a sixth copy of a five-copy book.
  4. Return on time and late.
  5. Run the five reports.
  6. Break the stock count, and catch it.

THE POINT

A library management system — the schema designed in practice.md Section C question 1 — exercising CRUD, aggregation and indexing together:

  1. Seed books, members and loans.
  2. Issue a book: insert a loan and decrement availableCopies conditionally (availableCopies: { $gt: 0 }), so a third copy of a two-copy book cannot be lent.

  3. Return it: set returned, compute any fine, increment the count back.

  4. Overdue report — { returned: null, due: { $lt: asAt } }.
  5. Most-borrowed — an aggregation with $match first.
  6. Check availableCopies is consistent with the count of unreturned loans.

That last check is the point of the experiment: the computed pattern speeds up the hottest query and introduces a value that can drift, and the only defence is to check it. The Python half asserts it at every step.

4. Programme

In mongosh, 20_case_study.js:

// Experiment 20 -- Case study: a library management system.
//
// Run with MongoDB 8.3.7 and mongosh 2.12.0, on a fresh server: it is typed
// into mongosh line by line, as you would at the prompt. What each line printed
// is on the lab page, and tools/data-science/capture_lab_outputs.py runs it
// again. The query logic is also executed and asserted in 20_case_study.py,
// through mongomock. (Until October 2026 mongod could not be installed where
// these labs are checked, and this file was desk-checked only.)
//
// The schema is the one designed in practice.md Section C question 1. Read
// that answer first: it justifies every embed and every reference, and this
// script only implements it.

use libraryDB
db.books.drop(); db.members.drop(); db.loans.drop()

// =============================================================================
// Step 1: Seed the books and members
// 1. SEED
// =============================================================================
db.books.insertMany([
  { _id: "978-1491954461", title: "MongoDB: The Definitive Guide",
    authors: ["Shannon Bradshaw", "Kristina Chodorow"],
    publisher: { name: "O'Reilly", year: 2019 },
    subjects: ["databases", "nosql"], totalCopies: 5, availableCopies: 5 },
  { _id: "978-0134685991", title: "Effective Java",
    authors: ["Joshua Bloch"], publisher: { name: "Addison-Wesley", year: 2018 },
    subjects: ["programming", "java"], totalCopies: 2, availableCopies: 2 },
  { _id: "978-1449355739", title: "Learning Python",
    authors: ["Mark Lutz"], publisher: { name: "O'Reilly", year: 2013 },
    subjects: ["programming", "python"], totalCopies: 3, availableCopies: 3 }
])

db.members.insertMany([
  { _id: "M2026001", name: "Asha Kumari", email: "asha@nri.ac.in",
    phones: ["9876543210"],
    address: { city: "Vijayawada", state: "AP", pin: "520010" },
    joined: ISODate("2026-07-01"), active: true, currentLoanCount: 0 },
  { _id: "M2026002", name: "Ravi Teja", email: "ravi@nri.ac.in",
    phones: ["9876500000"],
    address: { city: "Guntur", state: "AP", pin: "522002" },
    joined: ISODate("2026-07-05"), active: true, currentLoanCount: 0 }
])

// =============================================================================
// Step 2: Create the indexes
// 2. INDEXES -- practice.md Step 4
// =============================================================================
db.books.createIndex({ title: "text", authors: "text" })
db.books.createIndex({ subjects: 1 })                 // multikey
db.loans.createIndex({ member_id: 1, returned: 1 })   // query 2
db.loans.createIndex({ returned: 1, due: 1 })         // query 4, ESR
db.loans.createIndex({ isbn: 1, issued: -1 })         // query 5
db.members.createIndex({ email: 1 }, { unique: true })

// =============================================================================
// Step 3: Issue books, with a conditional decrement
// 3. ISSUE -- two writes, and a CONDITIONAL decrement
// =============================================================================
// A fixed "today", as in 20_case_study.py, so that a loan can be late.
const TODAY = ISODate("2026-08-26")
const DAY = 24 * 60 * 60 * 1000
// [Changed: TODAY was added, and issue() takes the day. Every loan was issued
// at new Date(), the moment the script ran, so no return could ever be late and
// the overdue report could never find anything.]

function issue(memberId, isbn, on = TODAY) {
  const book   = db.books.findOne({ _id: isbn })
  const member = db.members.findOne({ _id: memberId })

  // The guard is IN THE FILTER, not in an if. Checking availableCopies > 0 and
  // then decrementing is two operations, and two concurrent borrowers both
  // pass the check. This is one operation, so only one of them can match.
  const dec = db.books.updateOne({ _id: isbn, availableCopies: { $gt: 0 } },
                                 { $inc: { availableCopies: -1 } })
  if (dec.modifiedCount === 0) return { ok: false, why: "no copies available" }

  const issued = on
  const due    = new Date(issued.getTime() + 14 * 24 * 60 * 60 * 1000)
  db.loans.insertOne({
    member_id: memberId, isbn,
    book_title:  book.title,      // EXTENDED REFERENCE -- practice.md Step 3
    member_name: member.name,
    issued, due, returned: null, fine: 0 })
  db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: 1 } })
  return { ok: true }
}

issue("M2026001", "978-1491954461")
issue("M2026001", "978-0134685991")
issue("M2026002", "978-0134685991")
issue("M2026002", "978-1449355739")

// The third copy of a two-copy book:
issue("M2026001", "978-0134685991")     // -> { ok: false, why: "no copies available" }

// =============================================================================
// Step 4: Return them, and charge the fines
// 4. RETURN -- set returned, compute the fine, put the copy back
// =============================================================================
function returnBook(memberId, isbn, on) {
  const loan = db.loans.findOne({ member_id: memberId, isbn, returned: null })
  if (!loan) return { ok: false, why: "no open loan" }

  const daysLate = Math.max(0, Math.ceil((on - loan.due) / (24 * 60 * 60 * 1000)))
  const fine = daysLate * 2                      // Rs 2 per day

  db.loans.updateOne({ _id: loan._id }, { $set: { returned: on, fine } })
  db.books.updateOne({ _id: isbn }, { $inc: { availableCopies: 1 } })
  db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: -1 } })
  return { ok: true, daysLate, fine }
}

returnBook("M2026002", "978-1449355739", new Date(TODAY.getTime() + 10 * DAY))   // day 10 of 14
returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 20 * DAY))   // day 20: 6 late
returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 21 * DAY))   // again: refused
// [Changed: the return was at new Date(); the late return and the second
// return, which 20_case_study.py asserts, were added.]

// =============================================================================
// Step 5: Run the five reports
// 5. THE FIVE QUERIES -- practice.md Step 5
// =============================================================================
// 1. availability
db.books.findOne({ _id: "978-1491954461" }, { title: 1, availableCopies: 1 })

// 2. a member's current loans
db.loans.find({ member_id: "M2026001", returned: null })

// 4. overdue -- the extended reference pays for itself here: no $lookup
const asAt = new Date(TODAY.getTime() + 20 * DAY)     // overdue as at day 20
db.loans.find({ returned: null, due: { $lt: asAt } }).sort({ due: 1 })
// [Corrected: the .sort(...) began its own line. Typed into mongosh, a line that
// starts with a dot does not continue the one above -- the shell ran the
// find(), unsorted, without it, then rejected ".sort(...)" as an invalid command.]

// 5. most borrowed -- $match FIRST
db.loans.aggregate([
  { $match:  { issued: { $gte: ISODate("2026-01-01") } } },
  { $group:  { _id: "$isbn", title: { $first: "$book_title" },
               times: { $sum: 1 } } },
  { $sort:   { times: -1, _id: 1 } },
  { $limit:  10 }
])

// subject report -- multikey + $unwind
db.books.aggregate([
  { $unwind: "$subjects" },
  { $group:  { _id: "$subjects", titles: { $push: "$title" },
               n: { $sum: 1 } } },
  { $sort:   { n: -1, _id: 1 } }
])

// =============================================================================
// Step 6: Check the stock counts against the loans
// 6. THE INTEGRITY CHECK -- the point of the whole experiment
// =============================================================================
// availableCopies is the COMPUTED pattern: it makes query 1 constant-time and
// introduces a number that can DRIFT. The only defence is to check it.
db.books.aggregate([
  { $lookup: {
      from: "loans", localField: "_id", foreignField: "isbn", as: "loans" } },
  { $project: {
      title: 1, totalCopies: 1, availableCopies: 1,
      out: { $size: { $filter: { input: "$loans", as: "l",
                                 cond: { $eq: ["$$l.returned", null] } } } } } },
  { $addFields: { expected: { $subtract: ["$totalCopies", "$out"] } } },
  { $match: { $expr: { $ne: ["$availableCopies", "$expected"] } } }
])
// This should return NOTHING. Anything it returns is a book whose stored count
// disagrees with its open loans -- run it nightly, and alert on any row.

In Python, through mongomock, 20_case_study.py:

"""Experiment 20 — Case study: a library management system.

The schema is the one designed in practice.md Section C question 1, and this
runs the whole workflow against it: seed, issue, return, the five reports, and
the integrity check.

The integrity check is the point of the experiment. availableCopies is the
COMPUTED pattern -- it makes the availability lookup constant-time and
introduces a number that can drift out of step with the loans. So it is
asserted after every single write, not once at the end. When the last section
deliberately breaks it, the same check catches it.
"""
import datetime as dt

import mongomock

# This experiment does NOT use fixtures.py. The other nineteen share the
# collegeDB sample data; this one designs its own schema from scratch, which
# is the exercise.

DAY = dt.timedelta(days=1)
LOAN_DAYS = 14
FINE_PER_DAY = 2

BOOKS = [
    {"_id": "978-1491954461", "title": "MongoDB: The Definitive Guide",
     "authors": ["Shannon Bradshaw", "Kristina Chodorow"],
     "publisher": {"name": "O'Reilly", "year": 2019},
     "subjects": ["databases", "nosql"], "totalCopies": 5, "availableCopies": 5},
    {"_id": "978-0134685991", "title": "Effective Java",
     "authors": ["Joshua Bloch"],
     "publisher": {"name": "Addison-Wesley", "year": 2018},
     "subjects": ["programming", "java"], "totalCopies": 2, "availableCopies": 2},
    {"_id": "978-1449355739", "title": "Learning Python",
     "authors": ["Mark Lutz"], "publisher": {"name": "O'Reilly", "year": 2013},
     "subjects": ["programming", "python"], "totalCopies": 3, "availableCopies": 3},
]

MEMBERS = [
    {"_id": "M2026001", "name": "Asha Kumari", "email": "asha@nri.ac.in",
     "phones": ["9876543210"],
     "address": {"city": "Vijayawada", "state": "AP", "pin": "520010"},
     "joined": dt.datetime(2026, 7, 1), "active": True, "currentLoanCount": 0},
    {"_id": "M2026002", "name": "Ravi Teja", "email": "ravi@nri.ac.in",
     "phones": ["9876500000"],
     "address": {"city": "Guntur", "state": "AP", "pin": "522002"},
     "joined": dt.datetime(2026, 7, 5), "active": True, "currentLoanCount": 0},
]

# A fixed "today", so the overdue report is reproducible instead of drifting.
TODAY = dt.datetime(2026, 8, 26)


def seed():
    db = mongomock.MongoClient().libraryDB
    db.books.insert_many([dict(b) for b in BOOKS])
    db.members.insert_many([dict(m) for m in MEMBERS])
    db.books.create_index("subjects")
    db.loans.create_index([("member_id", 1), ("returned", 1)])
    db.loans.create_index([("returned", 1), ("due", 1)])
    db.loans.create_index([("isbn", 1), ("issued", -1)])
    db.members.create_index("email", unique=True)
    return db


# =============================================================================
# The two operations
# =============================================================================

def issue(db, member_id, isbn, on=TODAY):
    """Insert a loan and decrement the stored count -- CONDITIONALLY."""
    book = db.books.find_one({"_id": isbn})
    member = db.members.find_one({"_id": member_id})
    if book is None or member is None:
        return {"ok": False, "why": "unknown book or member"}

    # The guard lives IN THE FILTER. Reading availableCopies, testing it, then
    # decrementing is two operations: two concurrent borrowers both pass the
    # test and both decrement, and the count goes negative. One operation
    # cannot -- only one of them matches.
    dec = db.books.update_one({"_id": isbn, "availableCopies": {"$gt": 0}},
                              {"$inc": {"availableCopies": -1}})
    if dec.modified_count == 0:
        return {"ok": False, "why": "no copies available"}

    db.loans.insert_one({
        "member_id": member_id, "isbn": isbn,
        "book_title": book["title"],        # extended reference
        "member_name": member["name"],      # extended reference
        "issued": on, "due": on + LOAN_DAYS * DAY,
        "returned": None, "fine": 0})
    db.members.update_one({"_id": member_id},
                          {"$inc": {"currentLoanCount": 1}})
    return {"ok": True}


def return_book(db, member_id, isbn, on=TODAY):
    loan = db.loans.find_one({"member_id": member_id, "isbn": isbn,
                              "returned": None})
    if loan is None:
        return {"ok": False, "why": "no open loan"}

    days_late = max(0, (on - loan["due"]).days)
    fine = days_late * FINE_PER_DAY

    db.loans.update_one({"_id": loan["_id"]},
                        {"$set": {"returned": on, "fine": fine}})
    db.books.update_one({"_id": isbn}, {"$inc": {"availableCopies": 1}})
    db.members.update_one({"_id": member_id},
                          {"$inc": {"currentLoanCount": -1}})
    return {"ok": True, "daysLate": days_late, "fine": fine}


# =============================================================================
# The integrity check -- run after EVERY write
# =============================================================================

def drifted(db):
    """Books whose stored availableCopies disagrees with their open loans.

    Empty is the only acceptable answer. This is the whole reason the computed
    pattern is safe to use: it is cheap to verify, so verify it.
    """
    bad = []
    for b in db.books.find():
        out = db.loans.count_documents({"isbn": b["_id"], "returned": None})
        expected = b["totalCopies"] - out
        if b["availableCopies"] != expected:
            bad.append({"isbn": b["_id"], "title": b["title"],
                        "stored": b["availableCopies"], "expected": expected,
                        "openLoans": out})
    return bad


def members_drifted(db):
    """The same check for currentLoanCount -- a second computed field."""
    bad = []
    for m in db.members.find():
        out = db.loans.count_documents({"member_id": m["_id"], "returned": None})
        if m["currentLoanCount"] != out:
            bad.append({"member": m["_id"], "stored": m["currentLoanCount"],
                        "expected": out})
    return bad


def consistent(db, step):
    assert drifted(db) == [], (step, drifted(db))
    assert members_drifted(db) == [], (step, members_drifted(db))


# =============================================================================
# The workflow
# =============================================================================

def the_happy_path(db):
    consistent(db, "after seeding")

    assert issue(db, "M2026001", "978-1491954461")["ok"]
    consistent(db, "after issue 1")
    assert issue(db, "M2026001", "978-0134685991")["ok"]
    consistent(db, "after issue 2")
    assert issue(db, "M2026002", "978-0134685991")["ok"]
    consistent(db, "after issue 3")
    assert issue(db, "M2026002", "978-1449355739")["ok"]
    consistent(db, "after issue 4")

    assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 0
    assert db.members.find_one({"_id": "M2026001"})["currentLoanCount"] == 2
    assert db.loans.count_documents({"returned": None}) == 4

    print("  4 issues, and availableCopies / currentLoanCount agree with the")
    print("  loans collection after every single one:")
    for b in db.books.find().sort("_id", 1):
        print(f"    {b['title']:34s} {b['availableCopies']}/{b['totalCopies']} available")


def the_sixth_copy_of_a_five_copy_book(db):
    """The conditional decrement, and why it is not an if statement."""
    # Effective Java has 2 copies and both are out.
    before = db.books.find_one({"_id": "978-0134685991"})["availableCopies"]
    assert before == 0
    result = issue(db, "M2026001", "978-0134685991")

    assert result == {"ok": False, "why": "no copies available"}, result
    assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 0, \
        "the count must not go NEGATIVE"
    assert db.loans.count_documents({"isbn": "978-0134685991"}) == 2, \
        "and no loan row was written for the refused issue"
    consistent(db, "after a refused issue")

    print("  a third issue of a 2-copy book -> refused, count stayed at 0,")
    print("  no loan row written")
    print("       the guard is { _id: isbn, availableCopies: { $gt: 0 } } in the")
    print("       FILTER. As an if-then-decrement it is two operations, and two")
    print("       concurrent borrowers both pass the test. As one update, they")
    print("       cannot: the second one matches nothing")


def returning_on_time_and_late(db):
    # On time: issued TODAY, due TODAY+14, returned TODAY+10.
    on_time = return_book(db, "M2026002", "978-1449355739", TODAY + 10 * DAY)
    assert on_time == {"ok": True, "daysLate": 0, "fine": 0}, on_time
    consistent(db, "after an on-time return")
    assert db.books.find_one({"_id": "978-1449355739"})["availableCopies"] == 3

    # Late: returned TODAY+20, six days past the due date.
    late = return_book(db, "M2026002", "978-0134685991", TODAY + 20 * DAY)
    assert late == {"ok": True, "daysLate": 6, "fine": 12}, late
    assert 6 * FINE_PER_DAY == 12
    consistent(db, "after a late return")
    assert db.books.find_one({"_id": "978-0134685991"})["availableCopies"] == 1

    # Returning something not on loan changes nothing.
    again = return_book(db, "M2026002", "978-0134685991", TODAY + 21 * DAY)
    assert again == {"ok": False, "why": "no open loan"}, again
    consistent(db, "after a duplicate return")

    print("  returned day 10 of a 14-day loan -> 0 days late, fine 0")
    print("  returned day 20 of a 14-day loan -> 6 days late, fine Rs 12")
    print("  returning it a second time       -> refused, nothing changed")
    print("       that last one matters: without the returned: null in the")
    print("       filter, a double return increments availableCopies twice and")
    print("       the library thinks it owns a copy it does not have")


def the_five_reports(db):
    # 1. availability -- one document, no join
    b = db.books.find_one({"_id": "978-1491954461"},
                          {"title": 1, "availableCopies": 1})
    assert b == {"_id": "978-1491954461",
                 "title": "MongoDB: The Definitive Guide",
                 "availableCopies": 4}, b

    # 2. a member's open loans
    asha = list(db.loans.find({"member_id": "M2026001", "returned": None}))
    assert len(asha) == 2
    assert {l["isbn"] for l in asha} == {"978-1491954461", "978-0134685991"}

    # 4. overdue, as at TODAY + 20 days
    as_at = TODAY + 20 * DAY
    overdue = list(db.loans.find({"returned": None, "due": {"$lt": as_at}})
                   .sort("due", 1))
    assert len(overdue) == 2, overdue
    # The extended reference is what makes this report joinless.
    assert all("book_title" in l and "member_name" in l for l in overdue)
    assert {l["member_name"] for l in overdue} == {"Asha Kumari"}

    # 5. most borrowed
    top = list(db.loans.aggregate([
        {"$match": {"issued": {"$gte": dt.datetime(2026, 1, 1)}}},
        {"$group": {"_id": "$isbn", "title": {"$first": "$book_title"},
                    "times": {"$sum": 1}}},
        {"$sort": {"times": -1, "_id": 1}},
        {"$limit": 10}]))
    assert [(t["title"], t["times"]) for t in top] == [
        ("Effective Java", 2),
        ("Learning Python", 1),
        ("MongoDB: The Definitive Guide", 1)], top

    # subjects -- multikey plus $unwind
    subs = list(db.books.aggregate([
        {"$unwind": "$subjects"},
        {"$group": {"_id": "$subjects", "n": {"$sum": 1}}},
        {"$sort": {"n": -1, "_id": 1}}]))
    assert {s["_id"]: s["n"] for s in subs} == \
        {"programming": 2, "databases": 1, "java": 1, "nosql": 1, "python": 1}

    print(f"  1. availability   MongoDB Definitive Guide {b['availableCopies']}/5")
    print(f"  2. Asha's loans   {len(asha)} open")
    print(f"  4. overdue at {as_at.date()}  {len(overdue)}, both Asha's, no $lookup")
    print("  5. most borrowed  " +
          ", ".join(f"{t['title'].split(':')[0]} x{t['times']}" for t in top))
    print("       report 4 reads ONE collection because book_title and")
    print("       member_name were copied onto the loan -- the extended")
    print("       reference pattern paying for itself")


def when_it_drifts_the_check_catches_it(db):
    """Break it on purpose. A check that has never failed is not a check."""
    assert drifted(db) == []

    # Exactly the bug the transaction in practice.md Step 5 prevents: the loan
    # was written and the decrement was not.
    db.loans.insert_one({"member_id": "M2026002", "isbn": "978-1491954461",
                         "book_title": "MongoDB: The Definitive Guide",
                         "member_name": "Ravi Teja",
                         "issued": TODAY, "due": TODAY + LOAN_DAYS * DAY,
                         "returned": None, "fine": 0})

    bad = drifted(db)
    assert len(bad) == 1, bad
    assert bad[0]["isbn"] == "978-1491954461"
    assert bad[0]["stored"] == 4 and bad[0]["expected"] == 3, bad
    assert members_drifted(db) == [{"member": "M2026002",
                                    "stored": 0, "expected": 1}], \
        members_drifted(db)

    print("  a loan written with the decrement missing -- a half-done issue:")
    print(f"    {bad[0]['title']}: stored {bad[0]['stored']}, "
          f"expected {bad[0]['expected']} ({bad[0]['openLoans']} open loans)")
    print("    M2026002: currentLoanCount 0, expected 1")
    print("       nothing errored. The reports still ran. Only the check found")
    print("       it -- which is why it runs nightly, and why practice.md puts")
    print("       the two writes in a TRANSACTION in the first place")

    # Repair, and confirm.
    db.books.update_one({"_id": "978-1491954461"},
                        {"$inc": {"availableCopies": -1}})
    db.members.update_one({"_id": "M2026002"},
                          {"$inc": {"currentLoanCount": 1}})
    consistent(db, "after repair")
    print("  repaired, and both checks are clean again")


def main():
    print("Experiment 20 -- Library management case study")
    print("  schema: practice.md Section C question 1")
    # ONE database, carried through the whole workflow -- each stage builds on
    # the state the last one left, exactly as the real application would.
    # Step 1: Seed the books, members and loans
    db = seed()
    # Step 2: Issue and return a book
    the_happy_path(db)
    # Step 3: Refuse a sixth copy of a five-copy book
    the_sixth_copy_of_a_five_copy_book(db)
    # Step 4: Return on time and late
    returning_on_time_and_late(db)
    # Step 5: Run the five reports
    the_five_reports(db)
    # Step 6: Break the stock count, and catch it
    when_it_drifts_the_check_catches_it(db)


if __name__ == "__main__":
    main()

5. Execution and Results

In mongosh, 20_case_study.js:

OUTPUT

test> use libraryDB
switched to db libraryDB
libraryDB> db.books.drop(); db.members.drop(); db.loans.drop()
true
libraryDB> db.books.insertMany([
...   { _id: "978-1491954461", title: "MongoDB: The Definitive Guide",
...     authors: ["Shannon Bradshaw", "Kristina Chodorow"],
...     publisher: { name: "O'Reilly", year: 2019 },
...     subjects: ["databases", "nosql"], totalCopies: 5, availableCopies: 5 },
...   { _id: "978-0134685991", title: "Effective Java",
...     authors: ["Joshua Bloch"], publisher: { name: "Addison-Wesley", year: 2018 },
...     subjects: ["programming", "java"], totalCopies: 2, availableCopies: 2 },
...   { _id: "978-1449355739", title: "Learning Python",
...     authors: ["Mark Lutz"], publisher: { name: "O'Reilly", year: 2013 },
...     subjects: ["programming", "python"], totalCopies: 3, availableCopies: 3 }
... ])
{
  acknowledged: true,
  insertedIds: { '0': '978-1491954461', '1': '978-0134685991', '2': '978-1449355739' }
}
libraryDB> db.members.insertMany([
...   { _id: "M2026001", name: "Asha Kumari", email: "asha@nri.ac.in",
...     phones: ["9876543210"],
...     address: { city: "Vijayawada", state: "AP", pin: "520010" },
...     joined: ISODate("2026-07-01"), active: true, currentLoanCount: 0 },
...   { _id: "M2026002", name: "Ravi Teja", email: "ravi@nri.ac.in",
...     phones: ["9876500000"],
...     address: { city: "Guntur", state: "AP", pin: "522002" },
...     joined: ISODate("2026-07-05"), active: true, currentLoanCount: 0 }
... ])
{
  acknowledged: true,
  insertedIds: { '0': 'M2026001', '1': 'M2026002' }
}
libraryDB> db.books.createIndex({ title: "text", authors: "text" })
title_text_authors_text
libraryDB> db.books.createIndex({ subjects: 1 })                 // multikey
subjects_1
libraryDB> db.loans.createIndex({ member_id: 1, returned: 1 })   // query 2
member_id_1_returned_1
libraryDB> db.loans.createIndex({ returned: 1, due: 1 })         // query 4, ESR
returned_1_due_1
libraryDB> db.loans.createIndex({ isbn: 1, issued: -1 })         // query 5
isbn_1_issued_-1
libraryDB> db.members.createIndex({ email: 1 }, { unique: true })
email_1
libraryDB> const TODAY = ISODate("2026-08-26")
libraryDB> const DAY = 24 * 60 * 60 * 1000
libraryDB> function issue(memberId, isbn, on = TODAY) {
...   const book   = db.books.findOne({ _id: isbn })
...   const member = db.members.findOne({ _id: memberId })
...
...   // The guard is IN THE FILTER, not in an if. Checking availableCopies > 0 and
...   // then decrementing is two operations, and two concurrent borrowers both
...   // pass the check. This is one operation, so only one of them can match.
...   const dec = db.books.updateOne({ _id: isbn, availableCopies: { $gt: 0 } },
...                                  { $inc: { availableCopies: -1 } })
...   if (dec.modifiedCount === 0) return { ok: false, why: "no copies available" }
...
...   const issued = on
...   const due    = new Date(issued.getTime() + 14 * 24 * 60 * 60 * 1000)
...   db.loans.insertOne({
...     member_id: memberId, isbn,
...     book_title:  book.title,      // EXTENDED REFERENCE -- practice.md Step 3
...     member_name: member.name,
...     issued, due, returned: null, fine: 0 })
...   db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: 1 } })
...   return { ok: true }
... }
[Function: issue]
libraryDB> issue("M2026001", "978-1491954461")
{ ok: true }
libraryDB> issue("M2026001", "978-0134685991")
{ ok: true }
libraryDB> issue("M2026002", "978-0134685991")
{ ok: true }
libraryDB> issue("M2026002", "978-1449355739")
{ ok: true }
libraryDB> issue("M2026001", "978-0134685991")     // -> { ok: false, why: "no copies available" }
{ ok: false, why: 'no copies available' }
libraryDB> function returnBook(memberId, isbn, on) {
...   const loan = db.loans.findOne({ member_id: memberId, isbn, returned: null })
...   if (!loan) return { ok: false, why: "no open loan" }
...
...   const daysLate = Math.max(0, Math.ceil((on - loan.due) / (24 * 60 * 60 * 1000)))
...   const fine = daysLate * 2                      // Rs 2 per day
...
...   db.loans.updateOne({ _id: loan._id }, { $set: { returned: on, fine } })
...   db.books.updateOne({ _id: isbn }, { $inc: { availableCopies: 1 } })
...   db.members.updateOne({ _id: memberId }, { $inc: { currentLoanCount: -1 } })
...   return { ok: true, daysLate, fine }
... }
[Function: returnBook]
libraryDB> returnBook("M2026002", "978-1449355739", new Date(TODAY.getTime() + 10 * DAY))   // day 10 of 14
{ ok: true, daysLate: 0, fine: 0 }
libraryDB> returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 20 * DAY))   // day 20: 6 late
{ ok: true, daysLate: 6, fine: 12 }
libraryDB> returnBook("M2026002", "978-0134685991", new Date(TODAY.getTime() + 21 * DAY))   // again: refused
{ ok: false, why: 'no open loan' }
libraryDB> db.books.findOne({ _id: "978-1491954461" }, { title: 1, availableCopies: 1 })
{
  _id: '978-1491954461',
  title: 'MongoDB: The Definitive Guide',
  availableCopies: 4
}
libraryDB> db.loans.find({ member_id: "M2026001", returned: null })
[
  {
    _id: ObjectId('6ac2155e99e081c7eead4d4c'),
    member_id: 'M2026001',
    isbn: '978-1491954461',
    book_title: 'MongoDB: The Definitive Guide',
    member_name: 'Asha Kumari',
    issued: ISODate('2026-08-26T00:00:00.000Z'),
    due: ISODate('2026-09-09T00:00:00.000Z'),
    returned: null,
    fine: 0
  },
  {
    _id: ObjectId('6ac2155e99e081c7eead4d4d'),
    member_id: 'M2026001',
    isbn: '978-0134685991',
    book_title: 'Effective Java',
    member_name: 'Asha Kumari',
    issued: ISODate('2026-08-26T00:00:00.000Z'),
    due: ISODate('2026-09-09T00:00:00.000Z'),
    returned: null,
    fine: 0
  }
]
libraryDB> const asAt = new Date(TODAY.getTime() + 20 * DAY)     // overdue as at day 20
libraryDB> db.loans.find({ returned: null, due: { $lt: asAt } }).sort({ due: 1 })
[
  {
    _id: ObjectId('6ac2155e99e081c7eead4d4c'),
    member_id: 'M2026001',
    isbn: '978-1491954461',
    book_title: 'MongoDB: The Definitive Guide',
    member_name: 'Asha Kumari',
    issued: ISODate('2026-08-26T00:00:00.000Z'),
    due: ISODate('2026-09-09T00:00:00.000Z'),
    returned: null,
    fine: 0
  },
  {
    _id: ObjectId('6ac2155e99e081c7eead4d4d'),
    member_id: 'M2026001',
    isbn: '978-0134685991',
    book_title: 'Effective Java',
    member_name: 'Asha Kumari',
    issued: ISODate('2026-08-26T00:00:00.000Z'),
    due: ISODate('2026-09-09T00:00:00.000Z'),
    returned: null,
    fine: 0
  }
]
libraryDB> db.loans.aggregate([
...   { $match:  { issued: { $gte: ISODate("2026-01-01") } } },
...   { $group:  { _id: "$isbn", title: { $first: "$book_title" },
...                times: { $sum: 1 } } },
...   { $sort:   { times: -1, _id: 1 } },
...   { $limit:  10 }
... ])
[
  { _id: '978-0134685991', title: 'Effective Java', times: 2 },
  { _id: '978-1449355739', title: 'Learning Python', times: 1 },
  {
    _id: '978-1491954461',
    title: 'MongoDB: The Definitive Guide',
    times: 1
  }
]
libraryDB> db.books.aggregate([
...   { $unwind: "$subjects" },
...   { $group:  { _id: "$subjects", titles: { $push: "$title" },
...                n: { $sum: 1 } } },
...   { $sort:   { n: -1, _id: 1 } }
... ])
[
  {
    _id: 'programming',
    titles: [ 'Effective Java', 'Learning Python' ],
    n: 2
  },
  { _id: 'databases', titles: [ 'MongoDB: The Definitive Guide' ], n: 1 },
  { _id: 'java', titles: [ 'Effective Java' ], n: 1 },
  { _id: 'nosql', titles: [ 'MongoDB: The Definitive Guide' ], n: 1 },
  { _id: 'python', titles: [ 'Learning Python' ], n: 1 }
]
libraryDB> db.books.aggregate([
...   { $lookup: {
...       from: "loans", localField: "_id", foreignField: "isbn", as: "loans" } },
...   { $project: {
...       title: 1, totalCopies: 1, availableCopies: 1,
...       out: { $size: { $filter: { input: "$loans", as: "l",
...                                  cond: { $eq: ["$$l.returned", null] } } } } } },
...   { $addFields: { expected: { $subtract: ["$totalCopies", "$out"] } } },
...   { $match: { $expr: { $ne: ["$availableCopies", "$expected"] } } }
... ])

In Python, through mongomock, 20_case_study.py:

OUTPUT

Experiment 20 -- Library management case study
  schema: practice.md Section C question 1
  4 issues, and availableCopies / currentLoanCount agree with the
  loans collection after every single one:
    Effective Java                     0/2 available
    Learning Python                    2/3 available
    MongoDB: The Definitive Guide      4/5 available
  a third issue of a 2-copy book -> refused, count stayed at 0,
  no loan row written
       the guard is { _id: isbn, availableCopies: { $gt: 0 } } in the
       FILTER. As an if-then-decrement it is two operations, and two
       concurrent borrowers both pass the test. As one update, they
       cannot: the second one matches nothing
  returned day 10 of a 14-day loan -> 0 days late, fine 0
  returned day 20 of a 14-day loan -> 6 days late, fine Rs 12
  returning it a second time       -> refused, nothing changed
       that last one matters: without the returned: null in the
       filter, a double return increments availableCopies twice and
       the library thinks it owns a copy it does not have
  1. availability   MongoDB Definitive Guide 4/5
  2. Asha's loans   2 open
  4. overdue at 2026-09-15  2, both Asha's, no $lookup
  5. most borrowed  Effective Java x2, Learning Python x1, MongoDB x1
       report 4 reads ONE collection because book_title and
       member_name were copied onto the loan -- the extended
       reference pattern paying for itself
  a loan written with the decrement missing -- a half-done issue:
    MongoDB: The Definitive Guide: stored 4, expected 3 (2 open loans)
    M2026002: currentLoanCount 0, expected 1
       nothing errored. The reports still ran. Only the check found
       it -- which is why it runs nightly, and why practice.md puts
       the two writes in a TRANSACTION in the first place
  repaired, and both checks are clean again

Changed: loans were issued at new Date(), the moment the script ran, so no return could be late and the overdue report could never find anything. The script now issues on a fixed day, 26 August 2026, as the Python half does, and returns one book on time and one six days late. Corrected: the overdue report's .sort(...) began its own line, which mongosh rejected. And this page said the refused loan was "the sixth copy of a five-copy book"; it is the third of Effective Java's two copies.

RESULT

A third loan of a two-copy book is refused; returns on day 10 and day 20 are fined Rs 0 and Rs 12; two loans are overdue at day 20; the integrity check finds no drift.


Lab examination

An hour, a dataset, one experiment number, then a viva.

What costs marks:

What earns them:

Each program, on its own page

The same experiments, one page each, so a program can be reached by what it does rather than by its number.

RUNS

Databases, collections, inserting documents in MongoDB

RUNS

Find() and the comparison operators in MongoDB

RUNS

Logical operators in MongoDB

RUNS

Update operators in MongoDB

RUNS

Deleting documents in MongoDB

RUNS

Projection in MongoDB

RUNS

Sorting, limiting and skipping in MongoDB

RUNS

An embedded data model in MongoDB

RUNS

A normalized model using document references in MongoDB

RUNS

One-to-one, one-to-many and many-to-many in MongoDB

RUNS

Schema validation with JSON Schema in MongoDB

RUNS

Single-field and compound indexes in MongoDB

RUNS

Text search and multikey indexes in MongoDB

RUNS

$match, $group, $project, $sort in MongoDB

RUNS

$lookup, $unwind and $bucket in MongoDB

RUNS

Case study: a library management system in MongoDB