Skip to the content
On this page
  1. Used by several courses
  2. Course 1 — Office Automation
  3. Course 2 — Problem Solving Using C
  4. Course 3 — Python Programming
  5. Course 4 — Statistical Foundations
  6. Course 5 — Database Management Systems
  7. Course 6 — Data Science with R
  8. Course 7 — Web Technologies
  9. Course 8 — Data Mining
  10. Course 9 — Python for Data Analysis
  11. Course 10 — Document Oriented Database
  12. Course 11 — Business Intelligence Tools
  13. Course 12 A — Machine Learning
  14. Course 12 B — Big Data Technologies
  15. Course 13 A — Artificial Intelligence
  16. Course 13 B — Cloud Computing
  17. Course 14 A — Deep Learning
  18. Course 14 B — Time Series
  19. Course 15 A — Natural Language Processing
  20. Course 15 B — Data Engineering and MLOps

One CSV per method, or close to it. Every file was generated from a known truth — the regression file from a slope of 6.0, the AR(2) series from phi = (0.6, −0.3), the three clusters from centres the generator chose — so you can score your answer, not merely produce one.

tools/check_datasets.py reads every file back off disk and recovers its planted truth. A dataset whose right answer nobody has checked is worse than no dataset, because a wrong answer then looks like a lesson.

python3 tools/make_datasets.py    # regenerate (deterministic)
python3 tools/check_datasets.py   # prove each truth is recoverable

50 datasets. Seeded, so regenerating gives byte-identical files; a diff after regenerating means something changed that should not have.


Used by several courses

data/shared/flowers.csv

90 rows · species · sepal_length · sepal_width · petal_length · petal_width

Practise: k-NN; decision trees (ID3, C4.5, CART); Naive Bayes; k-Means; hierarchical clustering; PCA; train/test split.

Built from three known centres. 'alba' separates cleanly; the other two overlap on purpose, so a perfect score means a leak.

What it was built from

data/shared/sales-transactions.csv

9 rows · product · region · date · quantity · unit_price · revenue

Practise: pivot tables; GROUP BY; MapReduce; DAX measures; Hive and Spark aggregation; OLAP roll-up.

The same nine rows Courses 1, 8, 9, 11, 12 B and 15 B all analyse. Six different engines reach the same South total, which is only meaningful because they read the same rows.

What it was built from

Course 1 — Office Automation

data/course-1-office/budget.csv

7 rows · category · amount

Practise: Goal Seek; Scenario Manager; one-variable data table.

A savings RATE of 30% is not linear in income: the answer is 33000/0.70, which is not a figure you can read off the sheet.

What it was built from

data/course-1-office/class-results.csv

20 rows · roll · name · maths · physics · chemistry · english · computers

Practise: IF / nested IF / IFS; AND, OR, IFERROR; MIN, MAX, COUNTIF; descriptive statistics.

Grade on the AVERAGE. Point the formula at the total and 19 of the 20 get an A, including the student who failed all five papers.

What it was built from

data/course-1-office/payroll.csv

6 rows · name · emp_id · department · basic_pay

Practise: SUM, AVERAGE, absolute references; VLOOKUP / XLOOKUP / INDEX+MATCH; conditional formatting.

Deduction is 10% of (Basic + DA), not of Basic -- which is why Net is 1.32 x Basic and not 1.35 x. Check one row and you have checked the sheet.

What it was built from

Course 2 — Problem Solving Using C

data/course-2-c/employee-records.csv

10 rows · emp_no · name · salary · years · department

Practise: struct and array of structs; fgets, sscanf, strtok; fopen / fscanf / fprintf; string functions (strlen, strcpy, strcmp); sorting an array of structs; linear and binary search.

Ten records for the file-handling and structure practicals. The names contain spaces on purpose: scanf("%s") reads 'Anitha' and leaves 'Rao' in the buffer, which is the bug every student writes once.

What it was built from

Course 3 — Python Programming

data/course-3-python/students.csv

25 rows · roll · name · python · maths · statistics

Practise: file handling; the csv module; dictionaries; list comprehensions; exception handling.

Read it with csv.DictReader, total per student, and handle a missing file with try/except -- the three things practical 11 asks for.

What it was built from

Course 4 — Statistical Foundations

data/course-4-stats/before-after.csv

20 rows · subject · before · after

Practise: paired t-test; one-sample t-test on the differences; Wilcoxon signed-rank test.

The pairing is the point: run an INDEPENDENT t-test on the same two columns and watch the evidence weaken, because the between-subject variation is no longer removed.

What it was built from

data/course-4-stats/fertiliser-yield.csv

36 rows · fertiliser · yield

Practise: one-way ANOVA; the F distribution; post-hoc comparison; CRD in Design of Experiments.

Three fertilisers, twelve plots each. A and B are close; C is clearly higher. ANOVA says 'not all equal' -- it does not say WHICH, which is why the post-hoc test exists.

What it was built from

data/course-4-stats/heights.csv

60 rows · student_id · height_cm

Practise: mean, median, mode; variance and standard deviation; skewness and kurtosis; the normal distribution; one-sample t-test against 165.

Drawn from N(165, 8). The sample mean will not be exactly 165 -- the gap between it and the population value IS the sampling error the course is about.

What it was built from

data/course-4-stats/preference-survey.csv

140 rows · gender · preference

Practise: chi-square test of independence; contingency tables; expected frequencies; Yates' correction.

One row per respondent, so you must build the contingency table yourself first -- which is the half of the question students skip.

What it was built from

data/course-4-stats/study-hours-marks.csv

40 rows · hours · marks

Practise: scatter plot; Karl Pearson's correlation coefficient; least-squares regression; the two regression lines; coefficient of determination.

Built from marks = 12 + 6 x hours + noise. Fit it and you should recover a slope near 6 -- and the two regression lines (y on x, x on y) will NOT coincide.

What it was built from

data/course-4-stats/treatment-groups.csv

50 rows · group · score

Practise: independent two-sample t-test; F-test for equal variances; confidence interval for a difference of means; Mann-Whitney U (non-parametric).

A real difference of 5 marks with sd 6 and n=25 per group. The test SHOULD reject -- if yours does not, check which tail you used.

What it was built from

Course 5 — Database Management Systems

data/course-5-dbms/assignments.csv

7 rows · emp_id · project_id · hours_per_week

Practise: many-to-many resolution; composite keys; EXISTS / NOT EXISTS; division queries.

The junction table. 'Which employees work on NO project?' is the NOT EXISTS question, and E106 is the answer.

What it was built from

data/course-5-dbms/departments.csv

4 rows · dept_id · dept_name · city

Practise: CREATE TABLE; primary keys; SELECT ... WHERE.

The one-side of the one-to-many with employees.

What it was built from

data/course-5-dbms/employees.csv

7 rows · emp_id · name · dept_id · salary · hired · manager_id

Practise: INNER / LEFT / RIGHT / FULL JOIN; self join; GROUP BY with HAVING; subqueries; referential integrity.

manager_id is a self-referencing foreign key and two rows are NULL. An INNER self-join loses those two; a LEFT join keeps them -- that difference is the exam question.

What it was built from

data/course-5-dbms/projects.csv

4 rows · project_id · project_name · dept_id · budget

Practise: aggregate functions; ORDER BY; correlated subqueries.

Every project belongs to a department, so a three-table join runs employees -> departments -> projects.

What it was built from

data/course-5-dbms/unnormalised-orders.csv

4 rows · order_id · order_date · customer_name · customer_city · customer_phone · items · quantities · unit_prices

Practise: 1NF, 2NF, 3NF, BCNF; functional dependencies; decomposition; update, insert and delete anomalies.

Normalise it to 3NF and count the tables. Then change Anitha's phone number in the ORIGINAL file and see how many rows you have to touch -- that is the update anomaly, not a definition.

What it was built from

Course 6 — Data Science with R

data/course-6-r/car-mileage.csv

50 rows · car_id · mpg · weight_t · cylinders · transmission · service_months

Practise: data frames and factors; read.csv and str(); is.na / na.omit; lm() multiple regression; aggregate and tapply; dplyr verbs; ggplot2 scatter with a fitted line.

Fit mpg ~ weight_t + cylinders and you should recover about -7.5 and -0.8. Three service_months are blank on purpose: read it without na.strings and R will make the whole column a factor.

What it was built from

Course 7 — Web Technologies

data/course-7-web/products.csv

8 rows · sku · name · category · price · stock · rating · status

Practise: rendering a table from JSON; the Fetch API; array filter / map / reduce; sorting a table by column; form validation against a list.

Convert it to JSON, render it as a table, then filter by category and sort by price -- experiments 14 and 16 in one file. Two rows are out of stock, so your filter has something to remove.

What it was built from

Course 8 — Data Mining

data/course-8-datamining/cluster-points.csv

85 rows · x · y · true_cluster

Practise: k-Means and the elbow method; k-Medoids; DBSCAN; hierarchical clustering and dendrograms; silhouette score; BIRCH.

true_cluster is the answer key -- drop it before you cluster, then score against it. The ten rows labelled -1 are noise: k-Means cannot say so, DBSCAN can.

What it was built from

data/course-8-datamining/market-basket.csv

30 rows · transaction_id · item

Practise: Apriori; FP-Growth; support, confidence and lift; candidate generation and pruning; Partition and DIC.

Twelve baskets, one strong rule. Compute the support of every 1-itemset by hand first -- Apriori's whole trick is that it never counts a 2-itemset whose halves failed.

What it was built from

data/course-8-datamining/warehouse-facts.csv

144 rows · month · region · city · category · product · quantity · revenue

Practise: star schema; roll-up and drill-down; slice and dice; pivot; OLAP cube operations; measures against dimensions.

State the grain before you aggregate anything. Roll up city -> region -> all and the totals must agree at every level; if they do not, you have double-counted a join.

What it was built from

Course 9 — Python for Data Analysis

data/course-9-python-da/messy-customers.csv

12 rows · customer_id · name · email · city · age · salary · joined

Practise: isnull and sum; dropna against fillna; drop_duplicates; str.strip, str.lower, str.contains; astype and to_datetime; IQR and z-score outlier detection; value_counts.

Six empty cells, one duplicated row, three spellings of Hyderabad, an age of 150 and a salary twenty times the next. Clean it and your row count should fall from 12 to 11 and your city count from 5 to 3.

What it was built from

data/course-9-python-da/monthly-sales.csv

72 rows · month · region · revenue

Practise: groupby and agg; pivot_table; melt and stack; merge and join; resample and rolling means; matplotlib, Seaborn and Plotly.

Long format on purpose. pivot_table it into a 24 x 3 grid, plot the three lines, then melt it back -- and check you get the same 72 rows you started with.

What it was built from

Course 10 — Document Oriented Database

data/course-10-mongodb/courses.csv

3 rows · course_id · title · credits · instructor · capacity

Practise: $lookup; normalised against embedded modelling; schema validation rules.

The referenced half of the model. Embed it into each student and then change an instructor's name -- count how many documents you must touch. That count is the argument for referencing.

What it was built from

data/course-10-mongodb/students.csv

6 rows · student_id · name · age · city · enrolled_courses · grades

Practise: insertMany; find with $eq, $gt, $in; $elemMatch on arrays; embedded against referenced models; aggregation $unwind, $group, $lookup; multikey indexes.

Two semicolon-separated columns become ONE array of subdocuments. S104 has no enrolments -- so $unwind will drop that student unless you pass preserveNullAndEmptyArrays.

What it was built from

Course 11 — Business Intelligence Tools

data/course-11-bi/dim-date.csv

4 rows · date_key · date · year · month · quarter

Practise: time intelligence; date hierarchies; grouping by month.

There is no March. Group by month and you get four rows, not five -- which is what breaks a month-on-month growth column.

What it was built from

data/course-11-bi/dim-product.csv

4 rows · product_key · product · category · supplier_key · unit_cost · list_price

Practise: dimensional modelling; star against snowflake; relationships and cardinality; Power Query.

Four products. The supplier column is the one edge that turns the star into a snowflake.

What it was built from

data/course-11-bi/dim-store.csv

3 rows · store_key · store · region · opened

Practise: slicers and cross-filtering; row-level security.

Two southern stores against one northern one -- so a naive average by region is not the same as a total by region.

What it was built from

data/course-11-bi/fact-sales.csv

9 rows · date_key · store_key · product_key · qty

Practise: SUM, COUNT, DISTINCTCOUNT; CALCULATE and filter context; measure against calculated column; fan and chasm traps.

Join it to all three dimensions and you have the flat table a BI tool builds internally. Revenue is qty x list_price: 12,880 in total, 10,360 of it South.

What it was built from

Course 12 A — Machine Learning

data/course-12a-ml/customer-segments.csv

120 rows · annual_spend · visits_per_year · online_ratio · true_segment

Practise: k-Means; the elbow method; silhouette score; feature scaling before distance-based methods; hierarchical clustering; PCA for visualisation.

Cluster it WITHOUT scaling first. annual_spend runs to five figures and online_ratio is under 1, so unscaled k-Means clusters on spend alone. Then scale and watch the answer change.

What it was built from

data/course-12a-ml/house-prices.csv

200 rows · area_sqft · bedrooms · age_years · price_lakh

Practise: simple and multiple linear regression; train/test split; MAE, MSE, RMSE, R-squared; feature scaling; polynomial regression; regularisation.

Fit it and compare your coefficients with the four above. Then scale the features and refit: the coefficients change, the predictions do not, and knowing why is the point.

What it was built from

data/course-12a-ml/loan-approval.csv

300 rows · income · debt · credit_score · approved

Practise: logistic regression; k-NN; decision tree; Naive Bayes; SVM; confusion matrix, precision, recall, F1; ROC and AUC; cross-validation.

Generated from a logistic rule, so there IS an irreducible error rate -- a model reporting 100% has leaked the label. Compare your coefficients with the true ones, and note which feature matters most: credit_score has by far the biggest RAW coefficient, but income has the biggest effect, because a coefficient means nothing until you multiply it by the spread of its variable.

What it was built from

Course 12 B — Big Data Technologies

data/course-12b-bigdata/web-logs.csv

1200 rows · timestamp · ip · path · status · bytes

Practise: MapReduce word count and its shape; the shuffle and sort phase; combiners; Hive GROUP BY; Pig FOREACH GENERATE; Spark RDD reduceByKey and DataFrame agg.

Big enough that counting by hand is out and a map-reduce is in. Count hits per path, bytes per IP and the error rate three ways -- a dict, a GROUP BY and reduceByKey -- and the answers must agree.

What it was built from

data/course-12b-bigdata/wordcount-corpus.csv

5 rows · doc_id · text

Practise: the canonical MapReduce word count; mapper, combiner, reducer; TF-IDF (Course 15 A uses the same file).

Small enough to count by hand, which is the point: work out the answer on paper, then make MapReduce agree with you.

What it was built from

Course 13 A — Artificial Intelligence

data/course-13a-ai/family-relations.csv

9 rows · parent · child

Practise: Prolog facts and rules; unification and backtracking; recursive rules (ancestor); first-order logic; forward and backward chaining.

Load it as parent/2 facts and define sibling, grandparent and a recursive ancestor. Six grandparent pairs -- count them by hand before you run it.

What it was built from

data/course-13a-ai/graph-edges.csv

16 rows · from_city · to_city · cost

Practise: BFS, DFS, uniform-cost search; depth-limited and iterative deepening; greedy best-first and A*; admissible heuristics.

The classic map. BFS finds a three-hop route costing 450; uniform-cost finds a four-hop route costing 418. Fewest steps and cheapest are different questions, and this file proves it.

What it was built from

data/course-13a-ai/map-colouring.csv

9 rows · region · neighbour

Practise: constraint satisfaction; backtracking search; forward checking and arc consistency (AC-3); minimum-remaining-values heuristic.

Tasmania touches nothing, so it takes any colour -- a free variable that MRV should pick last. WA, NT and SA form a triangle, which is why two colours cannot work.

What it was built from

Course 13 B — Cloud Computing

data/course-13b-cloud/iam-policies.csv

7 rows · principal · action · resource · effect

Practise: IAM policy evaluation; explicit deny against implicit deny; least privilege; wildcards in resource ARNs.

Carol has s3:* on the bucket AND an explicit Deny on delete. Explicit Deny wins -- so a wildcard Allow is not the same as unrestricted access, and dave, who appears nowhere, is denied by default.

What it was built from

data/course-13b-cloud/storage-costs.csv

3 rows · tier · temperature · gb_stored · price_per_gb_month · retrieval_per_gb · egress_per_gb

Practise: storage tiers and lifecycle policies; total cost of ownership; egress charges; capex against opex.

Work out the monthly bill, then the bill if you had to read every byte back once. The cheapest tier to STORE is the most expensive to READ, and that reversal is the exam answer.

What it was built from

Course 14 A — Deep Learning

data/course-14a-deeplearning/sensor-failures.csv

400 rows · temperature_c · vibration_mm_s · failed

Practise: binary classification with a neural net; why depth helps; sigmoid output and binary cross-entropy; overfitting, dropout, early stopping.

The boundary is an ellipse, so a linear model cannot do well however long you train it. Five per cent of labels are flipped, so 0.95 is the ceiling -- anything above it is a leak.

What it was built from

data/course-14a-deeplearning/xor.csv

4 rows · x1 · x2 · y

Practise: the perceptron and its limit; activation functions; one hidden layer; backpropagation by hand.

Four rows that ended an AI winter. Train a single-layer perceptron until you are convinced it cannot exceed 3 of 4, then add one hidden layer.

What it was built from

Course 14 B — Time Series

data/course-14b-timeseries/ar2-series.csv

300 rows · t · value

Practise: stationarity; ACF and PACF read together; AR, MA and ARMA; Yule-Walker and MLE estimation; AIC and BIC; the Ljung-Box test; ADF and KPSS.

Built from phi = (0.6, -0.3) after a 200-point burn-in. The PACF cutting off at lag 2 is how you would have identified the order without being told.

What it was built from

data/course-14b-timeseries/macro-indicators.csv

250 rows · t · rates · inflation · unrelated

Practise: VAR models; Granger causality; impulse response; state-space form and the Kalman filter; cointegration.

Causality is planted in ONE direction: rates move inflation, inflation does not move rates. Test both ways -- a Granger test that fires in both directions has found correlation, not cause. The third column is a control that should fire in neither.

What it was built from

data/course-14b-timeseries/seasonal-sales.csv

72 rows · month · sales

Practise: decomposition, additive against multiplicative; STL; seasonal differencing; SARIMA; Holt-Winters; forecast intervals.

A linear trend of +8 a month under a 12-month season. Difference once at lag 12 and the season goes; difference again at lag 1 and the trend goes. Doing it in the wrong order is the classic error.

What it was built from

Course 15 A — Natural Language Processing

data/course-15a-nlp/ner-sentences.csv

5 rows · sentence · expected_entities

Practise: named entity recognition; POS tagging; chunking; evaluating NER against gold labels; precision and recall for extraction.

The gold labels are in the second column, so you can SCORE the tagger rather than eyeball it. Expect the model to get the cities right and to struggle with 'Andhra Pradesh' and 'Krishna'.

What it was built from

data/course-15a-nlp/sentiment-reviews.csv

20 rows · text · label

Practise: tokenization; stopword removal; stemming and lemmatization; bag of words and TF-IDF; Naive Bayes and logistic regression for sentiment; train/test split on text.

Twenty reviews, balanced, labelled by hand so the accuracy you compute means something. One review is deliberately mixed.

What it was built from

Course 15 B — Data Engineering and MLOps

data/course-15b-mlops/loan-current.csv

400 rows · batch · income · debt · credit_score · approved

Practise: population stability index; the Kolmogorov-Smirnov test; data drift against concept drift; retraining triggers and the metric gate; monitoring.

ONE feature moved. Detect which, and resist retraining on reflex: the inputs shifted but the input-to-label relationship did not, so a retrain buys almost nothing. Knowing that is the Unit 5 answer.

What it was built from

data/course-15b-mlops/loan-reference.csv

400 rows · batch · income · debt · credit_score · approved

Practise: training a baseline; MLflow experiment tracking; model registry; DVC data versioning.

Train on this one and register it. It is the reference every later batch is compared against.

What it was built from