Useful for UGC NET · CSIR NET · ASRB NET
Part of the machine-learning path: Machine Learning, Artificial Intelligence, Neural Networks and Deep Learning and Natural Language Processing.
Machine learning is not "the model". It is the pipeline, and the model is the easy part.
A student's first instinct is to reach for RandomForestClassifier. The
accuracy comes out at 0.94, and the exercise feels finished. It is not, because
the questions that decide whether that number means anything have not been
asked: what is the base rate? was the test set really held out? is 94% better
than always predicting the majority class?
| The part that feels like the work | The part that is the work |
|---|---|
| Choosing an algorithm | Framing the problem — what is being predicted, and for whom |
| Fitting the model | Preparing the data — Unit 2, and 70% of the effort |
| Getting a high accuracy | Knowing whether that accuracy is good — the base rate |
| Adding features | Not leaking the target into them |
| The training score | The test score, on data the model has never seen |
Unit 2 is the most important unit in this course and the one students skip. Units 3, 4 and 5 are a catalogue of algorithms, each three lines of scikit-learn. Unit 2 is what makes any of them mean something.
This course is a convergence point — more of the catalogue meets here than anywhere else.
| From | You have | Used here |
|---|---|---|
| Statistical Foundations for Data Science | Regression, correlation, hypothesis testing, distributions | Unit 3 is Statistical Foundations for Data Science's regression, refitted as prediction rather than explanation. §3.1 says exactly what changed |
| Data Mining | Decision trees (ID3, C4.5, CART), Naive Bayes, k-NN, K-Means, DBSCAN | Units 4 and 5 repeat these. §4.1 and §5.1 say which parts are revision, so you do not study them twice |
| Python for Data Analysis and Visualization | NumPy, pandas, cleaning, feature engineering, matplotlib | Every lab. Unit 2's preprocessing is Python for Data Analysis and Visualization Unit 3 with a purpose |
| Python Programming and Data Structures | Python, functions, classes | scikit-learn's fit/predict API |
| Data Science with R | The data science lifecycle, model evaluation, ROC | Unit 2's evaluation section, in Python instead of R |
Data Mining taught decision trees, Naive Bayes, k-NN, K-Means, hierarchical clustering and DBSCAN as hand-traced algorithms. This course teaches the same six as tools you fit and evaluate.
That difference is the point. Data Mining asked how does ID3 choose a split? — arithmetic on paper. This course asks is this tree overfitting, and how would you know? Same algorithm, a different question, and both are examined.
If you took Data Mining, budget your time on Units 2 and 3, which are new.
Understand fundamental concepts, types, and applications of machine learning.
Develop, evaluate, and optimize machine learning models through preprocessing, training, and feature engineering techniques.
Apply supervised and unsupervised learning algorithms to real-world problems using appropriate tools and methods.
NOTE
There are only three objectives, and four outcomes. Every other course's syllabus has five of each. Nothing appears to be missing — the three objectives do cover the five units between them — but if an examiner asks for "the fourth course objective", the document does not have one. Recorded in SYLLABUS-REVIEW.md.
Types of human learning and their machine analogues; Mitchell's definition; machine learning against traditional programming, and when not to use it; supervised, unsupervised, semi-supervised and reinforcement learning; the ML pipeline and where the effort goes; types of data and how each should be encoded; the feature matrix and the curse of dimensionality.
UNIT 2Preprocessing in order, and why splitting comes first; missing values, outliers, encoding and scaling; the three-way split and cross-validation; bias and variance; interpretability; why accuracy lies, the confusion matrix, precision, recall, F1 and AUC; performance enhancement and class imbalance; feature engineering, target leakage, subset selection and PCA.
UNIT 3What changes when explanation becomes prediction; simple linear regression and the LINE assumptions; multiple regression, multicollinearity and adjusted R²; polynomial regression and the conditioning trap; logistic regression, the sigmoid and the odds ratio; maximum likelihood estimation; Ridge, Lasso and elastic net.
UNIT 4The classification pipeline; binary, multi-class and multi-label; Naïve Bayes, the naive assumption and Laplace smoothing; k-Nearest Neighbour, distance metrics and choosing k; decision trees and pruning; support vector machines, the margin, support vectors and the kernel trick; random forest, its two sources of randomness and out-of-bag error.
UNIT 5Unsupervised against supervised, and why evaluation is the hard part; clustering types; K-Means, WCSS and the elbow; k-Medoids and robustness; hierarchical clustering and linkage; DBSCAN, core, border and noise points; internal against external validation metrics; case studies in image and speech recognition, spam filtering and fraud detection.
PRACTICEExam-style questions with fully worked solutions.
LABEvery prescribed lab experiment, with code and expected output.
ALSOA second, deeper treatment of the same subject, written separately. It is organised not by the five syllabus units but by the kind of supervision signal an algorithm learns from — supervised, unsupervised, semi-supervised, reinforcement. Each of its 23 algorithms gets its mathematics, its assumptions and failure modes, three worked examples from finance, agriculture and medicine, and runnable Python and R. Use the five units above for the syllabus; use this when you want to understand an algorithm properly.