Part of the machine-learning path: Machine Learning, Artificial Intelligence, Neural Networks and Deep Learning and Natural Language Processing.
Every other course in this catalogue works on data that was already numeric, or was numeric once you encoded it. Language is not, and every technique here is an answer to the question "how do I turn this text into numbers without throwing away what it means?"
The course is essentially four answers to that question, in increasing order of power and cost:
| Answer | Represents | Cannot represent |
|---|---|---|
| Bag of words | which words appear | word order, or that two words mean the same |
| N-grams | short local order | anything beyond n |
| Word embeddings (Word2Vec, GloVe) | that "fee" and "cost" are related | that "bank" has two senses |
| Contextual embeddings (BERT) | a different vector per occurrence | — but needs enormous pre-training |
THE BIG IDEA
Ambiguity is not an edge case in language; it is the normal condition. "I saw the man with the telescope" has two grammatical parses and nothing in the sentence chooses between them. "The bank was closed" has 18 WordNet senses for one word.
Every NLP technique is, at bottom, a way of choosing among readings — and the ones that work choose using context, which is why the field ended up at contextual embeddings.
Experiment 9 compares four documents where the answer is known in advance. Two of them are paraphrases; a third is about the monsoon. Bag-of-words scores the paraphrase at 0.0925 and the unrelated document at 0.3077 — it ranks the wrong one higher, because documents 0 and 3 share "the", "and" and "are" and nothing else.
A representation that ranks by function words is not measuring meaning at all, and that single table is the argument for everything in Units 3–5.
Eleven of the fourteen experiments run against real corpora and real models.
| Tool | What it is here |
|---|---|
| NLTK 3.10 | Brown (1.16M words), Reuters (10,788 docs), Penn Treebank (3,914 parsed sentences), movie_reviews (2,000 labelled), Gutenberg, WordNet |
spaCy 3.8 with en_core_web_sm |
real NER, real dependency parsing, real tokenization |
| scikit-learn 1.9 | the classification pipeline, TF-IDF, cross-validation |
| PyTorch | a character LSTM and a transformer encoder, both trained here |
All three fail for the same reason: huggingface.co is refused by this
environment's egress policy with a 403 at the gateway.
| # | Experiment | File | Its runnable half |
|---|---|---|---|
| 12 | BERT masked-word prediction | 12_bert_mlm.md |
a bidirectional transformer encoder trained here on Brown with a masked-LM objective |
| 13 | Abstractive summarization | 13_summarization.md |
extractive TextRank on a real Reuters article, scored against lead-3 |
| 14 | FAQ chatbot on embeddings | 14_faq_chatbot.md |
the same retriever on TF-IDF, scored against hand-labelled answers |
The runnable halves are not filler. In each case the architecture is identical and only the embedding function differs — which is precisely what makes the comparison instructive. Training the small version yourself is how you find out what pre-training on billions of words actually buys.
tools/data-science/run_nlp_labs.py asserts all three *** NOT EXECUTED *** markers are
still present.
IN DEPTH
Where an experiment could have printed "it worked", this course scores it against hand-labelled truth instead:
| Experiment | What is graded |
|---|---|
| 2 — regex | precision, recall and F1 against a labelled contact list containing deliberate near-misses |
| 8 — NER | spaCy's output against 13 hand-assigned entity labels |
| 9 — similarity | four documents whose correct ranking is known |
| 14 — retrieval | six queries with known answers, none copying an FAQ question |
Three of those scores came out worse than expected, and the notes explain them rather than hide them. That is the point of grading.
Introduce the foundations of Natural Language Processing and its applications in real-world tasks.
Familiarize students with text preprocessing, linguistic analysis, and parsing techniques.
Equip learners with methods for information extraction, word representations, and sentiment classification.
Explore deep learning techniques for NLP, including RNNs, LSTMs, GRUs, and Transformers.
Provide hands-on experience with modern NLP tools (NLTK, spaCy, Hugging Face) for implementing applications such as chatbots, summarization, and document classification.
| Unit | Topic | Notes | Hardest part |
|---|---|---|---|
| 1 | NLP fundamentals, ambiguity, regex | unit-1.md | the three kinds of ambiguity, told apart |
| 2 | Preprocessing, morphology, grammar, parsing | unit-2.md | CYK, and why top-down parsers loop |
| 3 | NER, embeddings, classification | unit-3.md | what Word2Vec learns and how |
| 4 | Deep learning for NLP | unit-4.md | RNN against CNN against transformer, honestly compared |
| 5 | Transformers and modern NLP | unit-5.md | BERT vs GPT, and summarization's failure modes |
Plus lab.md — all fourteen experiments with their measured output — and practice.md — exam questions with worked solutions.
labs/course-15a-nlp/ — the code, and the runner that asserts every figure
these notes quote
data/course-15a-nlp/ — practice datasets, CSV: ner-sentences.csv, sentiment-reviews.csv.
Every one was generated from a known truth, so you can score your answer
rather than just produce one; data/README.md lists what each was built
from, data/PRACTICE-QUESTIONS.md sets questions on each with a computed
answer key, and tools/data-science/check_datasets.py proves every one of those
answers against the file.
| Course | What it gives you here |
|---|---|
| Python Programming and Data Structures (Python) | string handling and regular expressions |
| Python for Data Analysis and Visualization (Python for Data Analysis) | the pipeline and the train/test discipline |
| Machine Learning (Machine Learning) | the classifier, the baselines, cross-validation |
| Neural Networks and Deep Learning (Deep Learning) | the RNN, LSTM and attention material is shared — study Unit 4 of both together |
| Data Mining (Data Mining) | TF-IDF and similarity measures |
NOTE
They cover the same architectures from two directions. Neural Networks and Deep Learning builds attention from scratch and measures it; this course applies it to language. Doing them in the same week roughly halves the work.
References: Siddiqui & Tiwary, Natural Language Processing and Information Retrieval · Kulkarni & Shivananda, Natural Language Processing Recipes, Apress, 2019.
WATCH OUT
The list reads "Reference Book: 1. … 2. 2. Natural Language Processing Recipes …" — a duplicated numeral, not a missing entry. There are two reference books, not three. See review finding D27.
Install NLTK and spaCy on your own machine, and download the models.
python -m spacy download en_core_web_sm is the step everyone forgets, and
nothing in Units 2–3 runs without it.
Tokenize a real paragraph in both, side by side. NLTK and spaCy disagree on contractions, hyphens and punctuation, and seeing where is worth more than the definition of a token.
Learn the difference between stemming and lemmatization by example.
studies → studi against study answers the exam question in one line.
Do the parsing by hand. CYK on a five-word sentence is tedious exactly once, and then the algorithm is obvious. This is the unit students skip and then lose ten marks on.
Score your own output. Every result in these notes is checked against hand-labelled truth rather than eyeballed — that is how it emerged that bag-of-words ranked an unrelated document above a paraphrase, and that spaCy mislabels two Indian state names while getting both cities right.
Do not wait for a GPU. Everything up to Unit 4 runs on a laptop CPU, and the transformer material can be understood from the mechanism before it is ever run at scale.
Unit 1's section on ambiguity, and then run
03_ambiguity_tokenize.py.
Six sentences, each ambiguous in a documented way, with WordNet sense counts and actual parse trees for the structural cases. Everything else in the course is a technique for resolving what those six sentences demonstrate, and the techniques make far more sense once you have seen the problem clearly.
What NLP is and the three properties that make language hard; the levels of analysis from phonology to discourse; applications, and why spam detection is adversarial; lexical, structural and contextual ambiguity told apart, with WordNet sense counts and actual parse trees; garden-path sentences; NLTK against spaCy and when to use each; regular expressions scored against hand-labelled truth, greedy against lazy quantifiers, and where a regex stops being the right tool.
UNIT 2Morphology, lexicon, orthographic rules; inflectional against derivational morphology; finite state transducers and why bidirectionality matters; sentence and word tokenization, and why two trained tokenisers disagree; stopword removal measured on the Brown corpus, and the sentences it destroys; stemming against lemmatization on 'ran', 'better' and 'university'/'universal'; context-free grammars and the Chomsky hierarchy; top-down, bottom-up and chart parsing; why left recursion kills a top-down parser; CYK and its O(n³); semantic analysis and meaning representation.
UNIT 3Named entity recognition, the BIO scheme, and spaCy scored against thirteen hand-assigned labels — including the two Indian state names it gets wrong and the gazetteer that fixes them; bag of words, TF-IDF and what IDF actually is; n-grams and the word order unigrams cannot see; the measurement where counts rank an unrelated document above a paraphrase; the distributional hypothesis, CBOW against skip-gram, negative sampling, GloVe; the classification pipeline and why the vectoriser is fitted on train only; comparing a model gap against the cross-validation spread; the ethics of preprocessing.
UNIT 4Why sequences break feedforward networks; RNN against CNN for text; the vanishing gradient as arithmetic, stated precisely enough to say that 3e-151 is not zero; LSTM gates and the additive cell path; GRU; the LSTM-RNN gap measured on both a constructed dataset and real IMDb, and why it narrows; perplexity and what value means no better than guessing; temperature; the O(T²) cost and why parallelism won; BERT against GPT; what pre-training buys, measured by training the small version; subword tokenization and the Hugging Face ecosystem.
UNIT 5Self-attention worked on checkable numbers; why the scores are divided by √d_k, measured against softmax saturation; multi-head attention, the encoder block, residuals and layer norm; positional encoding; encoder-decoder and what cross-attention does; BERT's pretraining, the full masking recipe including the 10% random and 10% unchanged, and why NSP was dropped; fine-tuning; GPT and hallucination; extractive, abstractive and hybrid summarization, the lead-3 baseline, and why regulated domains stay extractive; document classification, retrieval against generative chatbots, and the threshold every retrieval bot needs.
PRACTICEExam-style questions with fully worked solutions.
LABEvery prescribed lab experiment, with code and expected output.