Skip to the content

Welcome

This is the complete study package for Data Science using Python (STS-208). Its first stated objective is not a topic but a pipeline — extract raw data from text, CSV, XML, JSON, HTML, SQL and NoSQL sources, clean it, transform it, load it and visualize it — so the practical page is written as one pipeline run end to end, in the order the stages have to happen, rather than as seven unrelated programs.

Two of the seven items are already taught in full, and are linked rather than repeated. NumPy arrays and the whole Pandas data-wrangling list are covered unit by unit, with worked programs, in Python for Data Analysis. The relational design and SQL behind item 4 are in Database Management Systems; the document model behind item 5 is in Document Databases; the language itself is in Python Programming and Data Structures; and NLTK, which the objectives name, is in Natural Language Processing.

What the practical page writes out is what those courses do not contain: HTML and XML parsing, the anomaly pass, binary files, the re module outside Pandas, and driving the two databases from Python — held together by the pipeline.

What ran, and what did not. Every listing on the practical page that could be executed was, and the block beneath it is the output it actually produced. That includes NumPy and Pandas. Three listings are marked NOT EXECUTED on their first line — BeautifulSoup, which is not installed, and PyMySQL and PyMongo, which need servers these pages are not built against. The syllabus names all three, so all three are shown, and each is paired with a listing that did run and gives the same answer.

Course Objectives and Outcomes

  1. Able to apply to the data set an Extract, Transform, Load pipeline which extracts raw data — text files, CSV files, XML files, JSON, HTML files, SQL databases, NoSQL databases — cleans the data, performs transformations on it, loads it and visualizes it.
  2. Able to be familiar and expert in the use of the Python libraries and modules Pandas, NumPy, Beautiful Soup, PyMySQL, PyMongo, NLTK and Matplotlib.

What is in this Course

PRACTICAL

One Pipeline, Eight Stages

The same six records parsed from delimited text, CSV, JSON, XML and HTML into one structure, with the naive split(",") shown shifting every column on a quoted field; an anomaly pass that accepts two of seven records and names the fault in each of the other five; struct, array and pickle, with the alignment that costs two bytes and the integer 1 read back as 16777216; the re module, with a greedy match that returns one wrong answer and a plausible pattern that reads a version number as money; a normalised schema whose constraints reject three illegal inserts and whose transaction is rolled back; the same data as documents, with replace_one discarding three fields and $unwind dropping a document; and NumPy and Pandas, with views against copies, axis=, the ddof the two libraries disagree about, and a join that loses a row.

REFERENCE

Official Syllabus

The prescribed objectives, the datasets note and the seven-item practical list, as printed.

Where Each Prescribed Item Is Covered

#PrescribedAlready taughtWritten out here
1Parse text, CSV, HTML, XML and JSON; check anomalies and missing values Unit 3.1, 3.2 for text, CSV and JSON XML and HTML parsing, and the whole anomaly pass
2Reading and writing binary files file handling in Unit 5 all of it — struct, array, pickle, byte order and alignment
3Searching, splitting and replacing by regular expression Unit 4.2 applies them to Pandas columns the re API itself, and the two mistakes that do not raise
4Design a relational database, populate it, CRUD with SQL Database Management Systems — design, normalisation, SQL driving it from Python, constraints enforced, and the transaction
5A MongoDB client with pymongo: insert, search, remove, update, replace, aggregate, index Document Databases the ETL role, and where the document model differs from SQL — with every call mirrored by runnable Python
6NumPy arrays: shapes, sources, reshape, slice, index, arithmetic, logic, aggregation Unit 1 views against copies, dtype truncation, and axis=
7Pandas: hierarchical indexing, missing data, column arithmetic, merge and aggregate, plotting, file I/O Units 2, 3 and 5 the three places the last stage of a pipeline leaks

The Three Python Courses Together

This course is one end of a range the catalogue runs from both directions.

PaperRuleWhat it builds
STS-105 Statistical Methods using Python no packages at all every distribution, every tail area and every algorithm written from arithmetic — so you know what a \(t\) table is
STS-108 Data Handling using R and STS-207 using SPSS the tool does the statistics choosing the method, checking its assumptions, and reading the output
STS-208the full stack getting the data into a state where any of that is possible — which is where most of the working time actually goes

The order is deliberate: a student who has written the incomplete beta function once, chosen between a sign test and a Wilcoxon once, and cleaned a batch with five kinds of fault in it once has met the three different skills the word “statistics” is used for.

Next course in learning order: Economics Allied courses