This is the complete study package for Data Science using Python (STS-208). Its first stated objective is not a topic but a pipeline — extract raw data from text, CSV, XML, JSON, HTML, SQL and NoSQL sources, clean it, transform it, load it and visualize it — so the practical page is written as one pipeline run end to end, in the order the stages have to happen, rather than as seven unrelated programs.
What the practical page writes out is what those courses do
not contain: HTML and XML parsing, the anomaly pass, binary files, the
re module outside Pandas, and driving the two databases from Python — held
together by the pipeline.
NOT EXECUTED on their
first line — BeautifulSoup, which is not installed, and PyMySQL and PyMongo, which need
servers these pages are not built against. The syllabus names all three, so all three are
shown, and each is paired with a listing that did run and gives the same answer.
The same six records parsed from delimited text, CSV, JSON, XML and HTML into one
structure, with the naive split(",") shown shifting every column on a quoted
field; an anomaly pass that accepts two of seven records and names the fault in each of the
other five; struct, array and pickle, with the
alignment that costs two bytes and the integer 1 read back as 16777216; the re
module, with a greedy match that returns one wrong answer and a plausible pattern that reads
a version number as money; a normalised schema whose constraints reject three illegal inserts
and whose transaction is rolled back; the same data as documents, with
replace_one discarding three fields and $unwind dropping a
document; and NumPy and Pandas, with views against copies, axis=, the
ddof the two libraries disagree about, and a join that loses a row.
The prescribed objectives, the datasets note and the seven-item practical list, as printed.
| # | Prescribed | Already taught | Written out here |
|---|---|---|---|
| 1 | Parse text, CSV, HTML, XML and JSON; check anomalies and missing values | Unit 3.1, 3.2 for text, CSV and JSON | XML and HTML parsing, and the whole anomaly pass |
| 2 | Reading and writing binary files | file handling in Unit 5 | all of it — struct, array,
pickle, byte order and alignment |
| 3 | Searching, splitting and replacing by regular expression | Unit 4.2 applies them to Pandas columns | the re API itself, and the two mistakes that do not raise |
| 4 | Design a relational database, populate it, CRUD with SQL | Database Management Systems — design, normalisation, SQL | driving it from Python, constraints enforced, and the transaction |
| 5 | A MongoDB client with pymongo: insert, search, remove, update, replace, aggregate, index | Document Databases | the ETL role, and where the document model differs from SQL — with every call mirrored by runnable Python |
| 6 | NumPy arrays: shapes, sources, reshape, slice, index, arithmetic, logic, aggregation | Unit 1 | views against copies, dtype truncation, and axis= |
| 7 | Pandas: hierarchical indexing, missing data, column arithmetic, merge and aggregate, plotting, file I/O | Units 2, 3 and 5 | the three places the last stage of a pipeline leaks |
This course is one end of a range the catalogue runs from both directions.
| Paper | Rule | What it builds |
|---|---|---|
| STS-105 Statistical Methods using Python | no packages at all | every distribution, every tail area and every algorithm written from arithmetic — so you know what a \(t\) table is |
| STS-108 Data Handling using R and STS-207 using SPSS | the tool does the statistics | choosing the method, checking its assumptions, and reading the output |
| STS-208 | the full stack | getting the data into a state where any of that is possible — which is where most of the working time actually goes |
The order is deliberate: a student who has written the incomplete beta function once, chosen between a sign test and a Wilcoxon once, and cleaned a batch with five kinds of fault in it once has met the three different skills the word “statistics” is used for.