Useful for ISS
This is the complete study package for Data Handling using R (STS-108). Its stated outcome is one sentence — able to carry out the statistical analysis and write a statistical report using R for any data set — and the syllabus is eight headings rather than a list of programs. So these pages follow the workflow end to end on a single data set of twenty records, and every stage is checked against the last.
Five of the eight prescribed topics are not in it at all, and those are what is written out here: measurement scales and what each permits; data transformations; the seven diagrams beyond the standard four; model building with cross-validation; and the evaluation of model performance. Three of the five named non-parametric tests are new as well.
Able to carry out the statistical analysis and write a statistical report using R for any data set.
Scales of measurement and the summaries each permits; the seven pre-processing steps in the order they must be taken; standardization, normalization and the log, root, square and circular transformations, each judged by what it does to the skewness; the full prescribed list of diagrams, including the five base R does not make obvious; polynomial fits scored by 5-fold and leave-one-out cross-validation, where the in-sample error falls all the way to degree 8 and the cross-validated error reaches \(5351\); the confusion matrix with accuracy, precision, recall, specificity, \(F_1\) and AUC computed from four counts; and the parametric and non-parametric tests run side by side on the same data, where the sign test gives \(p = 0.146\) and Wilcoxon \(p = 0.0096\).
REFERENCEThe prescribed objective and the eight headings of the data-handling list, as printed, with the two notes on data sources and the practical record.
| Prescribed topic | Already taught | What this course adds |
|---|---|---|
| 1. Understanding the data set | Unit 1 — import, missing values, subsetting, merging | measurement scales and what each permits; dependence and independence among variables; the pre-processing order; the shape of the report |
| 2. Data transformations | — | all of it: standardization, normalization, log, sine, cosine, square, square root and exponential, with the skewness before and after |
| 3. Descriptive statistics | Unit 2 — the arithmetic | choosing the summary by the variable's scale, and naming R's conventions |
| 4. Data visualization | Unit 3 — histogram, boxplot, scatter, bar | pie, line, frequency polygon, ogive, area, Gantt, heat map and the correlation matrix |
| 5. Model building | Unit 5 — fitting and diagnostics | over-fitting and under-fitting measured; train/test, \(k\)-fold and leave-one-out; cross-tabs |
| 6. Evaluation of model performance | — | all of it: RMSE, MAE, \(R^{2}\) and adjusted \(R^{2}\) for quantitative outcomes; the confusion matrix, its five ratios and AUC for qualitative ones |
| 7. Parametric tests | Unit 4 — \(z\), \(\chi^{2}\), \(t\), \(F\), one-way ANOVA | the order they must be run in, and one worked example of a test that cannot fail |
| 8. Non-parametric tests | Unit 4 — Wilcoxon | the sign test, Mood's median test, Mann–Whitney \(U\) and the runs test, and what each one throws away |