Skip to the content

Useful for ISS

Welcome

This is the complete study package for Data Handling using R (STS-108). Its stated outcome is one sentence — able to carry out the statistical analysis and write a statistical report using R for any data set — and the syllabus is eight headings rather than a list of programs. So these pages follow the workflow end to end on a single data set of twenty records, and every stage is checked against the last.

R is already on this site, and this course does not teach it again. Computational Statistics and R Programming covers the environment and the language, import and export, missing values, subsetting and merging, the summary measures, the four standard diagrams, the distributions, the \(t\), \(\chi^{2}\), \(F\) and Wilcoxon tests, correlation and regression, and residual diagnostics. All five of its units are linked from the practical page at the point where they are needed.

Five of the eight prescribed topics are not in it at all, and those are what is written out here: measurement scales and what each permits; data transformations; the seven diagrams beyond the standard four; model building with cross-validation; and the evaluation of model performance. Three of the five named non-parametric tests are new as well.

Course Objective and Outcome

Able to carry out the statistical analysis and write a statistical report using R for any data set.

What is in this Course

PRACTICAL

The Whole Workflow, on One Data Set

Scales of measurement and the summaries each permits; the seven pre-processing steps in the order they must be taken; standardization, normalization and the log, root, square and circular transformations, each judged by what it does to the skewness; the full prescribed list of diagrams, including the five base R does not make obvious; polynomial fits scored by 5-fold and leave-one-out cross-validation, where the in-sample error falls all the way to degree 8 and the cross-validated error reaches \(5351\); the confusion matrix with accuracy, precision, recall, specificity, \(F_1\) and AUC computed from four counts; and the parametric and non-parametric tests run side by side on the same data, where the sign test gives \(p = 0.146\) and Wilcoxon \(p = 0.0096\).

REFERENCE

Official Syllabus

The prescribed objective and the eight headings of the data-handling list, as printed, with the two notes on data sources and the practical record.

Where Each Topic Comes From

Prescribed topicAlready taughtWhat this course adds
1. Understanding the data set Unit 1 — import, missing values, subsetting, merging measurement scales and what each permits; dependence and independence among variables; the pre-processing order; the shape of the report
2. Data transformations— all of it: standardization, normalization, log, sine, cosine, square, square root and exponential, with the skewness before and after
3. Descriptive statistics Unit 2 — the arithmetic choosing the summary by the variable's scale, and naming R's conventions
4. Data visualization Unit 3 — histogram, boxplot, scatter, bar pie, line, frequency polygon, ogive, area, Gantt, heat map and the correlation matrix
5. Model building Unit 5 — fitting and diagnostics over-fitting and under-fitting measured; train/test, \(k\)-fold and leave-one-out; cross-tabs
6. Evaluation of model performance— all of it: RMSE, MAE, \(R^{2}\) and adjusted \(R^{2}\) for quantitative outcomes; the confusion matrix, its five ratios and AUC for qualitative ones
7. Parametric tests Unit 4 — \(z\), \(\chi^{2}\), \(t\), \(F\), one-way ANOVA the order they must be run in, and one worked example of a test that cannot fail
8. Non-parametric tests Unit 4 — Wilcoxon the sign test, Mood's median test, Mann–Whitney \(U\) and the runs test, and what each one throws away
On the numbers. R is not run on this site. Every figure on the practical page was computed independently, in exact or double-precision arithmetic, and is printed beside the R line that produces it so that a run can be checked against it.

Next course in learning order: Statistical Methods using Python Statistical computing