Source document. This page reproduces the syllabus this course was written to, as published — its semesters, credits and paper numbers are that document’s, not this site’s. The course itself is studied on its own, in any order.
Course Information
Title
Computational Statistics and R Programming
Theory Credits
3 (3 hrs/week)
Practical Credits
1 (2 hrs/week)
Course Outcomes
Students will be able to import and preprocess statistical datasets in R efficiently.
Students will demonstrate proficiency in basic data-manipulation techniques for statistical preparation.
Students will apply R functions to handle real-world data issues like outliers and inconsistencies.
Students will compute and interpret descriptive statistics using R commands and packages.
Students will analyse data distributions and identify patterns or anomalies statistically.
Students will generate summary reports for datasets to support preliminary statistical insights.
Theory — Five Units
Unit 1: Basics of R for Statistical Data Handling
R environment — installation, command prompt, basic data types (vectors, matrices, data frames). Data import / export — reading from CSV, Excel, databases; handling missing values. Basic operations — sub-setting, merging datasets, applying functions.
Unit 2: Descriptive Statistics & Data Summarization in R
Measures of central tendency (mean, median, mode) and variability (variance, standard deviation, range). Skewness, kurtosis and quantiles; summary functions summary() and describe(). Handling categorical data — frequency tables, cross-tabulations and contingency tables.
Data import and preprocessing — import a dataset (e.g., iris.csv), handle missing values, perform basic sub-setting.
Descriptive statistics computation — calculate mean, variance and skewness for a sample dataset and interpret results.
Frequency analysis — create contingency tables and chi-square tests on categorical data.
Basic visualisations — generate histograms and boxplots to assess data normality.
Hypothesis testing — conduct t-tests and ANOVA on experimental data, reporting p-values.
Non-parametric tests — apply the Wilcoxon tests and compare with parametric alternatives.
Linear regression modelling — fit a simple linear model, plot residuals and predict new values.
Note: Practical problems must be done in the Computer Lab at least 4 hours per month using R. Use real datasets (e.g., UCI ML Repository) or built-in R datasets such as mtcars and airquality.