Skip to the content

Source document. This page reproduces the syllabus this course was written to, as published — its semesters, credits and paper numbers are that document’s, not this site’s. The course itself is studied on its own, in any order.

Course Information

TitleComputational Statistics and R Programming
Theory Credits3 (3 hrs/week)
Practical Credits1 (2 hrs/week)

Course Outcomes

  1. Students will be able to import and preprocess statistical datasets in R efficiently.
  2. Students will demonstrate proficiency in basic data-manipulation techniques for statistical preparation.
  3. Students will apply R functions to handle real-world data issues like outliers and inconsistencies.
  4. Students will compute and interpret descriptive statistics using R commands and packages.
  5. Students will analyse data distributions and identify patterns or anomalies statistically.
  6. Students will generate summary reports for datasets to support preliminary statistical insights.

Theory — Five Units

Unit 1: Basics of R for Statistical Data Handling

R environment — installation, command prompt, basic data types (vectors, matrices, data frames). Data import / export — reading from CSV, Excel, databases; handling missing values. Basic operations — sub-setting, merging datasets, applying functions.

Open Unit 1 →

Unit 2: Descriptive Statistics & Data Summarization in R

Measures of central tendency (mean, median, mode) and variability (variance, standard deviation, range). Skewness, kurtosis and quantiles; summary functions summary() and describe(). Handling categorical data — frequency tables, cross-tabulations and contingency tables.

Open Unit 2 →

Unit 3: Data Visualization for Statistical Insights in R

Base R graphics — histograms, boxplots, scatter plots, bar charts, residual plots for assumption checking in statistics.

Open Unit 3 →

Unit 4: Inferential Statistics & Hypothesis Testing in R

Probability distributions in R — Normal, Binomial, Poisson; random sampling (rnorm(), dbinom()). Hypothesis testing — t-tests, chi-square tests, ANOVA; p-values and confidence intervals. Non-parametric tests — Wilcoxon.

Open Unit 4 →

Unit 5: Regression Modeling in R

Karl Pearson correlation coefficient, Spearman rank correlation coefficient and simple linear regression (lm()).

Open Unit 5 →

Practical — List of Experiments (7)

  1. Data import and preprocessing — import a dataset (e.g., iris.csv), handle missing values, perform basic sub-setting.
  2. Descriptive statistics computation — calculate mean, variance and skewness for a sample dataset and interpret results.
  3. Frequency analysis — create contingency tables and chi-square tests on categorical data.
  4. Basic visualisations — generate histograms and boxplots to assess data normality.
  5. Hypothesis testing — conduct t-tests and ANOVA on experimental data, reporting p-values.
  6. Non-parametric tests — apply the Wilcoxon tests and compare with parametric alternatives.
  7. Linear regression modelling — fit a simple linear model, plot residuals and predict new values.

Note: Practical problems must be done in the Computer Lab at least 4 hours per month using R. Use real datasets (e.g., UCI ML Repository) or built-in R datasets such as mtcars and airquality.

Open practical course material →

Text Books / References

  1. J. Chambers (2008) — Software for Data Analysis: Programming with R, Springer.
  2. M. J. Crawley (2017) — The R Book, John Wiley & Sons.
  3. N. Matloff (2011) — The Art of R Programming, No Starch Press.
  4. Mark Gardener (2012) — Beginning R — The Statistical Programming Language, John Wiley & Sons.
  5. Purohit, Gore & Deshmukh (2008) — Statistics Using R, Narosa Publishing House.
  6. W. N. Venables & D. M. Smith — An Introduction to R (R Core Team, 2013, online).
  7. Nathan Yau (2011) — Visualize This, Wiley.
  8. Zumel & Mount (2014) — Practical Data Science with R, Manning Publications.

Suggested Co-Curricular Activities

  1. Training of students by related industrial experts.
  2. Assignments including technical assignments, if any.
  3. Seminars, group discussions, quiz, debates etc. on related topics.
  4. Preparation of audio and videos on tools of diagrammatic and graphical representations.
  5. Collection of material / figures / photos of related topics.
  6. Invited lectures and presentations of stalwarts on those topics.
  7. Visits / field trips of firms, research organizations etc.
UnitTopicApprox. Weightage
1R Basics & Data Handling20 %
2Descriptive Statistics in R20 %
3Data Visualization15 %
4Inferential Statistics & Tests25 %
5Regression Modeling20 %

Quick Reference — Core R Functions

TaskR Function
Read CSV / Excelread.csv() / readxl::read_excel()
Handle missingna.omit() / na.rm = TRUE
Summarysummary() / psych::describe()
Central tendencymean / median
Dispersionvar / sd / IQR
Skewness / Kurtosismoments::skewness / kurtosis
Frequency tabletable() / prop.table()
Histogram / Boxplothist() / boxplot()
Scatterplot(x, y)
Distributionsd/p/q/r + norm/binom/pois/t/chisq/f
Random samplesample() / rnorm()
t-testt.test()
Chi-squarechisq.test()
ANOVAaov(y ~ x)
Wilcoxonwilcox.test()
Correlationcor() / cor.test()
Linear regressionlm(y ~ x)
Predictionpredict(fit, new, interval = "prediction")
Residual diagnosticsplot(fit)