Skip to the content

EXECUTED WITH R 4.3.3

This script is run with R 4.3.3, by tools/data-science/run_r_equivalents.py, and the lab page shows what it printed and the plots it drew. Every number in its comments was checked against that output.

The code

Straight from labs/course-6-r/10_missing_outliers.R, unchanged.

# =====================================================================
# Run with R 4.3.3 (Rscript --vanilla). What it prints, and the plots it
# draws, are on the lab page, and tools/data-science/run_r_equivalents.py
# runs it again. (Until October 2026 R could not be installed where these
# labs are checked, so this file was desk-checked only; every number in its
# comments has since been checked against R's own output.)
# =====================================================================
# Experiment 10: Handle missing data and detect outliers
# Python equivalent: python/10_missing_outliers.py

# Step 1: Enter the data, with missing values and an outlier
x <- c(45, 67, NA, 52, 89, 91, NA, 64, 58, 82,
       76, 69, 71, 250, 60, 55, 93, 48, 79, NA)   # 250 is a planted outlier

# --- DETECTING MISSING VALUES ---
# Step 2: Find the missing values
is.na(x)              # logical vector
sum(is.na(x))         # 3
which(is.na(x))       # 3 7 20  -- positions, 1-based
mean(is.na(x)) * 100  # 15% missing

# NEVER write x == NA. Comparing with an unknown value yields NA, never TRUE.
# This is the same trap as SQL's "= NULL" from Course 5.

# --- HANDLING ---
# Step 3: Impute them by the mean or the median
mean(x, na.rm = TRUE)          # 79.35 -- drag upward from the 250
median(x, na.rm = TRUE)        # 69.00 -- resistant
# [Corrected: these said 79.06 and 68.00; R and the Python version both give
# 79.35 and 69.00.]
clean <- na.omit(x)            # drop the NAs

x_mean_imputed   <- ifelse(is.na(x), mean(x, na.rm = TRUE), x)
x_median_imputed <- ifelse(is.na(x), median(x, na.rm = TRUE), x)
# With an outlier present, MEDIAN imputation is the safer choice -- the mean
# has already been distorted by the very value you are trying to work around.

# --- OUTLIERS: the IQR rule ---
# Step 4: Find outliers by the IQR rule
q <- quantile(clean, c(0.25, 0.75))
iqr <- IQR(clean)
lower <- q[1] - 1.5 * iqr        # 22.00
upper <- q[2] + 1.5 * iqr        # 118.00
clean[clean < lower | clean > upper]     # 250
boxplot(clean)$out                       # same answer, drawn

# --- OUTLIERS: the z-score rule ---
# Step 5: Find outliers by the z-score rule
z <- (clean - mean(clean)) / sd(clean)
clean[abs(z) > 3]                # 250 here, z = 3.677

# MASKING -- why the IQR rule is preferred when outliers may cluster:
# add a second extreme value and the sd inflates enough that NEITHER is
# flagged by the z-score rule, while the IQR rule still catches both.
# Step 6: Add a second outlier, and see masking
masked <- c(clean, 260)
zm <- (masked - mean(masked)) / sd(masked)
masked[abs(zm) > 3]              # returns NOTHING -- both outliers masked
qm <- quantile(masked, c(0.25, 0.75)); im <- IQR(masked)
masked[masked < qm[1] - 1.5*im | masked > qm[2] + 1.5*im]   # 250 260 -- caught

Where this sits

The same thing in Python: Missing Data and Outlier Detection in Python.

The theory behind it is in this course’s units; the whole lab, with all 18 experiments, is on the lab page.