Skip to the content
On this page
  1. Section A — Predict the output
  2. Section B — Find and fix
  3. Section C — Write the code
  4. Section D — Long answers
  5. Quick self-test

Every numeric answer here was produced by the executed Python equivalents in labs/course-6-r/python/.


Section A — Predict the output

Q1

v <- c(10, 20, 30, 40, 50)
print(v[-2])
print(v[c(TRUE, FALSE)])
print(length(v[-1]))
Show answer
  • v[-2] excludes the 2nd element. R has no negative-from-the-end indexing; tail(v, 1) gives the last element.

  • c(TRUE, FALSE) is recycled to length 5 as T F T F T, keeping positions 1, 3 and 5.

  • One element removed from five leaves four.

Q2

x <- c(1, 2, 3, NA, 5)
print(mean(x))
print(mean(x, na.rm = TRUE))
print(x == NA)
Show answer

Any arithmetic involving NA yields NA — hence the first. na.rm = TRUE drops it, leaving mean(1,2,3,5) = 11/4 = 2.75. And == NA compares with an unknown value, so every result is unknown. Use is.na(x).

Q3

print(class(10))
print(class(10L))
print(1:3 + c(10, 20))
Show answer

10 is a double; 10L forces an integer. The third recycles c(10,20) to 10 20 10, giving 11, 22, 13 — and warns because 3 is not a multiple of 2. That warning is a bug signal; check your lengths when you see it.


Section B — Find and fix

Q4

Show answer
df$grade <- if (df$marks >= 40) "Pass" else "Fail"

Error: if takes a single condition; df$marks is a whole column. Modern R errors; older R silently used only the first row. Fix: df$grade <- ifelse(df$marks >= 40, "Pass", "Fail")

Q5

Show answer
library(dplyr)
result <- df %>% filter(marks > 50) + select(name, marks)

Error: + between dplyr verbs. + combines ggplot2 layers; %>% chains data operations. Fix: df %>% filter(marks > 50) %>% select(name, marks)

Q6

Show answer
km <- kmeans(customers, centers = 3)

Errors: no scaling and no nstart. Fix: km <- kmeans(scale(customers), centers = 3, nstart = 25)

Without scale() a variable in rupees swamps one in years — the variance ratio in the lab data is about 3.5 billion to 1 (corrected from 4 billion: R gives 3.48 billion). Without nstart you get whatever single random start produced.

Q7

Show answer
f <- factor(c("10", "20", "30"))
values <- as.numeric(f)

Error: returns the level codes 1 2 3, not 10 20 30. Fix: as.numeric(as.character(f))


Section C — Write the code

Q8

Top 3 sections by average marks, in dplyr

Show answer
students %>%
  filter(!is.na(marks)) %>%
  group_by(section) %>%
  summarise(n = n(), avg = mean(marks), .groups = "drop") %>%
  arrange(desc(avg)) %>%
  slice_head(n = 3)

The same six operations as the SQL you know from Database Management Systems: SELECT section, AVG(marks) … GROUP BY section ORDER BY … LIMIT 3.

Q9

Impute missing marks with the median, but only where sensible

Show answer
impute_median <- function(x, max_missing = 0.3) {
  pct <- mean(is.na(x))
  if (pct > max_missing) {
    warning(sprintf("%.0f%% missing -- too much to impute", pct * 100))
    return(x)
  }
  ifelse(is.na(x), median(x, na.rm = TRUE), x)
}

Median, not mean, because an outlier has already distorted the mean — the lab data shows mean 79.35 against median 69.00 with one 250 present (corrected from 79.06 and 68.00, which neither R nor the Python version gives). And the guard matters: imputing a column that is 60% missing invents most of it.

Q10

ggplot2: boxplot of marks by section, faceted by gender

Show answer
ggplot(students, aes(x = section, y = marks, fill = section)) +
  geom_boxplot(alpha = 0.7, outlier.colour = "red") +
  facet_wrap(~ gender) +
  labs(title = "Marks by section", x = "Section", y = "Marks") +
  theme_minimal() +
  theme(legend.position = "none")

fill (not colour) for the box interior, and the legend removed because the x-axis already labels the sections.


Section D — Long answers

Q11

A classifier reports 94% accuracy. Evaluate it.

Show answer

Given TP = 80, FP = 20, FN = 40, TN = 860:

Metric Value
Accuracy 0.940
Precision 0.800
Recall 0.667
Specificity 0.977
F1 0.727

The verdict: 94% is misleading. 88% of these cases are negative, so "always predict negative" already scores 0.88 — the model adds only 0.06. And recall of 0.667 means it misses 40 of 120 real cases.

For disease screening that is unacceptable: lower the threshold, accept more false positives, raise recall. Report the confusion matrix and per-class recall, never accuracy alone.

(These are the numbers in 14_evaluation.py, which asserts them.)

Q12

Explain the ARIMA workflow on a trending seasonal series

Show answer
  1. Plot it. Growing seasonal swing → multiplicative → take logs.
  2. Decompose — decompose() or stl() to see trend, seasonal, remainder.
  3. Test stationarity — adf.test() (H₀ = non-stationary) and kpss.test() (H₀ = stationary). Opposite nulls; write down which you are using.

  4. Difference — diff() for the trend, diff(lag = 12) for seasonality. ndiffs()/nsdiffs() say how many.

  5. Identify orders — ACF and PACF of the differenced series. PACF gives p, ACF gives q.

  6. Fit — auto.arima() searches by AIC.

  7. Check residuals — checkresiduals(). They must be white noise; you want a large Ljung-Box p-value.

  8. Forecast — forecast(fit, h = 24), back-transform with exp().

A point worth making that most answers miss: step 5 must come after step

  1. On the raw airline series the ACF decays slowly — the trend swamps everything — and the 12-month seasonality is only a ripple on that decay (0.66 at lag 8, back up to 0.76 at lag 12). Only after differencing does the seasonality stand out: after one difference the ACF peaks at lag 12, at 0.83. 16_arima.py shows the same on a series of its own, where the raw ACF decays monotonically and the differenced one troughs at lag 6 and peaks at lag 12, and asserts both patterns.

Corrected: this said the airline series' raw ACF decays monotonically, with the seasonality invisible, and that the differenced ACF troughs at lag 6. Those hold for the Python version's series; R on the airline data gives the figures above.


Quick self-test

  1. Why does v[-1] not give the last element?
  2. What does mean(x >= 40) compute for a numeric vector x?
  3. Which is MARGIN = 2 in apply() — rows or columns?
  4. Why must you scale() before kmeans()?
  5. State the null hypothesis of the ADF test.
  6. Which plot gives the AR order, and which the MA order?
  7. Why is geom_col() sometimes needed instead of geom_bar()?
  8. Why is filtered() written with parentheses in Shiny?

Answers: 1. Negative indexing excludes; use tail(v,1). · 2. The proportion at or above 40, since TRUE counts as 1. · 3. Columns. ·

  1. Distance is scale-dependent; an unscaled large-range variable dominates. ·
  2. Non-stationary (a unit root is present). · 6. PACF gives p, ACF gives q. · 7. geom_bar() counts rows; geom_col() uses a value you supply. ·

  3. It is a reactive — the parentheses fetch its current value.