Skip to the content

Topics Covered

Probability Distributions d/p/q/r prefixes Random Sampling t-tests χ² test ANOVA p-values & CIs Wilcoxon
On this page
  1. 1. Probability Distributions in R
  2. 2. Random Sampling
  3. 3. t-Tests
  4. 4. Chi-Square Tests
  5. 5. ANOVA
  6. 6. p-values and Confidence Intervals
  7. 7. Wilcoxon (Non-parametric) Tests
  8. 8. Putting it Together — Standard Workflow
  9. Key Take-aways

1. Probability Distributions in R

R uses a uniform d/p/q/r naming convention:

PrefixMeaningExample (Normal)
dDensity / PMFdnorm(0) = 0.3989
pCDF P(X ≤ q)pnorm(1.96) = 0.975
qQuantile (inverse CDF)qnorm(0.975) = 1.96
rRandom samplernorm(10)

Replace norm with any other distribution name: binom, pois, t, chisq, f, exp, unif, gamma, beta, geom, nbinom, hyper, etc.

1.1 Normal Distribution

# Density at x = 1.5 for N(0, 1)
dnorm(1.5)                            # 0.1295

# Density at x for N(60, 10)
dnorm(70, mean = 60, sd = 10)

# P(X <= 1.96)
pnorm(1.96)                           # 0.975

# P(50 <= X <= 70) for N(60, 10)
pnorm(70, 60, 10) - pnorm(50, 60, 10) # 0.6827

# 95th percentile for N(0,1)
qnorm(0.95)                           # 1.645

# Random sample of 1000 from N(70, 5)
set.seed(123)
x <- rnorm(1000, mean = 70, sd = 5)
hist(x, breaks = 30, col = "lightblue", probability = TRUE,
     main = "1000 samples ~ N(70, 5)")
curve(dnorm(x, 70, 5), add = TRUE, col = "red", lwd = 2)
The d / p / q / r family on the normal curve pnorm(q) = P(X ≤ q) dnorm(q) — curve height q = qnorm(p) rnorm(n) — random draws from the curve (green ticks)
d gives the curve height, p the area to the left of q, q goes backwards from a probability to the cut-off, and r draws random values. The same picture applies to every distribution family.

1.2 Binomial Distribution

# P(X = 3) for B(10, 0.5)
dbinom(3, size = 10, prob = 0.5)      # 0.117

# P(X <= 3)
pbinom(3, 10, 0.5)                    # 0.172

# 95th percentile (smallest k with P(X ≤ k) ≥ 0.95)
qbinom(0.95, 10, 0.5)                 # 8

# Random sample
rbinom(15, size = 10, prob = 0.4)     # 15 simulated counts

1.3 Poisson Distribution

# P(X = 5) when λ = 3
dpois(5, lambda = 3)                  # 0.1008

# P(X <= 2)
ppois(2, 3)                           # 0.4232

# 90th percentile
qpois(0.90, 3)                        # 5

# Random sample
rpois(100, lambda = 4)

1.4 Other Common Distributions

dexp(2, rate = 0.5)        # Exponential
punif(0.7)                 # Uniform(0,1)
dt(2, df = 10)             # Student's t
pchisq(11, df = 5)         # Chi-square
qf(0.95, df1 = 4, df2 = 20)  # F-distribution
EXAMPLE 1 — Probabilities under N(μ, σ)
# Marks ~ N(60, 15)
# P(marks > 80)
1 - pnorm(80, mean = 60, sd = 15)                 # 0.0912

# P(45 < marks < 75)
pnorm(75, 60, 15) - pnorm(45, 60, 15)             # 0.683
# i.e., 68.3% of students fall within ±1σ — empirical rule
EXAMPLE 2 — Binomial probability
# A coin tossed 20 times.  P(at most 8 heads) =
pbinom(8, 20, 0.5)                                # 0.2517

2. Random Sampling

# Always set seed first for reproducibility
set.seed(42)

# Sample from N(0,1)
rnorm(5)

# Sample of size 5 from 1..100 without replacement
sample(1:100, 5)

# Sample with replacement
sample(1:6, 10, replace = TRUE)         # 10 die rolls

# Weighted sampling
sample(c("A","B","C"), 50, replace = TRUE, prob = c(0.5, 0.3, 0.2))

Simulating a Sampling Distribution (CLT illustration)

set.seed(1)
N    <- 5000          # number of samples
n    <- 30            # sample size
mu   <- 50; sigma <- 10
means <- replicate(N, mean(rnorm(n, mu, sigma)))

hist(means, breaks = 30, freq = FALSE, col = "skyblue",
     main = "Sampling Distribution of the Mean (n = 30)",
     xlab = expression(bar(X)))
curve(dnorm(x, mu, sigma/sqrt(n)), add = TRUE, col = "red", lwd = 2)
# The histogram tracks the theoretical N(50, 10/sqrt(30)) curve — CLT verified.
Population N(50, 10) Means of samples (n = 30) ≈ N(50, 10/√30 = 1.83) μ = 50 x̄
The CLT simulation drawn: individual values spread with σ = 10 (orange), but sample means (histogram) pile tightly around μ with σ/√n ≈ 1.83 (red). Averaging shrinks spread by √n.

3. t-Tests

3.1 One-sample t-test

data <- c(48, 52, 49, 53, 51, 47, 55, 50, 49, 52)
t.test(data, mu = 50)            # H0: mu = 50
One Sample t-test data: data t = 0.7717, df = 9, p-value = 0.4601 alternative hypothesis: true mean is not equal to 50 95 percent confidence interval: 48.84 52.36 sample estimates: mean of x: 50.6

p > 0.05 ⇒ fail to reject H₀; the mean is consistent with 50. The 95 % CI (48.84, 52.36) contains 50.

Where did our t land? (df = 9, α = 0.05, two-tailed) −2.26 +2.26 reject H₀ reject H₀ observed t = 0.77 p = 0.46 ⇒ keep H₀ 0
To reject H₀ at α = 0.05 the t statistic must fall in a red tail (beyond ±2.26 for 9 df). Our t = 0.77 sits deep in the middle — hence the large p-value of 0.46.

3.2 Independent Two-sample t-test

group1 <- c(72, 75, 78, 80, 76)
group2 <- c(65, 68, 70, 72, 67)

# Default: Welch's t-test (unequal variances)
t.test(group1, group2)

# Pooled (equal variances) — set var.equal = TRUE
t.test(group1, group2, var.equal = TRUE)

# One-tailed
t.test(group1, group2, alternative = "greater")

3.3 Paired t-test

before <- c(72, 78, 69, 80, 85, 76, 82, 74)
after  <- c(70, 76, 67, 78, 83, 74, 79, 72)

t.test(before, after, paired = TRUE)
# Equivalent to: t.test(before - after)
EXAMPLE 1 — Two-sample on iris
setosa     <- iris$Sepal.Length[iris$Species == "setosa"]
versicolor <- iris$Sepal.Length[iris$Species == "versicolor"]
t.test(setosa, versicolor)              # p < 2.2e-16 ⇒ reject H0
EXAMPLE 2 — Paired (sleep dataset)
data(sleep)                              # 10 subjects, 2 drugs
t.test(extra ~ group, data = sleep, paired = TRUE)

4. Chi-Square Tests

4.1 Goodness of Fit

# Die rolled 60 times, observed face counts
observed <- c(8, 11, 9, 12, 10, 10)
chisq.test(observed)                    # default: equal expected probabilities
Chi-squared test for given probabilities data: observed X-squared = 1, df = 5, p-value = 0.9626

p ≈ 0.96 ⇒ fail to reject H₀; the die appears fair.

With explicit expected probabilities (Mendelian 9:3:3:1):

observed <- c(460, 140, 130, 70)
chisq.test(observed, p = c(9, 3, 3, 1)/16)

4.2 Independence Test (Contingency Table)

tbl <- matrix(c(60, 40,
                20, 80), nrow = 2, byrow = TRUE,
              dimnames = list(c("Smoker","Non-smoker"),
                              c("Cancer","No Cancer")))
chisq.test(tbl)
Pearson's Chi-squared test with Yates' continuity correction data: tbl X-squared = 31.688, df = 1, p-value = 1.81e-08

p < 0.001 ⇒ reject independence; smoking and cancer are associated.

For small expected counts (any cell < 5), use Fisher's Exact Test: fisher.test(tbl).

4.3 χ² Test for Single Variance

# H0: sigma^2 = sigma0^2; manual calculation
n        <- 25
s2       <- 18
sigma0_2 <- 16
chi2     <- (n - 1) * s2 / sigma0_2
chi2                                    # 27
# Two-tailed p-value
2 * min(pchisq(chi2, n - 1), 1 - pchisq(chi2, n - 1))

5. ANOVA

5.1 One-way ANOVA

# Iris: Sepal.Length differs across 3 Species?
fit <- aov(Sepal.Length ~ Species, data = iris)
summary(fit)
Df Sum Sq Mean Sq F value Pr(>F) Species 2 63.21 31.606 119.3 <2e-16 *** Residuals 147 38.96 0.265 --- Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

p < 0.001 ⇒ at least one species mean differs.

between-group variation (Sum Sq 63.2) within-group (Sum Sq 39.0) 5.01 5.94 6.59 setosa versicolor virginica
What the F test compares: how far the three group means (5.01, 5.94, 6.59) sit apart between groups versus the spread within each group. F = 31.61 / 0.265 = 119.3 — the separation dwarfs the noise.

5.2 Post-hoc — Tukey's HSD

TukeyHSD(fit)
diff lwr upr p adj versicolor-setosa 0.930 0.686 1.174 0.000 virginica-setosa 1.582 1.338 1.826 0.000 virginica-versicolor 0.652 0.408 0.896 0.000

All three pairwise differences are significant.

5.3 Two-way ANOVA

data(ToothGrowth)             # built-in
fit2 <- aov(len ~ supp * dose, data = ToothGrowth)   # main + interaction
summary(fit2)

# Without interaction
aov(len ~ supp + factor(dose), data = ToothGrowth) |> summary()

5.4 F-test for Equality of Variances

var.test(group1, group2)        # equivalent to F-test
# Or
var.test(extra ~ group, data = sleep)
EXAMPLE 1 — ANOVA on mtcars (mpg by cyl)
fit <- aov(mpg ~ factor(cyl), data = mtcars)
summary(fit)                       # F = 39.7, p < 0.001 ⇒ cyl matters
TukeyHSD(fit)
EXAMPLE 2 — Two-way with interaction
data(ToothGrowth)
fit <- aov(len ~ supp * factor(dose), data = ToothGrowth)
summary(fit)
# main effects of supp & dose AND their interaction tested in one shot

6. p-values and Confidence Intervals

All major R test functions return a list containing statistic, p.value, conf.int, estimate etc., which you can extract programmatically:

res <- t.test(group1, group2)
res$statistic       # t value
res$parameter       # degrees of freedom
res$p.value         # p-value
res$conf.int        # 95 % CI for the difference of means
res$estimate        # sample means

You can also compute CIs manually:

# 95 % CI for the mean
x  <- rnorm(50, 100, 15)
m  <- mean(x); se <- sd(x)/sqrt(length(x))
m + c(-1, 1) * qt(0.975, df = length(x) - 1) * se

7. Wilcoxon (Non-parametric) Tests

Use when normality assumption of t-test is violated.

7.1 Wilcoxon Signed-Rank (one sample / paired)

# One-sample
data <- c(48, 52, 49, 53, 51, 47, 55, 50, 49, 52)
wilcox.test(data, mu = 50)              # H0: median = 50

# Paired (Wilcoxon signed-rank)
wilcox.test(before, after, paired = TRUE)

7.2 Wilcoxon Rank-Sum / Mann-Whitney U (two independent)

wilcox.test(group1, group2)
# Equivalent to Mann-Whitney U test
wilcox.test(extra ~ group, data = sleep)

Comparison with Parametric Counterpart

# Both on the same data — does the conclusion agree?
t.test(group1, group2)$p.value
wilcox.test(group1, group2)$p.value
EXAMPLE 1 — Wilcoxon paired vs paired t
before <- c(72, 78, 69, 80, 85, 76, 82, 74)
after  <- c(70, 76, 67, 78, 83, 74, 79, 72)

t.test(before, after, paired = TRUE)$p.value      # tiny
wilcox.test(before, after, paired = TRUE)$p.value # also tiny — both reject H0
EXAMPLE 2 — Two-sample on highly skewed data
income_men   <- c(20, 22, 28, 30, 35, 50, 80, 120)
income_women <- c(18, 19, 25, 27, 30, 40, 60, 90)

shapiro.test(income_men)$p.value      # < 0.05 ⇒ normality doubtful
# So prefer Wilcoxon over t-test
wilcox.test(income_men, income_women)

8. Putting it Together — Standard Workflow

  1. Define hypotheses (H₀ and H₁).
  2. Check assumptions:
    • Normality: shapiro.test(), Q-Q plot.
    • Equal variances: var.test() or bartlett.test().
  3. Choose appropriate test (parametric if assumptions OK, else non-parametric).
  4. Run the test, examine p-value and CI.
  5. State conclusion in plain language.
# Example workflow
shapiro.test(group1); shapiro.test(group2)
var.test(group1, group2)
if (both_p > 0.05) {
  t.test(group1, group2, var.equal = TRUE)
} else {
  wilcox.test(group1, group2)
}

Key Take-aways