| Statistic | R Function | Notes |
|---|---|---|
| Arithmetic Mean | mean(x) | mean(x, trim = 0.1) for trimmed mean |
| Median | median(x) | Robust to outliers |
| Mode | no built-in | See user function below |
| Geometric Mean | psych::geometric.mean(x) | or exp(mean(log(x))) |
| Harmonic Mean | psych::harmonic.mean(x) | or 1/mean(1/x) |
| Weighted Mean | weighted.mean(x, w) |
R has no built-in mode function (the function mode() returns data type). Define your own:
get_mode <- function(v) {
uv <- unique(v)
uv[which.max(tabulate(match(v, uv)))]
}
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
mean(marks) # 54.42
median(marks) # 53.5
get_mode(marks) # 49
# Trimmed mean (drop top & bottom 10%)
mean(marks, trim = 0.10) # 54.8 (drops lowest 33 and highest 72)
trim = exists for mean().growth <- c(1.10, 1.20, 1.30) # 10%, 20%, 30%
# install.packages("psych")
library(psych)
geometric.mean(growth) # 1.197 → avg growth ≈ 19.7%
harmonic.mean(c(40, 60)) # 48 — avg speed for equal distances
| Statistic | R Function | Divisor |
|---|---|---|
| Range | range(x) returns min and max; diff(range(x)) | — |
| Variance (sample) | var(x) | n − 1 |
| Standard Deviation (sample) | sd(x) | n − 1 |
| IQR | IQR(x) | — |
| MAD (median abs. dev) | mad(x) | — |
| Coefficient of Variation | sd(x) / mean(x) * 100 | % |
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
diff(range(marks)) # 39
var(marks) # 157.17
sd(marks) # 12.54
IQR(marks) # 16.75
sd(marks)/mean(marks)*100 # CV ≈ 23.04%
data(iris)
sapply(iris[, 1:4],
function(x) sd(x)/mean(x)*100)
# Sepal.Length Sepal.Width Petal.Length Petal.Width
# 14.17 14.26 46.97 63.55
Petal.Width has the highest relative variability.
var, sd, IQR and CV measure exactly this — recall Petal.Width had the largest CV in iris.quantile(marks) # default 0, 25, 50, 75, 100%
quantile(marks, probs = c(0.1, 0.9)) # 10th & 90th percentiles
fivenum(marks) # Tukey's five-number summary
# 33.0 47.0 53.5 65.5 72.0 (Tukey hinges — can differ slightly from quantile())
summary(marks) # min, Q1, median, mean, Q3, max
marks drawn as a boxplot: the box spans Q1–Q3, the red line is the median, whiskers reach min and max. Note fivenum() uses Tukey hinges (47, 65.5), slightly different from quantile()'s 48, 64.75.Base R does not include these — load the moments or psych package.
# install.packages("moments")
library(moments)
skewness(marks) # γ₁
kurtosis(marks) # β₂ (note: returns kurtosis, not excess)
income example below is the right-hand picture.x <- c(5, 6, 7, 7, 8, 8, 8, 9, 9, 10)
skewness(x) # −0.30 (mild negative / left skew)
kurtosis(x) # 2.37 → β₂ < 3 ⇒ platykurtic
income <- c(20, 22, 25, 28, 30, 35, 50, 80, 120)
skewness(income) # 1.41 → strong positive skew
kurtosis(income) # 3.66 → mildly leptokurtic
moments::kurtosis() returns β₂ itself — subtract 3 for excess kurtosis γ₂.summary()data(iris)
summary(iris)
# Sepal.Length Sepal.Width Petal.Length Petal.Width Species
# Min. :4.30 Min. :2.00 Min. :1.00 Min. :0.10 setosa :50
# 1st Qu.:5.10 1st Qu.:2.80 1st Qu.:1.60 1st Qu.:0.30 versicolor:50
# Median :5.80 Median :3.00 Median :4.35 Median :1.30 virginica :50
# Mean :5.84 Mean :3.06 Mean :3.76 Mean :1.20
# 3rd Qu.:6.40 3rd Qu.:3.30 3rd Qu.:5.10 3rd Qu.:1.80
# Max. :7.90 Max. :4.40 Max. :6.90 Max. :2.50
describe()library(psych)
describe(iris[, 1:4])
# vars n mean sd median trimmed mad min max range skew kurt se
# Sepal.Length 1 150 5.84 0.83 5.80 5.81 1.04 4.30 7.90 3.60 0.31 -0.61 0.07
# Sepal.Width 2 150 3.06 0.44 3.00 3.04 0.44 2.00 4.40 2.40 0.31 0.14 0.04
# Petal.Length 3 150 3.76 1.77 4.35 3.76 1.85 1.00 6.90 5.90 -0.27 -1.42 0.14
# Petal.Width 4 150 1.20 0.76 1.30 1.18 1.04 0.10 2.50 2.40 -0.10 -1.36 0.06
describeBy() gives group-wise summaries (describe.by() is the older, deprecated name).
# By Species in iris
describeBy(iris[, 1:4], group = iris$Species)
# Using base R
aggregate(. ~ Species, data = iris, FUN = mean)
# Species Sepal.Length Sepal.Width Petal.Length Petal.Width
# 1 setosa 5.006 3.428 1.462 0.246
# 2 versicolor 5.936 2.770 4.260 1.326
# 3 virginica 6.588 2.974 5.552 2.026
# Using dplyr
library(dplyr)
iris %>%
group_by(Species) %>%
summarise(mean_PL = mean(Petal.Length),
sd_PL = sd(Petal.Length),
n = n())
# One-way table
gender <- c("M","F","F","M","M","F","M","F","F","M")
table(gender)
# F M
# 5 5
# Proportions
prop.table(table(gender))
# F M
# 0.5 0.5
# Percent
round(prop.table(table(gender)) * 100, 1)
# F M
# 50 50
result <- c("Pass","Pass","Fail","Pass","Fail","Pass","Pass","Fail","Pass","Pass")
table(gender, result)
# result
# gender Fail Pass
# F 2 3
# M 1 4
# Row percentages
prop.table(table(gender, result), margin = 1) * 100
# Column percentages: margin = 2
table(gender, result) cross-tab as side-by-side bars. Row percentages (margin = 1) compare pass rates directly: F = 3/5 = 60%, M = 4/5 = 80%.section <- c("A","A","B","B","A","B","A","B","A","B")
table(gender, result, section)
# , , section = A …
xtabs() (formula style)df <- data.frame(gender, result, section)
xtabs(~ gender + result, data = df)
xtabs(~ gender + result + section, data = df)
tbl <- table(gender, result)
addmargins(tbl)
# result
# gender Fail Pass Sum
# F 2 3 5
# M 1 4 5
# Sum 3 7 10
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
bins <- cut(marks,
breaks = c(0, 35, 50, 65, 80, 100),
labels = c("F","D","C","B","A"),
right = TRUE)
table(bins)
# F D C B A
# 1 4 4 3 0
data(iris)
big_petal <- ifelse(iris$Petal.Length > 4, "Big", "Small")
table(iris$Species, big_petal)
# big_petal
# Big Small
# setosa 0 50
# versicolor 34 16
# virginica 50 0
observed <- c(rndYellow = 460, rndGreen = 140, wriYellow = 130, wriGreen = 70)
expected <- c(450, 150, 150, 50)
chi2 <- sum((observed - expected)^2 / expected) # 11.56
1 - pchisq(chi2, df = 3) # p-value ≈ 0.009
data(mtcars)
# Step 1: Structure
str(mtcars)
# Step 2: Quick summary
summary(mtcars)
# Step 3: Detailed
library(psych)
describe(mtcars)
# Step 4: By group (cylinders)
describeBy(mtcars$mpg, group = mtcars$cyl)
# Step 5: Correlations between numeric variables
round(cor(mtcars[, c("mpg","hp","wt","disp")]), 2)
# mpg hp wt disp
# mpg 1.00 -0.78 -0.87 -0.85
# hp -0.78 1.00 0.66 0.79
# wt -0.87 0.66 1.00 0.89
# disp -0.85 0.79 0.89 1.00
mean, median, var, sd, IQR, quantile for the basics — all accept na.rm = TRUE.summary() for any data frame gives min/Q1/median/mean/Q3/max per column; for factors it shows counts.psych::describe() gives n, mean, sd, median, trimmed, mad, min, max, range, skew, kurt, se — all in one shot.kurtosis() returns β₂ (not excess); subtract 3 to get excess kurtosis (γ₂).table() for frequencies; prop.table() for percentages; xtabs() for formula-style cross-tabs.aggregate(), describeBy(), or dplyr::group_by + summarise.