Skip to the content

Topics Covered

Central Tendency Variability Skewness / Kurtosis Quantiles summary() describe() Frequency Tables Cross-tabulations Contingency Tables
On this page
  1. 1. Measures of Central Tendency
  2. 2. Measures of Variability
  3. 3. Quantiles & Percentiles
  4. 4. Skewness & Kurtosis
  5. 5. Quick Summary Functions
  6. 6. Handling Categorical Data
  7. 7. A Mini Descriptive-Statistics Workflow
  8. Key Take-aways

1. Measures of Central Tendency

StatisticR FunctionNotes
Arithmetic Meanmean(x)mean(x, trim = 0.1) for trimmed mean
Medianmedian(x)Robust to outliers
Modeno built-inSee user function below
Geometric Meanpsych::geometric.mean(x)or exp(mean(log(x)))
Harmonic Meanpsych::harmonic.mean(x)or 1/mean(1/x)
Weighted Meanweighted.mean(x, w)

R has no built-in mode function (the function mode() returns data type). Define your own:

get_mode <- function(v) {
  uv <- unique(v)
  uv[which.max(tabulate(match(v, uv)))]
}
EXAMPLE 1
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
mean(marks)                  # 54.42
median(marks)                # 53.5
get_mode(marks)              # 49

# Trimmed mean (drop top & bottom 10%)
mean(marks, trim = 0.10)     # 54.8  (drops lowest 33 and highest 72)
mean (x̄) median Original marks data 53.5 x̄ 54.4 304050 6070 Replace 72 with 150 (one outlier) outlier 150 still 53.5 x̄ jumps to 60.9 304050 6070
Why the median is robust: one extreme value drags the mean from 54.4 to 60.9, but the median does not move. This is also why trim = exists for mean().
EXAMPLE 2 — Growth rates
growth <- c(1.10, 1.20, 1.30)        # 10%, 20%, 30%
# install.packages("psych")
library(psych)
geometric.mean(growth)               # 1.197 → avg growth ≈ 19.7%
harmonic.mean(c(40, 60))             # 48 — avg speed for equal distances

2. Measures of Variability

StatisticR FunctionDivisor
Rangerange(x) returns min and max; diff(range(x))—
Variance (sample)var(x)n − 1
Standard Deviation (sample)sd(x)n − 1
IQRIQR(x)—
MAD (median abs. dev)mad(x)—
Coefficient of Variationsd(x) / mean(x) * 100%
EXAMPLE 1
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
diff(range(marks))           # 39
var(marks)                   # 157.17
sd(marks)                    # 12.54
IQR(marks)                   # 16.75
sd(marks)/mean(marks)*100    # CV ≈ 23.04%
EXAMPLE 2 — Comparing CV across variables
data(iris)
sapply(iris[, 1:4],
       function(x) sd(x)/mean(x)*100)
# Sepal.Length  Sepal.Width  Petal.Length  Petal.Width
#      14.17       14.26         46.97        63.55

Petal.Width has the highest relative variability.

smaller SD → smaller CV larger SD → larger CV same mean
Central tendency alone can hide the story: both curves have the same mean and differ only in spread. var, sd, IQR and CV measure exactly this — recall Petal.Width had the largest CV in iris.

3. Quantiles & Percentiles

quantile(marks)                        # default 0, 25, 50, 75, 100%
quantile(marks, probs = c(0.1, 0.9))   # 10th & 90th percentiles

fivenum(marks)                         # Tukey's five-number summary
# 33.0 47.0 53.5 65.5 72.0   (Tukey hinges — can differ slightly from quantile())
summary(marks)                         # min, Q1, median, mean, Q3, max
> summary(marks) Min. 1st Qu. Median Mean 3rd Qu. Max. 33.00 48.00 53.50 54.42 64.75 72.00
Median 53.5 Min 33 Q1 48 Q3 64.75 Max 72 IQR = 64.75 − 48 = 16.75 (middle 50% of the data) 304050 6070
The five-number summary of marks drawn as a boxplot: the box spans Q1–Q3, the red line is the median, whiskers reach min and max. Note fivenum() uses Tukey hinges (47, 65.5), slightly different from quantile()'s 48, 64.75.

4. Skewness & Kurtosis

Base R does not include these — load the moments or psych package.

# install.packages("moments")
library(moments)

skewness(marks)              # γ₁
kurtosis(marks)              # β₂ (note: returns kurtosis, not excess)
mean median mode Left-skewed (γ₁ < 0) Mean < Median < Mode Symmetric (γ₁ ≈ 0) Mean = Median = Mode Right-skewed (γ₁ > 0) Mode < Median < Mean
The sign of skewness γ₁ tells you which way the long tail points — and the tail is what drags the mean away from the mode. The income example below is the right-hand picture.
EXAMPLE 1
x <- c(5, 6, 7, 7, 8, 8, 8, 9, 9, 10)
skewness(x)        # −0.30 (mild negative / left skew)
kurtosis(x)        # 2.37   → β₂ < 3 ⇒ platykurtic
EXAMPLE 2 — Right-skewed income
income <- c(20, 22, 25, 28, 30, 35, 50, 80, 120)
skewness(income)   # 1.41 → strong positive skew
kurtosis(income)   # 3.66 → mildly leptokurtic
Leptokurtic — β₂ > 3 Mesokurtic — β₂ = 3 Platykurtic — β₂ < 3
Kurtosis β₂ compares peakedness and tail weight against the normal (mesokurtic) curve: leptokurtic = sharp peak with heavy tails, platykurtic = flat top with light tails. moments::kurtosis() returns β₂ itself — subtract 3 for excess kurtosis γ₂.

5. Quick Summary Functions

5.1 Base R — summary()

data(iris)
summary(iris)
#  Sepal.Length    Sepal.Width    Petal.Length   Petal.Width        Species
#  Min.   :4.30   Min.   :2.00   Min.   :1.00   Min.   :0.10   setosa    :50
#  1st Qu.:5.10   1st Qu.:2.80   1st Qu.:1.60   1st Qu.:0.30   versicolor:50
#  Median :5.80   Median :3.00   Median :4.35   Median :1.30   virginica :50
#  Mean   :5.84   Mean   :3.06   Mean   :3.76   Mean   :1.20
#  3rd Qu.:6.40   3rd Qu.:3.30   3rd Qu.:5.10   3rd Qu.:1.80
#  Max.   :7.90   Max.   :4.40   Max.   :6.90   Max.   :2.50

5.2 psych — describe()

library(psych)
describe(iris[, 1:4])
#              vars   n mean   sd median trimmed  mad  min  max range  skew  kurt   se
# Sepal.Length    1 150 5.84 0.83  5.80    5.81 1.04 4.30 7.90  3.60  0.31 -0.61 0.07
# Sepal.Width     2 150 3.06 0.44  3.00    3.04 0.44 2.00 4.40  2.40  0.31  0.14 0.04
# Petal.Length    3 150 3.76 1.77  4.35    3.76 1.85 1.00 6.90  5.90 -0.27 -1.42 0.14
# Petal.Width     4 150 1.20 0.76  1.30    1.18 1.04 0.10 2.50  2.40 -0.10 -1.36 0.06

describeBy() gives group-wise summaries (describe.by() is the older, deprecated name).

5.3 Group-wise Summaries

# By Species in iris
describeBy(iris[, 1:4], group = iris$Species)

# Using base R
aggregate(. ~ Species, data = iris, FUN = mean)
#      Species Sepal.Length Sepal.Width Petal.Length Petal.Width
# 1    setosa        5.006       3.428        1.462       0.246
# 2 versicolor       5.936       2.770        4.260       1.326
# 3  virginica       6.588       2.974        5.552       2.026

# Using dplyr
library(dplyr)
iris %>%
  group_by(Species) %>%
  summarise(mean_PL = mean(Petal.Length),
            sd_PL   = sd(Petal.Length),
            n       = n())

6. Handling Categorical Data

6.1 Frequency Tables

# One-way table
gender <- c("M","F","F","M","M","F","M","F","F","M")
table(gender)
# F M
# 5 5

# Proportions
prop.table(table(gender))
#   F   M
# 0.5 0.5

# Percent
round(prop.table(table(gender)) * 100, 1)
#   F   M
# 50  50

6.2 Two-way (Cross-tabulation)

result <- c("Pass","Pass","Fail","Pass","Fail","Pass","Pass","Fail","Pass","Pass")
table(gender, result)
#       result
# gender Fail Pass
#      F    2    3
#      M    1    4

# Row percentages
prop.table(table(gender, result), margin = 1) * 100
# Column percentages: margin = 2
Fail Pass 12 34 2 3 1 4 F M
The table(gender, result) cross-tab as side-by-side bars. Row percentages (margin = 1) compare pass rates directly: F = 3/5 = 60%, M = 4/5 = 80%.

6.3 Three-way (Contingency Cube)

section <- c("A","A","B","B","A","B","A","B","A","B")
table(gender, result, section)
# , , section = A …

6.4 Using xtabs() (formula style)

df <- data.frame(gender, result, section)
xtabs(~ gender + result, data = df)
xtabs(~ gender + result + section, data = df)

6.5 Margins & Adding Totals

tbl <- table(gender, result)
addmargins(tbl)
#       result
# gender Fail Pass Sum
#      F    2    3   5
#      M    1    4   5
#      Sum  3    7  10

6.6 Frequency of a Continuous Variable (binning)

marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72, 55, 49)
bins  <- cut(marks,
             breaks = c(0, 35, 50, 65, 80, 100),
             labels = c("F","D","C","B","A"),
             right  = TRUE)
table(bins)
# F D C B A
# 1 4 4 3 0
EXAMPLE 1 (Cross-tabulation iris)
data(iris)
big_petal <- ifelse(iris$Petal.Length > 4, "Big", "Small")
table(iris$Species, big_petal)
#              big_petal
#              Big Small
# setosa         0    50
# versicolor    34    16
# virginica     50     0
EXAMPLE 2 (Mendelian ratio)
observed <- c(rndYellow = 460, rndGreen = 140, wriYellow = 130, wriGreen = 70)
expected <- c(450, 150, 150, 50)

chi2 <- sum((observed - expected)^2 / expected)   # 11.56
1 - pchisq(chi2, df = 3)                          # p-value ≈ 0.009

7. A Mini Descriptive-Statistics Workflow

data(mtcars)

# Step 1: Structure
str(mtcars)

# Step 2: Quick summary
summary(mtcars)

# Step 3: Detailed
library(psych)
describe(mtcars)

# Step 4: By group (cylinders)
describeBy(mtcars$mpg, group = mtcars$cyl)

# Step 5: Correlations between numeric variables
round(cor(mtcars[, c("mpg","hp","wt","disp")]), 2)
#       mpg    hp    wt  disp
# mpg  1.00 -0.78 -0.87 -0.85
# hp  -0.78  1.00  0.66  0.79
# wt  -0.87  0.66  1.00  0.89
# disp -0.85  0.79  0.89  1.00

Key Take-aways