R offers three major plotting systems:
The generic plot() function dispatches on the class of the first argument:
| Parameter | Meaning | Example |
|---|---|---|
main | Title | main = "Histogram of Marks" |
xlab, ylab | Axis labels | xlab = "Hours" |
col | Colour | col = "skyblue" or col = 2 |
pch | Plotting character (0–25) | pch = 19 (solid dot) |
lty | Line type | lty = 2 (dashed) |
lwd | Line width | lwd = 2 |
xlim, ylim | Axis ranges | xlim = c(0, 100) |
cex | Symbol/text size | cex = 1.2 |
las | Axis labels orientation | las = 1 (horizontal) |
Use hist() to display the distribution of a continuous variable.
data(mtcars)
hist(mtcars$mpg,
breaks = 10,
col = "skyblue",
border = "white",
main = "Distribution of Fuel Economy",
xlab = "Miles per Gallon",
ylab = "Frequency")
# With density curve overlay
hist(mtcars$mpg,
freq = FALSE, # density instead of counts
col = "lightyellow",
main = "MPG Density")
lines(density(mtcars$mpg), col = "red", lwd = 2)
rug(mtcars$mpg) # tick marks under axis
hist(mtcars$mpg) with the actual bin counts, a density() overlay (red) and rug() tick marks showing each individual car under the axis.breaks = 15 — approximately 15 bins.breaks = c(10,15,20,25,30,35) — explicit boundaries.breaks = "Sturges" (default), "Scott", "FD" — algorithms.# Compare two histograms side by side
par(mfrow = c(1, 2))
hist(iris$Sepal.Length, col = "lightblue", main = "Sepal Length")
hist(iris$Petal.Length, col = "lightgreen", main = "Petal Length")
par(mfrow = c(1, 1)) # reset
x <- rnorm(500, mean = 60, sd = 10)
hist(x, freq = FALSE, breaks = 25, col = "lavender",
main = "Sample vs Theoretical Normal", xlab = "Value")
lines(density(x), lwd = 2) # sample density (black)
curve(dnorm(x, mean = 60, sd = 10),
add = TRUE, col = "red", lwd = 2) # theoretical curve (red)
legend("topright", c("Sample density","N(60, 10)"),
col = c("black","red"), lty = 1, lwd = 2, bty = "n")
Display the median, quartiles and outliers. The box spans Q1–Q3; each whisker extends to the most extreme observation still within 1.5 × IQR of the box; anything beyond that is drawn as an individual outlier point.
boxplot(mtcars$mpg,
col = "lightgreen",
main = "Boxplot of MPG",
ylab = "MPG")
# By group (formula syntax)
boxplot(mpg ~ cyl, data = mtcars,
col = c("salmon","skyblue","palegreen"),
main = "MPG by Number of Cylinders",
xlab = "Cylinders", ylab = "MPG",
notch = TRUE) # adds notches around medians
Interpretation: circles beyond the whiskers are outliers — points below Q1 − 1.5 × IQR or above Q3 + 1.5 × IQR (the fences are measured from the quartiles, not the median). Non-overlapping notches suggest the medians differ significantly.
boxplot(): the fences at Q1 − 1.5×IQR and Q3 + 1.5×IQR (dashed) decide which points are flagged as outliers.# All four iris numeric variables on one plot
boxplot(iris[, 1:4],
col = c("pink","skyblue","lightgreen","plum"),
main = "Iris Measurements",
ylab = "cm")
boxplot(Sepal.Length ~ Species, data = iris,
col = "lightyellow",
main = "Sepal Length by Species")
plot(mtcars$wt, mtcars$mpg,
pch = 19, col = "navy",
main = "Fuel Economy vs Weight",
xlab = "Weight (1000 lbs)",
ylab = "MPG")
# Add a regression line
fit <- lm(mpg ~ wt, data = mtcars)
abline(fit, col = "red", lwd = 2)
# Add LOESS smooth
lines(lowess(mtcars$wt, mtcars$mpg), col = "blue", lty = 2, lwd = 2)
legend("topright", c("Linear fit","LOESS"),
col = c("red","blue"), lty = c(1,2), lwd = 2, bty = "n")
abline(fit) (mpg = 37.29 − 5.34 × wt) and the dashed blue curve is the lowess() smooth, which reveals the relationship flattening for heavy cars.pairs(iris[, 1:4], col = iris$Species, pch = 19,
main = "Iris Scatter Plot Matrix")
pch) Referenceplot(1:25, rep(1, 25), pch = 1:25, cex = 2,
main = "Plotting characters 1–25", ylab = "", yaxt = "n")
text(1:25, rep(1.2, 25), labels = 1:25)
pch = 19 (large solid dot) is the usual choice for scatter plots; 21–25 accept a separate fill colour via bg =.plot(Petal.Length ~ Petal.Width, data = iris,
col = iris$Species, pch = 19,
main = "Iris: Petal length vs width")
legend("topleft", levels(iris$Species),
col = 1:3, pch = 19, bty = "n")
data(airquality)
plot(airquality$Day[1:31],
airquality$Temp[1:31],
type = "o", col = "darkred",
main = "May Temperatures (NY)",
xlab = "Day", ylab = "Temp (°F)")
grid()
counts <- table(mtcars$cyl)
barplot(counts,
col = c("salmon","skyblue","palegreen"),
main = "Number of Cars by Cylinders",
xlab = "Cylinders", ylab = "Count")
# Horizontal
barplot(counts, horiz = TRUE)
# Stacked / grouped (two-way table)
tbl <- table(mtcars$gear, mtcars$cyl)
barplot(tbl, beside = TRUE, # set FALSE for stacked
col = c("lightblue","lightgreen","lightyellow"),
main = "Cylinders by Gears",
legend.text = TRUE)
beside = FALSE to stack the bars instead.pie(counts,
col = c("salmon","skyblue","palegreen"),
labels = paste(names(counts), counts),
main = "Cars by Cylinders")
Note: Bar charts are usually preferable to pie charts for comparison.
fruits <- c("Apple","Banana","Cherry","Date")
prices <- c(80, 40, 200, 150)
barplot(prices, names.arg = fruits,
col = "lightcoral",
main = "Fruit Prices (₹/kg)", ylab = "₹/kg")
# Iris counts by species (one per row, but illustrative)
sp <- table(iris$Species)
barplot(sp, col = c("plum","skyblue","palegreen"),
main = "Iris Sample by Species")
After fitting a model with lm() or glm(), R provides four diagnostic plots through plot(model).
fit <- lm(mpg ~ wt + hp, data = mtcars)
par(mfrow = c(2, 2))
plot(fit)
par(mfrow = c(1, 1))
The four panels are:
plot(fit) shows when the model assumptions hold. Trouble signs: a curved red line in panel 1 (non-linearity — see Example 2), S-shaped points in panel 2 (non-normality), a wedge in panel 3 (heteroscedasticity), or points past the Cook's bands in panel 4.# Histogram of residuals
hist(resid(fit), breaks = 12, col = "lightgray",
main = "Distribution of Residuals", xlab = "Residual")
# Q-Q plot for any data
qqnorm(resid(fit)); qqline(resid(fit), col = "red", lwd = 2)
# Shapiro-Wilk test for normality
shapiro.test(resid(fit))
fit <- lm(Sepal.Length ~ Petal.Length, data = iris)
par(mfrow = c(2, 2)); plot(fit); par(mfrow = c(1, 1))
shapiro.test(resid(fit)) # p > 0.05 ⇒ residuals plausibly normal
x <- 1:50
y <- x + x^2/40 + rnorm(50, 0, 3) # quadratic relation + noise
fit <- lm(y ~ x)
par(mfrow = c(1, 2))
plot(x, y, main = "Data + Linear Fit"); abline(fit, col = "red")
plot(fit, which = 1) # curved residual pattern signals non-linearity
par(mfrow = c(1, 1))
# PDF
pdf("myplot.pdf", width = 8, height = 6)
hist(mtcars$mpg, col = "skyblue", main = "MPG")
dev.off()
# PNG
png("myplot.png", width = 800, height = 600, res = 100)
boxplot(mpg ~ cyl, data = mtcars, col = "lightyellow")
dev.off()
# JPEG, SVG, TIFF — similar
svg("myplot.svg"); plot(1:10); dev.off()
viridis package).hist() for distributions; boxplot() for spread & outliers; plot() for scatter; barplot() for categorical.main, xlab, ylab, col, pch, lty, lwd.par(mfrow = c(r, c)).plot(lm_object) auto-generates four diagnostic plots for residual assumption checking.pdf(), png(), etc., followed by dev.off().