R is a free, open-source language and environment for statistical computing and graphics, created by Ross Ihaka and Robert Gentleman (1993) and maintained by the R Core Team. It is the de-facto standard for academic statistics and data science.
Features of R:
ggplot2, dplyr…).Rich package ecosystem (a package for almost every task): data manipulation
(dplyr, data.table), machine learning (caret, xgboost),
text mining (tm), web scraping (rvest), spatial data (sf). R also
supports dynamic reporting — R Markdown weaves code, text and plots into
HTML/PDF/Word, and shiny builds interactive web apps — and interoperates with
other languages (reticulate for Python, Rcpp for C++). These make R a single tool
for academic research, business analytics, and data-science projects.
R is the engine; RStudio is an IDE (integrated development environment) that makes R far easier to use. Its four panes are: Source editor (write/save scripts), Console (run commands), Environment/History (see objects), and Files/Plots/Packages/Help. Install R first (from CRAN), then RStudio.
Installing (all platforms follow the same two steps):
https://cran.r-project.org/): on Windows run the
downloaded .exe; on macOS the .pkg; on Linux use the package manager, e.g.
sudo apt-get install r-base.sudo gdebi rstudio-*.deb).A typical RStudio workflow: set the working directory with
setwd("path/to/project") → write code in the Source pane and run it with
Ctrl+Enter (Windows) / Cmd+Enter (Mac) → inspect objects with
summary(), str() and View() → make plots (they appear in the Plots
tab) → write up results reproducibly with R Markdown.
The assignment operator is <- (preferred) or =. Whatever is on the right is
stored in the name on the left.
# assignment
x <- 10 # leftward (preferred): store 10 in x
y = 5 # equals sign also works (less common)
15 -> z # rightward: value on the left goes into z
w <- x + y # w becomes 15
print(w) # 15
w # typing the name also prints it
., followed by letters,
digits, . or _. It cannot start with a digit.Var and var are different.if, else, TRUE, …)
as names — see §4.name <- "John" # valid
age <- 30 # valid
is_student <- TRUE # valid
2var <- 5 # ERROR: starts with a digit
| Mode | Example | Check with |
|---|---|---|
| numeric (double) | 3.14 | is.numeric() |
| integer | 5L | is.integer() |
| character | "stats" | is.character() |
| logical | TRUE, FALSE | is.logical() |
| complex | 2+3i | is.complex() |
class(3.14) # "numeric"
class(5L) # "integer"
class("stats") # "character"
class(TRUE) # "logical"
# convert (coerce) between types
as.integer("7") # 7
as.character(42) # "42"
A sixth common type is the factor, used for categorical data (it stores the distinct
levels once and records each value as a code):
gender <- factor(c("Male","Female","Female")).
Every object also has a mode — the broad category of what it stores. Modes split into two families:
numeric,
character, logical, complex, and raw (raw bytes, e.g.
charToRaw("R")).list (holds objects of
different modes), function, expression (an unevaluated expression), and
call (an unevaluated function call).mode(5) # "numeric"
mode("Data Science") # "character"
mode(list(1, 2, 3)) # "list"
# class() gives the finer type; mode() gives the storage family
class(5L) # "integer" (but mode(5L) is "numeric")
| Type | Operators |
|---|---|
| Arithmetic | + - * / ^ %% %/% (modulus, integer division) |
| Relational | < > <= >= == != |
| Logical | & | ! && || |
| Assignment | <- = -> |
17 %% 5 # 2 (remainder)
17 %/% 5 # 3 (integer quotient)
2 ^ 10 # 1024
(5 > 3) & (2 > 4) # FALSE
(5 > 3) | (2 > 4) # TRUE
Check if a year is a leap year: (2024 %% 4 == 0) & (2024 %% 100 != 0 | 2024 %% 400 == 0)
returns TRUE.
%% tests divisibility: 15 %% 3 == 0 is TRUE (15 is a multiple of 3);
16 %% 3 == 0 is FALSE.
Miscellaneous and matrix operators:
| Operator | Meaning | Example → Result |
|---|---|---|
: | sequence generation | 1:5 → 1 2 3 4 5 |
%in% | membership test | 3 %in% c(1,2,3) → TRUE |
%*% | matrix multiplication | A %*% B → matrix product |
t() | matrix transpose | t(A) |
When several operators appear in one expression, R evaluates them in this order (highest first):
()^+ and - (sign)* / %/% and %%+ -< > <= >= == !=!& &&| ||3 + 5 * 2 ^ 2 is evaluated as \(3 + 5 \times (2^2) = 3 + 5\times 4 = 3 + 20 = 23\) —
exponent first, then multiply, then add. Use parentheses to force a different order:
(3 + 5) * 2 ^ 2 = 32.
| Value | Meaning |
|---|---|
Inf, -Inf | positive / negative infinity (e.g. 1/0) |
NaN | "Not a Number" (e.g. 0/0) |
NA | missing value (Not Available) |
NULL | the empty / null object (absence of a value) |
TRUE / FALSE | logicals (also T / F); behave as 1 / 0 in arithmetic |
1/0 # Inf
0/0 # NaN
is.na(NA) # TRUE
sum(c(TRUE, TRUE, FALSE)) # 2 (logicals count as 1/0)
Some keywords have fixed meaning and cannot be used as variable names: the control-structure words
if, else, for, while, repeat,
break, next, function, return; the logical constants
TRUE, FALSE, NA, NULL; and Inf,
NaN.
Each special value has a matching test, and helpers exist to clean them out:
is.na(NA) # TRUE – missing
is.null(NULL) # TRUE – empty object
is.infinite(1/0) # TRUE – Inf / -Inf
is.nan(0/0) # TRUE – Not a Number
is.finite(5) # TRUE – neither NA, Inf, nor NaN
na.omit(c(1, NA, 3)) # drops the NA, leaving 1 and 3
| Symbol | Use |
|---|---|
$ | access a named element of a list / column of a data frame (df$age) |
@ | access a slot of an S4 object |
[ ] | extract elements from a vector, matrix or data frame (vec[2]) |
:: | use a function from a package without loading it (stats::sd) |
... | a variable number of arguments passed to a function |
TRUE/FALSE have the shorthands T/F (avoid redefining
them). Because logicals act as 1/0, they are ideal for filtering: a logical vector kept
inside [ ] selects the elements where it is TRUE.
vec <- c(10, 20, 30, 40)
condition <- vec > 20 # FALSE FALSE TRUE TRUE
vec[condition] # 30 40 (logical indexing)
all(vec > 5) # TRUE – every element passes
any(vec > 35) # TRUE – at least one passes
sqrt(81) # 9
abs(-7) # 7
round(3.14159, 2) # 3.14
seq(1, 10, by = 2) # 1 3 5 7 9
rep("hi", 3) # "hi" "hi" "hi"
# getting help
?mean # open help for mean()
help(sd) # same as ?sd
example(sum) # run the documented examples
args(rnorm) # see a function's arguments
Define and call your own function:
area_circle <- function(r) {
pi * r^2
}
area_circle(7) # 153.938
A function with a default argument:
power <- function(x, n = 2) x^n
power(5) # 25 (uses default n = 2)
power(5, 3) # 125
| Structure | Dimensions | Holds |
|---|---|---|
| Vector | 1-D | elements of one type |
| Matrix | 2-D | one type |
| Array | n-D | one type |
| List | 1-D | mixed types |
| Data frame | 2-D | columns of (possibly) different types |
| Factor | 1-D | categorical data with levels |
Vectors, matrices and data frames are covered here and in Unit 5.
# if - else
marks <- 72
if (marks >= 40) {
print("Pass")
} else {
print("Fail")
}
# for loop
for (i in 1:5) print(i^2) # 1 4 9 16 25
# while loop
n <- 1
while (n <= 3) { print(n); n <- n + 1 }
# repeat with break
k <- 1
repeat { print(k); k <- k + 1; if (k > 3) break }
apply family (sapply, lapply) for speed and clarity.
A vector is an ordered collection of elements of the same mode. It is the fundamental data structure in R — even a single number is a vector of length 1.
v1 <- c(4, 8, 15, 16, 23, 42) # combine
v2 <- 1:10 # 1 2 ... 10
v3 <- seq(0, 1, by = 0.25) # 0 0.25 0.50 0.75 1
v4 <- rep(c(1, 2), times = 3) # 1 2 1 2 1 2
length(v1) # 6
v1[1] # 4 (R indexes from 1, not 0)
v1[c(2, 4)] # 8 16
v1[-1] # drop the 1st element
v1[v1 > 15] # 16 23 42 (logical indexing)
names(v1) <- c("a","b","c","d","e","f")
v1["c"] # 15 (index by name)
v <- c(10, 20, 30)
v <- c(v, 40) # add at end -> 10 20 30 40
v <- v[-2] # remove 2nd -> 10 30 40
v + 5 # 15 35 45 (added to each)
v * 2 # 20 60 80
sum(v); mean(v); max(v)
When two vectors of unequal length are combined, R recycles the shorter one to match the longer. If lengths are not multiples, R still recycles but gives a warning.
c(1, 2, 3, 4) + c(10, 20) # 11 22 13 24 (10,20 recycled)
c(1, 2, 3) * 2 # 2 4 6 (scalar recycled to each)
5 %in% c(2, 5, 8) # TRUE (membership)
x <- c(-2, 4, -6, 8)
ifelse(x > 0, "pos", "neg") # "neg" "pos" "neg" "pos"
a <- c(1, 2, 3); b <- c(1, 2, 4)
a == b # TRUE TRUE FALSE (element-wise)
identical(a, b) # FALSE (whole-object comparison)
all(a == b) # FALSE
w <- c(3, NA, 7, NA, 12)
is.na(w) # FALSE TRUE FALSE TRUE FALSE
mean(w) # NA (NA propagates)
mean(w, na.rm = TRUE) # 7.333 (ignore NAs)
sum(is.na(w)) # 2 (count missing)
y <- c(1, NULL, 2) # NULL is dropped -> 1 2
scores <- c(35, 78, 52, 90, 41, 66)
scores[scores >= 50] # 78 52 90 66 (passed)
which(scores >= 50) # 2 3 4 6 (their positions)
scores[scores >= 50 & scores < 80] # 78 52 66
marks <- c(46, 54, 45, 34, 55, 64)
cat("Mean =", mean(marks), "\n") # Mean = 49.67
cat("SD =", sd(marks), "\n") # SD = 10.27
cat("Max =", max(marks), "\n") # Max = 64
sort(marks) # 34 45 46 54 55 64
v1 <- 1:20
v1 <- v1 + 2 # add 2 to every element (recycled scalar)
v1 <- v1 / 5 # divide every element by 5
even <- v1[(1:20) %% 2 == 0] # keep elements at even positions
even
Demonstrates the practical-syllabus task "add 2 to every element, then divide by 5" using recycling, followed by positional filtering.
<-; modes are numeric, integer, character, logical, complex.Inf, NaN, NA (missing), NULL (empty).?fn, help(), example().ifelse, and
na.rm for missing data.which()) is the everyday tool of data cleaning.