R is a free, open-source language and environment for statistical computing and graphics, first released in 1993. It is maintained by the R Core Team and the wider community via CRAN (Comprehensive R Archive Network).
https://cran.r-project.org; download the installer for your operating system.https://posit.co/download/rstudio-desktop/ — an integrated development environment (IDE) for R.When R starts, you see a prompt >. Type any expression; press Enter to evaluate.
> 2 + 3
[1] 5
> sqrt(64)
[1] 8
> "Hello, R!"
[1] "Hello, R!"
To assign: x <- 10 or x = 10. To list objects: ls(). To remove: rm(x).
?mean or help("mean") — opens help page for a function.example(mean) — runs the documented examples.apropos("plot") — list all functions containing "plot".Before working with the data structures in the next section, it helps to know how R stores values, the operators it provides, and its special values.
Assignment stores a value in a name. R offers leftward <- (preferred), the equals sign
=, and rightward ->.
x <- 10 # leftward (preferred)
y = 5 # equals sign (valid, less common)
15 -> z # rightward
z # 15
Naming rules: a name must start with a letter or a period ., followed by
letters, digits, . or _; it cannot start with a digit; names are case-sensitive
(Var and var differ); and reserved words (§2.5) cannot be used as names.
| Type | Operators |
|---|---|
| Arithmetic | + - * / ^, %% (modulus), %/% (integer division) |
| Relational | < > <= >= == != (return TRUE/FALSE) |
| Logical | & | (element-wise), && || (single value), ! (not) |
| Assignment | <- = -> |
| Miscellaneous | : (sequence), %in% (membership) |
| Matrix | %*% (matrix product), t() (transpose) |
17 %% 5 # 2 (remainder)
17 %/% 5 # 3 (integer quotient)
2 ^ 10 # 1024
3 %in% c(1, 2, 3) # TRUE
(5 > 3) & (2 > 4) # FALSE
Leap-year test with logical operators:
(2024 %% 4 == 0) & (2024 %% 100 != 0 | 2024 %% 400 == 0) returns TRUE.
In a compound expression R evaluates operators in this order (highest first):
()^+ -* / %/% %%+ -< > <= >= == !=! & && | ||3 + 5 * 2 ^ 2 gives \(3 + 5\times(2^2) = 3 + 20 = 23\) — exponent, then multiply, then add.
Parentheses override: (3 + 5) * 2 ^ 2 = 32.
Every object has a mode — the family of what it stores. Atomic modes hold
elements of a single basic type: numeric, character, logical,
complex, raw. Non-atomic modes hold richer objects: list
(mixed contents), function, expression, call. Use mode()
to check; the individual data types are detailed in the next section.
mode(5) # "numeric"
mode("stats") # "character"
mode(list(1, "a")) # "list"
| Value | Meaning | Test |
|---|---|---|
Inf, -Inf | infinity, e.g. 1/0 | is.infinite() |
NaN | Not a Number, e.g. 0/0 | is.nan() |
NA | missing value (detailed later in the Handling Missing Values section) | is.na() |
NULL | the empty object (absence of a value) | is.null() |
Reserved words cannot be used as names: if else for while repeat break next
function return TRUE FALSE NA NULL Inf NaN.
Special characters: $ (element of a list / column of a data frame,
df$age), @ (slot of an S4 object), [ ] (extract elements),
:: (a function from a package, stats::sd), and ... (a variable number
of function arguments).
TRUE/FALSE (shorthands T/F) behave as 1/0 in
arithmetic, so they drive filtering: a logical vector inside [ ] keeps the
TRUE positions.
vec <- c(10, 20, 30, 40)
vec[vec > 20] # 30 40 (logical indexing)
all(vec > 5) # TRUE – every element passes
any(vec > 35) # TRUE – at least one passes
sum(c(TRUE, TRUE, FALSE)) # 2
The fundamental data structure. All elements must be of the same type.
| Type | Example | Check function |
|---|---|---|
| numeric (double) | 3.14 | is.numeric() |
| integer | 5L | is.integer() |
| character | "R" | is.character() |
| logical | TRUE / FALSE | is.logical() |
| complex | 1+2i | is.complex() |
x <- c(2, 4, 6, 8, 10) # numeric vector
y <- 1:10 # sequence 1..10
z <- seq(0, 1, by = 0.25) # 0.00 0.25 0.50 0.75 1.00
w <- rep("A", 5) # "A" "A" "A" "A" "A"
b <- c(TRUE, FALSE, TRUE) # logical
length(x) # 5
class(x) # "numeric"
x <- c(1, 2, 3, 4)
x * 2 # 2 4 6 8
x + c(10, 20) # recycles: 11 22 13 24
sum(x); mean(x); sd(x)
Create a vector of 10 student marks and compute the mean & max:
marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72)
mean(marks) # 54.9
max(marks) # 72
Convert temperatures from Celsius to Fahrenheit:
celsius <- c(0, 10, 20, 30, 40)
fahrenheit <- celsius * 9/5 + 32
fahrenheit # 32 50 68 86 104
Two-dimensional, all elements same type.
M <- matrix(1:12, nrow = 3, ncol = 4)
M
# [,1] [,2] [,3] [,4]
# [1,] 1 4 7 10
# [2,] 2 5 8 11
# [3,] 3 6 9 12
dim(M) # 3 4
M[2, 3] # 8
M[ , 2] # second column → 4 5 6
M[1, ] # first row → 1 4 7 10
t(M) # transpose
M %*% t(M) # matrix multiplication
Can contain elements of different types.
person <- list(name = "Aarav", age = 19, marks = c(85, 90, 78))
person$name # "Aarav"
person$marks[2] # 90
Workhorse for statistical data — like a spreadsheet table; columns can be of different types.
df <- data.frame(
ID = 1:5,
Name = c("Aarav","Bhavna","Chetan","Divya","Esha"),
Marks = c(78, 82, 65, 90, 88),
Pass = c(TRUE, TRUE, TRUE, TRUE, TRUE)
)
df
str(df) # structure
head(df, 3) # first 3 rows
tail(df, 2) # last 2 rows
nrow(df); ncol(df); dim(df)
gender <- factor(c("M", "F", "F", "M", "M"))
levels(gender) # "F" "M"
table(gender) # F:2 M:3
# Ordered factor
grade <- factor(c("B","A","C","A","B"), levels = c("C","B","A"), ordered = TRUE)
# Reading
data <- read.csv("students.csv") # comma-separated
data <- read.csv("students.csv", header = TRUE, na.strings = "")
# Writing
write.csv(data, "output.csv", row.names = FALSE)
# install.packages("readxl")
library(readxl)
data <- read_excel("students.xlsx", sheet = 1)
data <- read_excel("students.xlsx", sheet = "Marks", range = "A1:D50")
# Writing (use writexl or openxlsx)
# install.packages("writexl")
library(writexl)
write_xlsx(data, "output.xlsx")
data <- read.table("data.txt", header = TRUE, sep = "\t")
data <- read.delim("data.txt") # tab default
# install.packages("DBI"); install.packages("RSQLite")
library(DBI); library(RSQLite)
con <- dbConnect(SQLite(), "school.db")
data <- dbGetQuery(con, "SELECT * FROM students WHERE marks > 60")
dbDisconnect(con)
data() # list all built-in datasets
data(iris)
head(iris) # 4 numeric features + Species factor
data(mtcars) # car attributes
data(airquality) # NY air-quality measurements
R represents missing data as NA.
x <- c(2, 4, NA, 8, NA, 12)
is.na(x) # FALSE FALSE TRUE FALSE TRUE FALSE
sum(is.na(x)) # 2 missing values
mean(x) # NA ← because NA propagates
mean(x, na.rm = TRUE) # 6.5 (NAs removed)
# Drop rows with any NA in a data frame
complete <- na.omit(df)
# Replace NA with the column mean
x[is.na(x)] <- mean(x, na.rm = TRUE)
Identify and replace missing values in airquality$Ozone:
data(airquality)
sum(is.na(airquality$Ozone)) # 37
airquality$Ozone[is.na(airquality$Ozone)] <- mean(airquality$Ozone, na.rm = TRUE)
sum(is.na(airquality$Ozone)) # 0
Drop rows with any missing values in a custom data frame:
df_complete <- na.omit(df)
nrow(df) - nrow(df_complete) # number of rows dropped
x <- c(10, 20, 30, 40, 50)
x[1] # 10
x[c(2, 4)] # 20 40
x[-1] # everything except element 1
x[x > 25] # logical filter → 30 40 50
# By column name
df$Marks # one column as vector
df[["Marks"]] # same
df[, "Marks"] # same
# Multiple columns
df[, c("Name", "Marks")]
# By row
df[1, ] # first row
df[df$Marks > 70, ] # rows where marks > 70
df[df$Pass == TRUE, "Name"] # passing students' names
# Using subset()
subset(df, Marks > 70, select = c(Name, Marks))
# install.packages("dplyr")
library(dplyr)
df %>%
filter(Marks > 70) %>%
select(Name, Marks) %>%
arrange(desc(Marks))
# Add rows (same columns)
combined <- rbind(df1, df2)
# Add columns (same number of rows)
expanded <- cbind(df1, age = c(19, 20, 21))
students <- data.frame(ID = 1:5, Name = c("A","B","C","D","E"))
scores <- data.frame(ID = c(2,3,4,6), Score = c(80, 75, 90, 60))
# Inner join: only IDs present in both
inner <- merge(students, scores, by = "ID")
# Left join: all from students
left <- merge(students, scores, by = "ID", all.x = TRUE)
# Right join: all from scores
right <- merge(students, scores, by = "ID", all.y = TRUE)
# Full outer join
outer <- merge(students, scores, by = "ID", all = TRUE)
Merge marks of Math and Science exams by Roll Number — left join keeps all students:
math <- data.frame(Roll = 1:10, Math = round(runif(10, 40, 100)))
science <- data.frame(Roll = c(1:5, 7, 9, 10), Science = round(runif(8, 40, 100)))
merged <- merge(math, science, by = "Roll", all.x = TRUE)
Stack two semesters' data with same columns:
sem1 <- data.frame(Name = c("A","B","C"), Grade = c(85, 90, 78))
sem2 <- data.frame(Name = c("D","E"), Grade = c(82, 88))
all <- rbind(sem1, sem2)
| Function | Use |
|---|---|
apply(M, MARGIN, FUN) | Apply FUN over rows (MARGIN = 1) or columns (MARGIN = 2) of a matrix / data frame |
lapply(list, FUN) | Apply over a list; returns a list |
sapply(list, FUN) | Like lapply but simplifies to vector / matrix |
tapply(x, group, FUN) | Apply by group (like SQL GROUP BY) |
mapply(FUN, ...) | Multivariate version of sapply |
M <- matrix(1:20, nrow = 4)
apply(M, 1, mean) # row means: 9 10 11 12
apply(M, 2, sum) # column sums: 10 26 42 58 74
scores <- c(80, 85, 75, 90, 95, 60, 70)
groups <- c("A","A","A","B","B","C","C")
tapply(scores, groups, mean) # A:80 B:92.5 C:65
# Define a function
cv <- function(x) {
sd(x) / mean(x) * 100
}
cv(c(10, 12, 15, 18, 20)) # coefficient of variation
# Anonymous (lambda) function inside apply
apply(M, 2, function(col) max(col) - min(col))
Use apply to get column-wise means of iris numeric columns:
data(iris)
apply(iris[, 1:4], 2, mean)
# Sepal.Length Sepal.Width Petal.Length Petal.Width
# 5.843 3.057 3.758 1.199
Mean Sepal.Length by Species using tapply:
tapply(iris$Sepal.Length, iris$Species, mean)
# setosa versicolor virginica
# 5.006 5.936 6.588
getwd() # current directory
setwd("C:/Users/ragha/Documents/R") # change directory
ls() # list all objects in memory
rm(x) # remove x
rm(list = ls()) # clear workspace
save(df, file = "df.RData") # save object
load("df.RData") # load it back
save.image() # save entire workspace
sessionInfo() # R version & loaded packages
install.packages("psych") # once
library(psych) # each new session
library() # list installed packages
read.csv, read_excel, read.table, dbGetQuery import data; write.csv, write_xlsx export.na.rm = TRUE or na.omit().[ ], $, subset(), or dplyr verbs.merge() implements SQL-style joins; rbind/cbind stack.apply, lapply, sapply, tapply are vectorised alternatives to loops.