Skip to the content

Topics Covered

R Installation Command Prompt Vectors Matrices Data Frames Data Import / Export Missing Values Subsetting Merging apply Family
On this page
  1. 1. The R Environment
  2. 2. R Language Basics
  3. 3. Basic Data Types
  4. 4. Data Import & Export
  5. 5. Handling Missing Values (NA)
  6. 6. Subsetting
  7. 7. Merging Datasets
  8. 8. Applying Functions
  9. 9. Working Directory & Workspace
  10. Key Take-aways

1. The R Environment

DEFINITION

R is a free, open-source language and environment for statistical computing and graphics, first released in 1993. It is maintained by the R Core Team and the wider community via CRAN (Comprehensive R Archive Network).

Installation

  1. Visit https://cran.r-project.org; download the installer for your operating system.
  2. Run the installer. R is available as a console application.
  3. Optional but highly recommended: install RStudio from https://posit.co/download/rstudio-desktop/ — an integrated development environment (IDE) for R.

The R Command Prompt

When R starts, you see a prompt >. Type any expression; press Enter to evaluate.

> 2 + 3
[1] 5

> sqrt(64)
[1] 8

> "Hello, R!"
[1] "Hello, R!"

To assign: x <- 10 or x = 10. To list objects: ls(). To remove: rm(x).

Getting Help

2. R Language Basics

Before working with the data structures in the next section, it helps to know how R stores values, the operators it provides, and its special values.

2.1 Assignment and Variable Names

Assignment stores a value in a name. R offers leftward <- (preferred), the equals sign =, and rightward ->.

x <- 10        # leftward (preferred)
y = 5          # equals sign (valid, less common)
15 -> z        # rightward
z              # 15

Naming rules: a name must start with a letter or a period ., followed by letters, digits, . or _; it cannot start with a digit; names are case-sensitive (Var and var differ); and reserved words (§2.5) cannot be used as names.

2.2 Operators

TypeOperators
Arithmetic+ - * / ^, %% (modulus), %/% (integer division)
Relational< > <= >= == != (return TRUE/FALSE)
Logical& | (element-wise), && || (single value), ! (not)
Assignment<- = ->
Miscellaneous: (sequence), %in% (membership)
Matrix%*% (matrix product), t() (transpose)
17 %% 5              # 2   (remainder)
17 %/% 5             # 3   (integer quotient)
2 ^ 10              # 1024
3 %in% c(1, 2, 3)   # TRUE
(5 > 3) & (2 > 4)    # FALSE
EXAMPLE 1

Leap-year test with logical operators: (2024 %% 4 == 0) & (2024 %% 100 != 0 | 2024 %% 400 == 0) returns TRUE.

2.3 Operator Precedence

In a compound expression R evaluates operators in this order (highest first):

  1. Parentheses ()
  2. Exponentiation ^
  3. Unary sign + -
  4. Multiply / divide * / %/% %%
  5. Add / subtract + -
  6. Relational < > <= >= == !=
  7. Logical ! & && | ||
EXAMPLE 2 (precedence)

3 + 5 * 2 ^ 2 gives \(3 + 5\times(2^2) = 3 + 20 = 23\) — exponent, then multiply, then add. Parentheses override: (3 + 5) * 2 ^ 2 = 32.

2.4 Modes

Every object has a mode — the family of what it stores. Atomic modes hold elements of a single basic type: numeric, character, logical, complex, raw. Non-atomic modes hold richer objects: list (mixed contents), function, expression, call. Use mode() to check; the individual data types are detailed in the next section.

mode(5)             # "numeric"
mode("stats")       # "character"
mode(list(1, "a"))  # "list"

2.5 Special Values, Reserved Words and Characters

ValueMeaningTest
Inf, -Infinfinity, e.g. 1/0is.infinite()
NaNNot a Number, e.g. 0/0is.nan()
NAmissing value (detailed later in the Handling Missing Values section)is.na()
NULLthe empty object (absence of a value)is.null()

Reserved words cannot be used as names: if else for while repeat break next function return TRUE FALSE NA NULL Inf NaN.

Special characters: $ (element of a list / column of a data frame, df$age), @ (slot of an S4 object), [ ] (extract elements), :: (a function from a package, stats::sd), and ... (a variable number of function arguments).

2.6 Logical Values

TRUE/FALSE (shorthands T/F) behave as 1/0 in arithmetic, so they drive filtering: a logical vector inside [ ] keeps the TRUE positions.

vec <- c(10, 20, 30, 40)
vec[vec > 20]    # 30 40   (logical indexing)
all(vec > 5)     # TRUE  – every element passes
any(vec > 35)    # TRUE  – at least one passes
sum(c(TRUE, TRUE, FALSE))  # 2

3. Basic Data Types

3.1 Atomic Vectors

The fundamental data structure. All elements must be of the same type.

TypeExampleCheck function
numeric (double)3.14is.numeric()
integer5Lis.integer()
character"R"is.character()
logicalTRUE / FALSEis.logical()
complex1+2iis.complex()

Creating Vectors

x <- c(2, 4, 6, 8, 10)        # numeric vector
y <- 1:10                     # sequence 1..10
z <- seq(0, 1, by = 0.25)     # 0.00 0.25 0.50 0.75 1.00
w <- rep("A", 5)              # "A" "A" "A" "A" "A"
b <- c(TRUE, FALSE, TRUE)     # logical
length(x)                     # 5
class(x)                      # "numeric"

Vector Operations (vectorised)

x <- c(1, 2, 3, 4)
x * 2          # 2 4 6 8
x + c(10, 20)  # recycles: 11 22 13 24
sum(x); mean(x); sd(x)
EXAMPLE 1

Create a vector of 10 student marks and compute the mean & max:

marks <- c(45, 52, 38, 71, 64, 49, 58, 33, 67, 72)
mean(marks)    # 54.9
max(marks)     # 72
EXAMPLE 2

Convert temperatures from Celsius to Fahrenheit:

celsius <- c(0, 10, 20, 30, 40)
fahrenheit <- celsius * 9/5 + 32
fahrenheit     # 32 50 68 86 104

3.2 Matrices

Two-dimensional, all elements same type.

M <- matrix(1:12, nrow = 3, ncol = 4)
M
#      [,1] [,2] [,3] [,4]
# [1,]    1    4    7   10
# [2,]    2    5    8   11
# [3,]    3    6    9   12

dim(M)        # 3 4
M[2, 3]       # 8
M[ , 2]       # second column → 4 5 6
M[1, ]        # first row → 1 4 7 10
t(M)          # transpose
M %*% t(M)    # matrix multiplication

3.3 Lists

Can contain elements of different types.

person <- list(name = "Aarav", age = 19, marks = c(85, 90, 78))
person$name           # "Aarav"
person$marks[2]       # 90

3.4 Data Frames

Workhorse for statistical data — like a spreadsheet table; columns can be of different types.

df <- data.frame(
  ID    = 1:5,
  Name  = c("Aarav","Bhavna","Chetan","Divya","Esha"),
  Marks = c(78, 82, 65, 90, 88),
  Pass  = c(TRUE, TRUE, TRUE, TRUE, TRUE)
)
df
str(df)        # structure
head(df, 3)    # first 3 rows
tail(df, 2)    # last 2 rows
nrow(df); ncol(df); dim(df)

3.5 Factors (Categorical Variables)

gender <- factor(c("M", "F", "F", "M", "M"))
levels(gender)        # "F" "M"
table(gender)         # F:2  M:3

# Ordered factor
grade <- factor(c("B","A","C","A","B"), levels = c("C","B","A"), ordered = TRUE)

4. Data Import & Export

4.1 CSV Files

# Reading
data <- read.csv("students.csv")             # comma-separated
data <- read.csv("students.csv", header = TRUE, na.strings = "")

# Writing
write.csv(data, "output.csv", row.names = FALSE)

4.2 Excel Files

# install.packages("readxl")
library(readxl)
data <- read_excel("students.xlsx", sheet = 1)
data <- read_excel("students.xlsx", sheet = "Marks", range = "A1:D50")

# Writing (use writexl or openxlsx)
# install.packages("writexl")
library(writexl)
write_xlsx(data, "output.xlsx")

4.3 Tab / Whitespace-Delimited Files

data <- read.table("data.txt", header = TRUE, sep = "\t")
data <- read.delim("data.txt")              # tab default

4.4 Databases

# install.packages("DBI"); install.packages("RSQLite")
library(DBI); library(RSQLite)
con <- dbConnect(SQLite(), "school.db")
data <- dbGetQuery(con, "SELECT * FROM students WHERE marks > 60")
dbDisconnect(con)

4.5 Built-in Datasets (for practice)

data()            # list all built-in datasets
data(iris)
head(iris)        # 4 numeric features + Species factor
data(mtcars)      # car attributes
data(airquality)  # NY air-quality measurements

5. Handling Missing Values (NA)

R represents missing data as NA.

x <- c(2, 4, NA, 8, NA, 12)
is.na(x)                 # FALSE FALSE TRUE FALSE TRUE FALSE
sum(is.na(x))            # 2 missing values

mean(x)                  # NA  ← because NA propagates
mean(x, na.rm = TRUE)    # 6.5 (NAs removed)

# Drop rows with any NA in a data frame
complete <- na.omit(df)

# Replace NA with the column mean
x[is.na(x)] <- mean(x, na.rm = TRUE)
EXAMPLE 1

Identify and replace missing values in airquality$Ozone:

data(airquality)
sum(is.na(airquality$Ozone))            # 37
airquality$Ozone[is.na(airquality$Ozone)] <- mean(airquality$Ozone, na.rm = TRUE)
sum(is.na(airquality$Ozone))            # 0
EXAMPLE 2

Drop rows with any missing values in a custom data frame:

df_complete <- na.omit(df)
nrow(df) - nrow(df_complete)   # number of rows dropped

6. Subsetting

6.1 Vectors

x <- c(10, 20, 30, 40, 50)
x[1]                # 10
x[c(2, 4)]          # 20 40
x[-1]               # everything except element 1
x[x > 25]           # logical filter → 30 40 50

6.2 Data Frames

# By column name
df$Marks                       # one column as vector
df[["Marks"]]                  # same
df[, "Marks"]                  # same

# Multiple columns
df[, c("Name", "Marks")]

# By row
df[1, ]                        # first row
df[df$Marks > 70, ]            # rows where marks > 70
df[df$Pass == TRUE, "Name"]    # passing students' names

# Using subset()
subset(df, Marks > 70, select = c(Name, Marks))

6.3 dplyr Style (Modern)

# install.packages("dplyr")
library(dplyr)
df %>%
  filter(Marks > 70) %>%
  select(Name, Marks) %>%
  arrange(desc(Marks))

7. Merging Datasets

7.1 rbind / cbind (stacking)

# Add rows (same columns)
combined <- rbind(df1, df2)

# Add columns (same number of rows)
expanded <- cbind(df1, age = c(19, 20, 21))

7.2 merge() — Like SQL JOIN

students <- data.frame(ID = 1:5, Name = c("A","B","C","D","E"))
scores   <- data.frame(ID = c(2,3,4,6), Score = c(80, 75, 90, 60))

# Inner join: only IDs present in both
inner    <- merge(students, scores, by = "ID")

# Left join: all from students
left     <- merge(students, scores, by = "ID", all.x = TRUE)

# Right join: all from scores
right    <- merge(students, scores, by = "ID", all.y = TRUE)

# Full outer join
outer    <- merge(students, scores, by = "ID", all = TRUE)
EXAMPLE 1

Merge marks of Math and Science exams by Roll Number — left join keeps all students:

math    <- data.frame(Roll = 1:10, Math    = round(runif(10, 40, 100)))
science <- data.frame(Roll = c(1:5, 7, 9, 10), Science = round(runif(8, 40, 100)))
merged  <- merge(math, science, by = "Roll", all.x = TRUE)
EXAMPLE 2

Stack two semesters' data with same columns:

sem1 <- data.frame(Name = c("A","B","C"), Grade = c(85, 90, 78))
sem2 <- data.frame(Name = c("D","E"),     Grade = c(82, 88))
all  <- rbind(sem1, sem2)

8. Applying Functions

8.1 The apply family

FunctionUse
apply(M, MARGIN, FUN)Apply FUN over rows (MARGIN = 1) or columns (MARGIN = 2) of a matrix / data frame
lapply(list, FUN)Apply over a list; returns a list
sapply(list, FUN)Like lapply but simplifies to vector / matrix
tapply(x, group, FUN)Apply by group (like SQL GROUP BY)
mapply(FUN, ...)Multivariate version of sapply

Examples

M <- matrix(1:20, nrow = 4)
apply(M, 1, mean)        # row means: 9 10 11 12
apply(M, 2, sum)         # column sums: 10 26 42 58 74

scores <- c(80, 85, 75, 90, 95, 60, 70)
groups <- c("A","A","A","B","B","C","C")
tapply(scores, groups, mean)   # A:80   B:92.5   C:65

8.2 Custom Functions

# Define a function
cv <- function(x) {
  sd(x) / mean(x) * 100
}

cv(c(10, 12, 15, 18, 20))    # coefficient of variation

# Anonymous (lambda) function inside apply
apply(M, 2, function(col) max(col) - min(col))
EXAMPLE 1

Use apply to get column-wise means of iris numeric columns:

data(iris)
apply(iris[, 1:4], 2, mean)
# Sepal.Length  Sepal.Width  Petal.Length  Petal.Width
#       5.843       3.057          3.758        1.199
EXAMPLE 2

Mean Sepal.Length by Species using tapply:

tapply(iris$Sepal.Length, iris$Species, mean)
#    setosa  versicolor   virginica
#     5.006       5.936       6.588

9. Working Directory & Workspace

getwd()                              # current directory
setwd("C:/Users/ragha/Documents/R")  # change directory

ls()                                 # list all objects in memory
rm(x)                                # remove x
rm(list = ls())                      # clear workspace

save(df, file = "df.RData")          # save object
load("df.RData")                     # load it back

save.image()                         # save entire workspace
sessionInfo()                        # R version & loaded packages

Installing & Loading Packages

install.packages("psych")         # once
library(psych)                    # each new session
library()                         # list installed packages

Key Take-aways