Skip to the content

Topics Covered

Key Publications & Outputs Wings of NSO Major Surveys Phases of Census Large Sample Surveys Arithmetic Operators Assignment Vectors Matrix Creation Matrix Operations Mathematical & Statistical Functions Probability Distributions (prefixes d/p/q/r )

Topic Overview — What & Why

Unit X has two distinct halves: knowledge of the institutional Indian statistical machinery and prominent Indian statisticians, and computational skills (R and LaTeX). NET questions on this unit are largely factual or syntax-based — precision matters.

  • MoSPI: Government of India's apex statistical ministry; oversees the entire statistical system. Two wings: National Statistical Office (statistical activities) and Programme Implementation (monitoring projects, MPLADS).
  • National Statistical Commission (NSC): autonomous policy body (2005); sets quality standards, recommends methodology revisions, coordinates with state statistical agencies.
  • National Statistical Office (NSO): formed in 2019 by merging CSO and NSSO; nodal agency for compilation and dissemination of official statistics.
  • Census & large sample surveys: Census every 10 years (Registrar General); NSS rounds, NFHS, SRS, ASI, PLFS — the data infrastructure of national policy making.
  • Indian Statisticians: Mahalanobis ($D^2$, ISI, NSS); Sukhatme (sampling in agriculture); Bose (combinatorics, BIBD, disproved Euler); S N Roy (multivariate); C R Rao (Cramér-Rao, Rao-Blackwell, Rao's score); Basu (theorem on independence).
  • R as a calculator: arithmetic, assignment, vector creation. The starting point for any R user.
  • Functions & matrix operations: matrix() construction, matrix multiplication via %*%, solve(), eigen(), det(), t().
  • Built-in functions, missing data, logical operators: standard math/stat functions, the d/p/q/r distribution prefixes, NA handling with na.rm, vectorised logicals (&, |).
  • Conditional execution & loops: if/else, for, while, repeat, break, next; the apply family for vectorised iteration.
  • Sequences, repeats, sorting, strings: seq, rep, sort, order, rank, plus string manipulation (paste, substr, gsub).
  • Lists, factors, formatting: heterogeneous containers, categorical encodings (ordered/unordered), display via print, cat, format, sprintf.
  • Data frames & I/O: tabular data; read.csv, write.csv, RDS files for serialised objects.
  • Graphics & plots: base R plotting (plot, hist, boxplot, qqnorm) plus device control for saving figures.
  • Scripts & functions: writing reusable code, default arguments, scoping rules, error handling with tryCatch.
  • LaTeX: the standard typesetting system for scientific documents; document classes, sectioning, math mode, tables, figures, bibliography. Other tools include MS Word, LibreOffice, Google Docs, R Markdown, Overleaf.

1. Ministry of Statistics & Programme Implementation (MoSPI)

Why this section? The Indian Statistical System has a clear hierarchy — from policy (NSC) to operations (NSO) to ground-level surveys. Knowing who does what is essential for the factual questions in this unit.

MoSPI is the Government of India ministry responsible for coverage and quality aspects of statistics released in the country and oversees implementation of statutory programmes. Created in 1999 by merging the Department of Statistics and the Department of Programme Implementation.

Two Wings

  1. National Statistical Office (NSO) — handles statistical activities. Includes:
    • Central Statistics Office (CSO).
    • National Sample Survey Office (NSSO).
    • Computer Centre.
  2. Programme Implementation Wing — monitors twenty-point programme, infrastructure projects, and MPLADS (Members of Parliament Local Area Development Scheme).

Key Publications & Outputs

EXAMPLE 1 The base year for CPI (Combined) was revised to 2012=100 by CSO under MoSPI to reflect updated consumption patterns.
EXAMPLE 2 The Annual Survey of Industries — conducted under the Collection of Statistics Act 2008 — covers all factories registered under sections 2(m)(i) and 2(m)(ii) of the Factories Act 1948.

🌍 Where it's used in real life

  1. Publishing India's GDP figures.
  2. Releasing inflation (CPI) data.
  3. Coordinating national surveys.
  4. Monitoring MPLADS and infrastructure projects.
  5. Setting national statistical standards.

2. National Statistical Commission (NSC)

NSC is an autonomous body constituted in 2005 (based on the recommendations of the Rangarajan Commission, 2001) to evolve policies and priorities for the Indian Statistical System.

Composition

Functions

EXAMPLE 1 NSC issues guidelines on revising base year of national accounts (e.g., 2011-12 series for GDP).
EXAMPLE 2 NSC examines methodology of CPI, IIP, GDP and recommends improvements (e.g., expanding IIP coverage).

🌍 Where it's used in real life

  1. Advising on survey methodology.
  2. Setting data-quality standards.
  3. Recommending base-year revisions.
  4. Coordinating centre and state statistics.
  5. Giving policy advice on official data.

3. National Statistical Office (NSO)

NSO formed in 2019 by merging the CSO and NSSO under one umbrella to streamline the statistical system. NSO is now the nodal agency for collection, compilation and dissemination of official statistics.

Wings of NSO

Major Surveys

EXAMPLE 1 NSO publishes PLFS annual report giving worker-population ratio, unemployment rate, labour force participation rate by age, sex, and rural/urban.
EXAMPLE 2 NSO uses stratified multi-stage random sampling: villages/UFS blocks (first stage), households (second stage).

🌍 Where it's used in real life

  1. Running the PLFS employment survey.
  2. Compiling the national accounts.
  3. Conducting consumer-expenditure surveys.
  4. Producing the Index of Industrial Production.
  5. Disseminating official statistics.

4. Census & Large Sample Surveys

Census of India — complete enumeration of population conducted every 10 years under the Census Act 1948 by the Office of the Registrar General & Census Commissioner (Ministry of Home Affairs). Provides demographic, socio-economic, housing data — basis for delimitation, planning, and resource allocation.

Phases of Census

  1. Houselisting and Housing Census — buildings, amenities, assets.
  2. Population Enumeration — individual demographic information.

Large Sample Surveys

EXAMPLE 1 The 2011 Census of India enumerated $\approx 1.21$ billion persons. Sex ratio: 940 females per 1000 males. Literacy rate: 74.04%.
EXAMPLE 2 NSS 75th round (2017-18) covered Health/Education; PLFS replaced the older "Employment-Unemployment" component from 2017-18.

🌍 Where it's used in real life

  1. The decadal population count for planning.
  2. NFHS health indicators.
  3. SRS birth and death rates.
  4. The agriculture census of landholdings.
  5. The economic census of businesses.

5. Contributions of Indian Statisticians

StatisticianMajor Contributions
Prasanta Chandra Mahalanobis (1893-1972) Mahalanobis $D^2$ distance; Founder of Indian Statistical Institute (ISI), 1931; National Sample Survey (1950); Pilot Surveys; "Mahalanobis Model" for second Five Year Plan; large-scale sample surveys.
P. V. Sukhatme (1911-1997) Sampling theory in agricultural statistics; book "Sampling Theory of Surveys"; FAO; food and nutrition statistics; protein requirements estimation.
R. C. Bose (1901-1987) Design of Experiments; combinatorial mathematics — disproved Euler's conjecture on Latin squares (with Shrikhande & Parker); BIBD; finite geometries; coding theory.
S. N. Roy (1906-1964) Multivariate analysis; Roy's Largest Root test; Roy's union-intersection principle; multivariate ANOVA; theory of probability distributions.
C. R. Rao (1920-2023) Cramér-Rao Inequality; Rao-Blackwell Theorem; Rao's Score Test; Fisher-Rao Theorem; orthogonal arrays; quadratic entropy; Padma Vibhushan; International Prize in Statistics 2023.
R. R. Bahadur Bahadur efficiency; large deviations theory; sufficient statistics.
D. Basu Basu's Theorem (independence of complete sufficient and ancillary statistics).
K. R. Parthasarathy Probability theory; quantum probability.
J. K. Ghosh Bayesian inference; high-dimensional asymptotics.
V. S. Huzurbazar Sufficient statistics; admissibility; founded Department of Statistics at Pune.
EXAMPLE 1 Mahalanobis $D^2=(\boldsymbol\mu_1-\boldsymbol\mu_2)^T\Sigma^{-1}(\boldsymbol\mu_1-\boldsymbol\mu_2)$ — used in discriminant analysis, multivariate outlier detection, and anthropometric classification.
EXAMPLE 2 Bose, Shrikhande and Parker (1959) constructed two mutually orthogonal Latin squares of order $n=22$, refuting Euler's 178-year-old conjecture that none exist for $n\equiv 2\pmod 4.$

🌍 Where it's used in real life

  1. Mahalanobis distance in classification and ML.
  2. The Cramér–Rao bound in estimation.
  3. Rao–Blackwell estimator improvement.
  4. Bose's designs in experiments and coding.
  5. Sukhatme's agricultural sampling methods.

6. R as a Calculator

R is a free, open-source software for statistical computing and graphics, developed by Ross Ihaka and Robert Gentleman (1993). Successor to S.

Arithmetic Operators

OperatorMeaningExampleResult
+Addition3 + 58
-Subtraction10 - 46
*Multiplication4 * 624
/Division15 / 43.75
^ or **Power2^101024
%%Modulus17 %% 52
%/%Integer division17 %/% 53

Assignment

x <- 5         # preferred
x = 5          # also valid
assign("x", 5) # functional
5 -> x         # rare
EXAMPLE 1
> (3 + 5) * 2 / 4
[1] 4
> sqrt(16) + log(exp(1))
[1] 5
EXAMPLE 2
> x <- c(2, 4, 6, 8, 10)
> mean(x)
[1] 6
> sum(x^2)
[1] 220

🌍 Where it's used in real life

  1. Quick statistical calculations.
  2. Teaching introductory statistics.
  3. Reproducible analysis scripts.
  4. Ad-hoc data exploration.
  5. Automating repetitive computations.

7. Functions and Matrix Operations

Vectors

v <- c(1, 2, 3, 4, 5)        # combine
seq(1, 10, by = 2)            # 1 3 5 7 9
rep(0, 5)                     # 0 0 0 0 0
1:10                          # 1..10

Matrix Creation

A <- matrix(1:6, nrow = 2, ncol = 3)        # column-major
B <- matrix(1:6, nrow = 2, byrow = TRUE)    # row-major
diag(3)                                      # 3×3 identity

Matrix Operations

OperationR syntax
Element-wise multiplyA * B
Matrix multiplyA %*% B
Transposet(A)
Inversesolve(A)
Determinantdet(A)
Tracesum(diag(A))
Eigenvalues / Eigenvectorseigen(A)
Rankqr(A)$rank
Solve $Ax=b$solve(A, b)
EXAMPLE 1
> A <- matrix(c(2, 1, 1, 3), nrow = 2)
> eigen(A)$values
[1] 3.618034 1.381966
> det(A)
[1] 5
EXAMPLE 2
> A <- matrix(c(1, 2, 3, 4), nrow = 2)   # column-major: rows (1,3) and (2,4)
> b <- c(7, 10)
> solve(A, b)
[1] 1 2

🌍 Where it's used in real life

  1. Solving systems of linear equations.
  2. Regression via matrix algebra.
  3. Eigenvalue and PCA computations.
  4. Portfolio covariance calculations.
  5. Setting up simulations.

8. Built-in Functions, Missing Data & Logical Operators

Mathematical & Statistical Functions

abs(x)        # absolute value
sqrt(x)       # square root
exp(x)        # e^x
log(x)        # natural log
log(x, base)  # arbitrary base
mean(x)       # arithmetic mean
median(x)     # median
var(x)        # variance (n-1 divisor)
sd(x)         # standard deviation
quantile(x)   # quartiles
summary(x)    # 5-num summary + mean
cor(x, y)     # correlation
cov(x, y)     # covariance

Probability Distributions (prefixes d/p/q/r)

dnorm(x, mean, sd)     # density
pnorm(q, mean, sd)     # CDF
qnorm(p, mean, sd)     # quantile
rnorm(n, mean, sd)     # random sample

# Same family for: binom, pois, exp, gamma, beta, t, chisq, f, unif

Missing Data

x <- c(1, NA, 3, NA, 5)
is.na(x)             # FALSE TRUE FALSE TRUE FALSE
mean(x)              # NA
mean(x, na.rm = TRUE) # 3
na.omit(x)           # drop NAs
complete.cases(x)    # logical: not NA

Logical Operators

OperatorMeaning
==Equal
!=Not equal
<, >, <=, >=Comparisons
& / &&Element-wise AND / Short-circuit AND
| / ||Element-wise OR / Short-circuit OR
!NOT
%in%Membership
EXAMPLE 1
> pnorm(1.96)
[1] 0.9750021
> qchisq(0.95, df = 5)
[1] 11.07050
EXAMPLE 2
> x <- c(2, 5, NA, 7, 1)
> x[!is.na(x) & x > 3]
[1] 5 7

🌍 Where it's used in real life

  1. Computing means and standard deviations.
  2. Probability and quantile lookups (pnorm, qnorm).
  3. Cleaning missing values.
  4. Filtering records by a condition.
  5. Random sampling for simulation.

9. Conditional Execution & Loops

If-Else

if (condition) {
  expr1
} else if (condition2) {
  expr2
} else {
  expr3
}

# Vectorised
ifelse(x > 0, "positive", "non-positive")

For Loop

for (i in 1:10) {
  print(i^2)
}

While Loop

x <- 1
while (x < 100) {
  x <- x * 2
}

Repeat / Break / Next

i <- 0
repeat {
  i <- i + 1
  if (i %% 7 == 0) next
  if (i > 30) break
  print(i)
}

Apply Family (preferred over loops)

apply(M, 1, sum)    # row sums
apply(M, 2, mean)   # column means
sapply(1:5, function(x) x^2)
lapply(list(...), FUN)
mapply(FUN, ...)
tapply(x, group, FUN)
EXAMPLE 1
fact <- function(n) {
  if (n <= 1) return(1)
  result <- 1
  for (i in 2:n) result <- result * i
  result
}
fact(5)   # 120
EXAMPLE 2
# Mean of each column of mtcars
sapply(mtcars, mean)

🌍 Where it's used in real life

  1. Automating repeated analyses.
  2. Batch-processing many files.
  3. Running simulation iterations.
  4. Conditional data cleaning.
  5. Building custom pipelines.

10. Data Management: Sequences, Repeats, Sorting, Ordering, Strings

Sequences and Repeats

seq(1, 20, by = 2)
seq(0, 1, length.out = 11)
rep(c(1, 2, 3), times = 3)   # 1 2 3 1 2 3 1 2 3
rep(c(1, 2, 3), each = 3)    # 1 1 1 2 2 2 3 3 3

Sorting and Ordering

x <- c(3, 1, 4, 1, 5, 9, 2, 6)
sort(x)                      # ascending
sort(x, decreasing = TRUE)   # descending
order(x)                     # indices to sort
rank(x)                      # ranks
rev(x)                       # reverse

Strings

nchar("statistics")            # 10
toupper("hello")               # "HELLO"
tolower("WORLD")               # "world"
paste("UGC", "NET", sep = "-") # "UGC-NET"
paste0("Unit", 1:3)            # "Unit1" "Unit2" "Unit3"
substr("statistics", 1, 4)     # "stat"
sub("a", "@", "banana")        # "b@nana"
gsub("a", "@", "banana")       # "b@n@n@"
strsplit("a,b,c", ",")
EXAMPLE 1
> x <- c(20, 5, 30, 10)
> order(x)
[1] 2 4 1 3
> x[order(x)]    # same as sort(x)
[1]  5 10 20 30
EXAMPLE 2
> sprintf("Mean = %.2f", 3.14159)
[1] "Mean = 3.14"

🌍 Where it's used in real life

  1. Generating date or index sequences.
  2. Sorting and ranking results.
  3. Cleaning text fields.
  4. Preparing labels and IDs.
  5. Recoding survey responses.

11. Lists, Factors, Display & Formatting

Lists

L <- list(name = "Asha", age = 25, scores = c(85, 90, 78))
L$name
L[["age"]]
L[[3]][2]   # 90

Factors

g <- factor(c("M", "F", "M", "F", "M"))
levels(g)            # "F" "M"
as.numeric(g)        # 2 1 2 1 2
table(g)             # frequency table

# Ordered factors
edu <- factor(c("BA", "MA", "PhD"),
              levels = c("BA", "MA", "PhD"),
              ordered = TRUE)

Display & Formatting

print(x)
cat("Result:", x, "\n")
format(3.14159, nsmall = 2)        # "3.14"
formatC(1234567, big.mark = ",")   # "1,234,567"
round(3.14159, 2)                  # 3.14
signif(1234567, 3)                 # 1230000
EXAMPLE 1
student <- list(roll = 101, marks = c(80, 75, 90),
                grade = factor("A", levels = c("F","C","B","A")))
mean(student$marks)   # 81.67
EXAMPLE 2
income <- factor(c("Low","High","Med","Low","High"),
                 levels = c("Low","Med","High"),
                 ordered = TRUE)
income[1] < income[2]   # TRUE

🌍 Where it's used in real life

  1. Storing mixed-type results together.
  2. Encoding categorical survey data.
  3. Ordered ratings (low/medium/high).
  4. Formatting numbers for reports.
  5. Building frequency tables.

12. Data Frames, Data Input and Output

Creating Data Frames

df <- data.frame(
  name = c("A", "B", "C"),
  age = c(25, 30, 22),
  score = c(85, 92, 78)
)
df$age            # column
df[1, ]           # row
df[df$age > 24, ] # subset
str(df)           # structure
summary(df)       # summary
nrow(df); ncol(df); dim(df)

Importing Data

read.csv("file.csv", header = TRUE)
read.table("file.txt", sep = "\t", header = TRUE)
readxl::read_excel("file.xlsx")        # excel
foreign::read.spss("file.sav")          # SPSS

Exporting

write.csv(df, "output.csv", row.names = FALSE)
write.table(df, "output.txt", sep = "\t")
saveRDS(df, "df.rds")
df <- readRDS("df.rds")
EXAMPLE 1
data(iris)
head(iris, 3)
mean(iris$Sepal.Length[iris$Species == "setosa"])
EXAMPLE 2
# Add a new column
df$pass <- ifelse(df$score >= 80, "Yes", "No")

🌍 Where it's used in real life

  1. Loading CSV survey data.
  2. Filtering and subsetting records.
  3. Merging two datasets.
  4. Exporting results to file.
  5. Summarising tabular data.

13. Graphics & Plots

Base R Plots

FunctionPlot type
plot(x, y)Scatter plot
hist(x)Histogram
boxplot(x ~ g)Boxplot
barplot(table(x))Bar chart
pie(x)Pie chart
qqnorm(x); qqline(x)Q-Q plot
pairs(df)Scatterplot matrix

Common Arguments

plot(x, y,
     main = "Title",
     xlab = "X-axis", ylab = "Y-axis",
     col  = "blue",
     pch  = 19,           # point character
     type = "b",          # both points and lines
     lty  = 2,            # dashed
     lwd  = 2)            # line width
abline(h = 0, v = 0)
abline(lm(y ~ x), col = "red")
legend("topright", legend = c("Group A","Group B"), col = c(1,2), pch = 19)

Multiple Plots

par(mfrow = c(2, 2))   # 2 x 2 grid
plot(...)              # 4 plots

Saving Plots

png("plot.png", width = 800, height = 600)
plot(...)
dev.off()
EXAMPLE 1
x <- rnorm(100)
hist(x, breaks = 20, col = "skyblue",
     main = "Histogram of N(0,1) sample")
abline(v = 0, col = "red", lwd = 2)
EXAMPLE 2
boxplot(Sepal.Length ~ Species, data = iris,
        col = c("red","green","blue"),
        main = "Sepal Length by Species")

🌍 Where it's used in real life

  1. Histograms of exam scores.
  2. Scatter plots for regression.
  3. Boxplots comparing groups.
  4. Time-series line charts.
  5. Saving figures for reports.

14. Scripts, Functions, and Programming Basics

Defining Functions

square <- function(x) {
  return(x^2)
}

# Multiple arguments with defaults
mysum <- function(x, y = 10) {
  x + y
}

# Variable arguments
myfun <- function(...) {
  args <- list(...)
  sum(unlist(args))
}

Scope

Error Handling

result <- tryCatch({
  log(-1)
}, warning = function(w) "Warning caught",
   error   = function(e) "Error caught")

Scripts

Save a sequence of R commands in a .R file. Run with source("script.R") from console.

Useful Packages

EXAMPLE 1
se <- function(x) sd(x) / sqrt(length(x))
x <- c(2, 4, 6, 8, 10)
se(x)         # 1.414214
EXAMPLE 2
# t-test wrapper
mytest <- function(x, mu0 = 0, alpha = 0.05) {
  n  <- length(x)
  t  <- (mean(x) - mu0) / (sd(x) / sqrt(n))
  p  <- 2 * pt(-abs(t), df = n - 1)
  list(t = t, p = p, reject = (p < alpha))
}
mytest(rnorm(30, 0.5))

🌍 Where it's used in real life

  1. Reusable analysis functions.
  2. Automating monthly reports.
  3. Building custom statistical tools.
  4. Error-handled data pipelines.
  5. Packaging code for a team.

15. LaTeX & Other Word Processing Software

LaTeX is a typesetting system based on the TeX engine (Donald Knuth, 1978; LaTeX by Leslie Lamport, 1985). Standard for mathematical and scientific documents — produces high-quality PDFs from plain-text source files (.tex).

Document Structure

\documentclass[12pt]{article}
\usepackage{amsmath, amssymb, graphicx}
\title{My Statistics Paper}
\author{Author Name}
\date{\today}
\begin{document}
\maketitle
\section{Introduction}
Content goes here...
\end{document}

Math Mode

ElementLaTeXRenders
Inline math$x^2 + y^2 = z^2$$x^2+y^2=z^2$
Display math\[ \int_0^1 x\,dx \]$\int_0^1 x\,dx$
Fraction\frac{a}{b}$\frac{a}{b}$
Square root\sqrt{x}$\sqrt{x}$
Sum\sum_{i=1}^n x_i$\sum_{i=1}^n x_i$
Integral\int_a^b f(x)dx$\int_a^b f(x)dx$
Greek\alpha,\beta,\sigma,\Sigma$\alpha,\beta,\sigma,\Sigma$
Matrix\begin{pmatrix}1&2\\3&4\end{pmatrix}$\begin{pmatrix}1&2\\3&4\end{pmatrix}$

Common Sections

\section{Title}
\subsection{Title}
\subsubsection{Title}
\paragraph{Title}

Lists, Tables, Figures

\begin{itemize}
  \item Bullet 1
  \item Bullet 2
\end{itemize}

\begin{enumerate}
  \item Numbered 1
  \item Numbered 2
\end{enumerate}

\begin{tabular}{|c|c|}
\hline
A & B \\
\hline
1 & 2 \\
\hline
\end{tabular}

\begin{figure}[h]
  \centering
  \includegraphics[width=0.8\linewidth]{plot.png}
  \caption{My Figure}
  \label{fig:plot}
\end{figure}

Bibliography (BibTeX)

\bibliographystyle{plain}
\bibliography{refs}        % refs.bib file
\cite{author2020}

Other Word Processing Software

EXAMPLE 1
The likelihood is
\[
L(\theta;x) = \prod_{i=1}^n f(x_i;\theta).
\]
The MLE satisfies
\[
\frac{\partial}{\partial\theta}\log L(\theta;x) = 0.
\]
Renders mathematical exposition with proper spacing.
EXAMPLE 2
R Markdown chunk:
```{r}
x <- rnorm(100)
mean(x); sd(x)
```
Knit produces report with code, output and plots — preferred for reproducible research.

🌍 Where it's used in real life

  1. Typesetting research papers and theses.
  2. Writing math-heavy notes.
  3. Preparing journal submissions.
  4. Reproducible reports with R Markdown.
  5. Professional CVs and slides.

Final Revision Checklist

Tip: Unit X questions are largely factual (names, years, contributions) and direct R syntax / output prediction — quick wins if memorized.