The word Statistics comes from:
All these words go back to the idea of the political state. In ancient times, kings used statistics to record population, wealth and military strength to run the state — hence the name.
| Period | Contribution |
|---|---|
| 3050 BC (Egypt) | Pharaohs maintained census records to build the pyramids. |
| Ancient India | Kautilya's Arthashastra mentions birth/death registration; Akbar's Ain-i-Akbari by Abul Fazl recorded agricultural statistics. |
| 17th Century | Captain John Graunt — "Father of Vital Statistics" — analysed Bills of Mortality. |
| 18th–19th Century | Laplace, Gauss developed the theory of errors and the Normal distribution. |
| 20th Century | Karl Pearson, R. A. Fisher, P. C. Mahalanobis (India) made statistics a rigorous science. |
The word "Statistics" is used in two senses:
Bowley: "Statistics are numerical statements of facts in any department of enquiry placed in relation to each other."
Yule and Kendall: "By statistics we mean quantitative data affected to a marked extent by multiplicity of causes."
Horace Secrist (the most accepted): "Statistics are aggregates of facts, affected to a marked extent by multiplicity of causes, numerically expressed, enumerated or estimated according to reasonable standards of accuracy, collected in a systematic manner for a predetermined purpose and placed in relation to each other."
Croxton and Cowden: "Statistics may be defined as the science of collection, presentation, analysis and interpretation of numerical data."
R. A. Fisher: "The science of statistics is essentially a branch of applied mathematics and may be regarded as the mathematics applied to observational data."
Statistics is widely applied across disciplines:
The four major functions are Collection, Presentation, Analysis, Interpretation.
Gathering numerical facts in a systematic and accurate manner. The first and most crucial step, because if data are wrong everything that follows will also be wrong.
Arranging collected data in tables, diagrams or graphs so that it becomes easy to understand.
Using statistical tools — averages, dispersion, correlation, regression — to extract meaningful information from data.
Drawing valid conclusions from analysis. Wrong interpretation defeats the entire purpose of a statistical study.
| Primary Data | Secondary Data |
|---|---|
| Collected first-hand by the investigator for the present purpose. | Already collected by someone else, used by the present investigator. |
| Original; specifically suited to the problem. | Saves time and money but may not exactly fit current needs. |
| Costly and time consuming. | Cheap and quickly available. |
A college canteen owner wants to know which snack is most preferred. He stands at the counter for one week and records every order — this is primary data (collected first hand).
The same owner reads a published nutrition report from the FSSAI to fix the price — that report is secondary data for him.
To study reading habits of 50 students in your class, the best method is direct personal interview (small group, high accuracy).
To study reading habits across all colleges in Andhra Pradesh, mailed questionnaires or schedules through enumerators are better — wide geographical area.
Classification means arranging data into groups or classes according to common characteristics. It is the first step of presentation.
Based on a numerical / measurable characteristic such as height, weight, income, marks.
Example: Students grouped by marks: 0–20, 20–40, 40–60, ...
Based on a non-measurable attribute such as gender, religion, literacy, blood group.
Example: Population grouped as Male/Female, or Literate/Illiterate.
Based on time: hourly, daily, monthly, yearly.
Example: Population of India in 1981, 1991, 2001, 2011, 2021.
Based on location / place: state-wise, country-wise.
Example: Rice production of Andhra Pradesh, Telangana, Tamil Nadu, Punjab.
"Number of patients admitted in KGH Hospital each day in May 2026."
Answer: The classification is temporal (by date) and at the same time the counts are quantitative. Two-way classification is allowed.
"Number of literate persons in 13 districts of Andhra Pradesh in 2021."
Answer: The classification is spatial (district-wise) and the attribute "literate" is qualitative; counts of literates are quantitative.
Two basic modes are prescribed in the syllabus: Textual and Tabular.
Data is described within sentences and paragraphs of running text.
"In a class of 60 students, 40 are boys and 20 are girls. Of the 40 boys, 32 passed and 8 failed. Of the 20 girls, 18 passed and 2 failed."
Notice that comparison is hard — that's why we use tables.
A statistical table is a systematic arrangement of data in rows and columns for easy reading and comparison.
Table 1: Result of 60 students, Class A, June 2026 [Source: Class Register]
| Gender (Stub) | Passed | Failed | Total |
|---|---|---|---|
| Boys | 32 | 8 | 40 |
| Girls | 18 | 2 | 20 |
| Total | 50 | 10 | 60 |
Footnote: Pass means securing ≥ 35% in the term-end exam.
Unit 1 is descriptive, but every summary you compute from Unit 3 onward rests on a small, fixed vocabulary of symbols. Fixing that notation now — while the data are still just tables and classes — makes the later formulae read as shorthand rather than a new language.
A data set of \(N\) values is written \(x_1, x_2, \ldots, x_N\); a single typical value is \(x_i\), where the index \(i\) runs \(i = 1, 2, \ldots, N\).
The Greek capital sigma is the "add-them-up" operator:
\[ \sum_{i=1}^{N} x_i \;=\; x_1 + x_2 + \cdots + x_N. \]
For the two pass counts \(x_1 = 32\) and \(x_2 = 18\) we would write \(\sum_{i=1}^{2} x_i = 32 + 18 = 50\) — exactly the "Total Passed" cell of the table above. Two rules you will use constantly:
\[ \sum_{i=1}^{N} (a x_i) = a\sum_{i=1}^{N} x_i, \qquad \sum_{i=1}^{N} (x_i + y_i) = \sum_{i=1}^{N} x_i + \sum_{i=1}^{N} y_i. \]
When data are classified (Section 8), each class \(i\) carries a frequency \(f_i\) (how many items fall in it). The total number of observations is
\[ N = \sum_{i=1}^{k} f_i \quad\text{over } k \text{ classes.} \]
The relative frequency (proportion) of class \(i\) is \(p_i = f_i / N\), and its percentage is \(100\,p_i\). A one-line derivation shows the proportions are guaranteed to close to 1:
\[ \sum_{i=1}^{k} p_i \;=\; \sum_{i=1}^{k} \frac{f_i}{N} \;=\; \frac{1}{N}\sum_{i=1}^{k} f_i \;=\; \frac{N}{N} \;=\; 1. \]
This is the arithmetic that makes a pie chart's slices (Unit 2) always fill the full circle.
Statistics deals with aggregates, so it constantly distinguishes the whole group from a part of it. The convention — worth memorising now — is Roman letters for a sample, Greek for a population:
| Quantity | Population (parameter) | Sample (statistic) |
|---|---|---|
| Size | \(N\) | \(n\) |
| Mean | \(\mu\) | \(\bar{x}\) |
| Variance | \(\sigma^{2}\) | \(s^{2}\) |
| Std. deviation | \(\sigma\) | \(s\) |
| Proportion | \(P\) | \(p\) |
A parameter is a fixed (usually unknown) property of the whole population; a statistic is computed from a sample and used to estimate it. Descriptive statistics (this course) summarises whatever data are in hand; inferential statistics later uses the sample statistic to reason about the population parameter.