Skip to the content

Topics Covered

Origin & History Definitions Importance Scope Limitations Functions Primary Data Secondary Data Classification Presentation Notation (Σ, μ, σ)
On this page
  1. 1. Origin and History of Statistics
  2. 2. Definitions of Statistics
  3. 3. Importance of Statistics
  4. 4. Scope of Statistics
  5. 5. Limitations of Statistics
  6. 6. Functions of Statistics
  7. 7. Collection of Data
  8. 8. Classification of Data
  9. 9. Presentation of Data
  10. 10. Statistical Notation — A Preview
  11. Key Take-aways from Unit 1

1. Origin and History of Statistics

ORIGIN

The word Statistics comes from:

All these words go back to the idea of the political state. In ancient times, kings used statistics to record population, wealth and military strength to run the state — hence the name.

Brief Historical Timeline

PeriodContribution
3050 BC (Egypt)Pharaohs maintained census records to build the pyramids.
Ancient IndiaKautilya's Arthashastra mentions birth/death registration; Akbar's Ain-i-Akbari by Abul Fazl recorded agricultural statistics.
17th CenturyCaptain John Graunt — "Father of Vital Statistics" — analysed Bills of Mortality.
18th–19th CenturyLaplace, Gauss developed the theory of errors and the Normal distribution.
20th CenturyKarl Pearson, R. A. Fisher, P. C. Mahalanobis (India) made statistics a rigorous science.
From counting the state to a mathematical science Descriptive / "state-craft" era Mathematical era 3050 BC Egypt — census for pyramids ~300 BC – 1590 AD India — Kautilya's Arthashastra; Ain-i-Akbari 17th century John Graunt — "Father of Vital Statistics" 18th–19th c. Laplace & Gauss — theory of errors, Normal curve 20th c. Pearson, Fisher, Mahalanobis — rigorous science
The discipline began as simple record-keeping for the state (hence the name) and only became a mathematical science in the last three centuries — the theory you meet in Units 3–5 belongs to that final, modern phase.

2. Definitions of Statistics

The word "Statistics" is used in two senses:

PLURAL DEFINITIONS

Bowley: "Statistics are numerical statements of facts in any department of enquiry placed in relation to each other."

Yule and Kendall: "By statistics we mean quantitative data affected to a marked extent by multiplicity of causes."

Horace Secrist (the most accepted): "Statistics are aggregates of facts, affected to a marked extent by multiplicity of causes, numerically expressed, enumerated or estimated according to reasonable standards of accuracy, collected in a systematic manner for a predetermined purpose and placed in relation to each other."

SINGULAR DEFINITIONS

Croxton and Cowden: "Statistics may be defined as the science of collection, presentation, analysis and interpretation of numerical data."

R. A. Fisher: "The science of statistics is essentially a branch of applied mathematics and may be regarded as the mathematics applied to observational data."

Characteristics implied by Secrist's definition (Important for exam)

  1. Aggregate of facts — single isolated figure is not statistics.
  2. Affected by multiplicity of causes — e.g. crop yield depends on rain, soil, seed.
  3. Numerically expressed — qualitative facts like "honesty" are not statistics.
  4. Enumerated or estimated — counted (population) or estimated (national income).
  5. Reasonable standards of accuracy.
  6. Collected systematically — not haphazardly.
  7. For a predetermined purpose.
  8. Placed in relation to each other — comparable.

3. Importance of Statistics

  1. Planning — Five Year Plans of India use statistical data.
  2. Administration — government decisions on tax, pricing, subsidies.
  3. Economics — supply, demand, GDP, inflation, unemployment.
  4. Business & Industry — production, sales forecasting, quality control.
  5. Banking & Insurance — premium calculation, risk analysis.
  6. Research — testing of hypothesis in social sciences and medicine.
  7. Astronomy, Biology, Physics — measurement and error theory.
  8. Education & Psychology — IQ tests, scaling.

4. Scope of Statistics

Statistics is widely applied across disciplines:

5. Limitations of Statistics

  1. It does not study qualitative phenomena like beauty, honesty, intelligence directly.
  2. It does not deal with individual measurements — only aggregates.
  3. Statistical laws are true on average, not exact.
  4. Statistics can be misused: "Figures don't lie, but liars figure."
  5. Requires expertise — wrong handling gives wrong results.
  6. Results are only as good as the data collected.

6. Functions of Statistics

The four major functions are Collection, Presentation, Analysis, Interpretation.

STEP 1 Collection gather raw facts STEP 2 Presentation tables, diagrams STEP 3 Analysis averages, dispersion… STEP 4 Interpretation valid conclusions If Step 1 is faulty, every later step inherits the error — "garbage in, garbage out."
The four functions form a one-way pipeline: each stage consumes the output of the one before it, which is why accurate collection is the foundation of the whole study.
FUNCTION 1

Collection of Data

Gathering numerical facts in a systematic and accurate manner. The first and most crucial step, because if data are wrong everything that follows will also be wrong.

FUNCTION 2

Presentation of Data

Arranging collected data in tables, diagrams or graphs so that it becomes easy to understand.

FUNCTION 3

Analysis of Data

Using statistical tools — averages, dispersion, correlation, regression — to extract meaningful information from data.

FUNCTION 4

Interpretation of Data

Drawing valid conclusions from analysis. Wrong interpretation defeats the entire purpose of a statistical study.

7. Collection of Data

DATA PRIMARY first-hand · original · costly SECONDARY re-used · cheap · quick Six methods of collecting primary data 1 · Direct personal interview 2 · Indirect oral investigation 3 · Local correspondents 4 · Mailed questionnaires 5 · Schedules via enumerators 6 · Direct observation Choice depends on area covered, cost, accuracy and literacy of respondents. Published RBI, NSSO, CSO, WHO, UN, journals Unpublished firm records, dissertations, files
A single rule decides the label: data collected first-hand for the current enquiry is primary; the very same figures become secondary the moment a different investigator re-uses them (see Example 1 below).

Two Sources of Data

Primary DataSecondary Data
Collected first-hand by the investigator for the present purpose. Already collected by someone else, used by the present investigator.
Original; specifically suited to the problem. Saves time and money but may not exactly fit current needs.
Costly and time consuming. Cheap and quickly available.

Methods of Collecting Primary Data

  1. Direct personal interview — investigator asks respondent face to face.
  2. Indirect oral investigation — information collected from third parties (police enquiries).
  3. Information through local correspondents — newspapers use this method.
  4. Mailed questionnaires — sent through post; cheap, wider area, but low response.
  5. Schedules sent through enumerators — used in census; enumerator fills the schedule.
  6. Direct observation — investigator records by observing without questioning.

Sources of Secondary Data

EXAMPLE 1

Identifying primary vs. secondary data

A college canteen owner wants to know which snack is most preferred. He stands at the counter for one week and records every order — this is primary data (collected first hand).

The same owner reads a published nutrition report from the FSSAI to fix the price — that report is secondary data for him.

EXAMPLE 2

Choosing a method of primary collection

To study reading habits of 50 students in your class, the best method is direct personal interview (small group, high accuracy).

To study reading habits across all colleges in Andhra Pradesh, mailed questionnaires or schedules through enumerators are better — wide geographical area.

8. Classification of Data

Classification means arranging data into groups or classes according to common characteristics. It is the first step of presentation.

Objectives

Types of Classification (as per syllabus)

CLASSIFICATION OF DATA Quantitative basis: a number you can measure height, income, marks (0–20, 20–40…) Qualitative basis: an attribute, not measurable gender, religion, blood group, literacy Temporal basis: time (chronological) 1991, 2001, 2011, 2021 census counts Spatial basis: place (geographical) rice output by state: AP, TS, TN, Punjab
The four bases are not mutually exclusive — one data set can be classified in several ways at once (Examples 1 and 2 combine a temporal or spatial axis with a quantitative count).
QUANTITATIVE

Based on a numerical / measurable characteristic such as height, weight, income, marks.

Example: Students grouped by marks: 0–20, 20–40, 40–60, ...

QUALITATIVE

Based on a non-measurable attribute such as gender, religion, literacy, blood group.

Example: Population grouped as Male/Female, or Literate/Illiterate.

TEMPORAL (Chronological)

Based on time: hourly, daily, monthly, yearly.

Example: Population of India in 1981, 1991, 2001, 2011, 2021.

SPATIAL (Geographical)

Based on location / place: state-wise, country-wise.

Example: Rice production of Andhra Pradesh, Telangana, Tamil Nadu, Punjab.

EXAMPLE 1

Classify the following

"Number of patients admitted in KGH Hospital each day in May 2026."

Answer: The classification is temporal (by date) and at the same time the counts are quantitative. Two-way classification is allowed.

EXAMPLE 2

Classify the following

"Number of literate persons in 13 districts of Andhra Pradesh in 2021."

Answer: The classification is spatial (district-wise) and the attribute "literate" is qualitative; counts of literates are quantitative.

9. Presentation of Data

Two basic modes are prescribed in the syllabus: Textual and Tabular.

9.1 Textual Presentation

Data is described within sentences and paragraphs of running text.

EXAMPLE 1

Textual presentation

"In a class of 60 students, 40 are boys and 20 are girls. Of the 40 boys, 32 passed and 8 failed. Of the 20 girls, 18 passed and 2 failed."

Notice that comparison is hard — that's why we use tables.

9.2 Tabular Presentation

A statistical table is a systematic arrangement of data in rows and columns for easy reading and comparison.

Essential Parts of a Good Statistical Table

  1. Table number — for reference.
  2. Title — short, clear, complete (what, where, when).
  3. Head note / Prefatory note — units of measurement (₹ in lakhs, kg, etc.).
  4. Caption — heading of columns.
  5. Stub — heading of rows.
  6. Body — actual numerical data.
  7. Footnote — clarifications, exceptions.
  8. Source note — where the data came from.
Memory hook: "T-T-H-C-S-B-F-S" — Title, Table No., Head note, Caption, Stub, Body, Footnote, Source.
① Table number ② Title ③ Head note (units) ④ Caption (column heads) ⑤ Stub (row heads) ⑥ Body (the figures) ⑦ Footnote ⑧ Source note Table 1. Result of 60 students, Class A, June 2026 (Head note: figures in number of students) GenderPassed FailedTotal Boys 32840 Girls 18220 Total 501060 Footnote: "Pass" = securing ≥ 35% in the term-end exam. Source: Class Register, Dept. of Statistics.
The eight parts of a well-built table, located on a real example. Caption labels the columns, the stub labels the rows — the two are easy to confuse, so anchor them here.
EXAMPLE 2

Convert the textual data of Example 1 into a table

Table 1: Result of 60 students, Class A, June 2026 [Source: Class Register]

Gender (Stub)PassedFailedTotal
Boys32840
Girls18220
Total501060

Footnote: Pass means securing ≥ 35% in the term-end exam.

Types of Tables

10. Statistical Notation — A Preview

Unit 1 is descriptive, but every summary you compute from Unit 3 onward rests on a small, fixed vocabulary of symbols. Fixing that notation now — while the data are still just tables and classes — makes the later formulae read as shorthand rather than a new language.

OBSERVATIONS & INDEXING

A data set of \(N\) values is written \(x_1, x_2, \ldots, x_N\); a single typical value is \(x_i\), where the index \(i\) runs \(i = 1, 2, \ldots, N\).

SUMMATION (Σ)

The Greek capital sigma is the "add-them-up" operator:

\[ \sum_{i=1}^{N} x_i \;=\; x_1 + x_2 + \cdots + x_N. \]

For the two pass counts \(x_1 = 32\) and \(x_2 = 18\) we would write \(\sum_{i=1}^{2} x_i = 32 + 18 = 50\) — exactly the "Total Passed" cell of the table above. Two rules you will use constantly:

\[ \sum_{i=1}^{N} (a x_i) = a\sum_{i=1}^{N} x_i, \qquad \sum_{i=1}^{N} (x_i + y_i) = \sum_{i=1}^{N} x_i + \sum_{i=1}^{N} y_i. \]

FREQUENCIES & PROPORTIONS

When data are classified (Section 8), each class \(i\) carries a frequency \(f_i\) (how many items fall in it). The total number of observations is

\[ N = \sum_{i=1}^{k} f_i \quad\text{over } k \text{ classes.} \]

The relative frequency (proportion) of class \(i\) is \(p_i = f_i / N\), and its percentage is \(100\,p_i\). A one-line derivation shows the proportions are guaranteed to close to 1:

\[ \sum_{i=1}^{k} p_i \;=\; \sum_{i=1}^{k} \frac{f_i}{N} \;=\; \frac{1}{N}\sum_{i=1}^{k} f_i \;=\; \frac{N}{N} \;=\; 1. \]

This is the arithmetic that makes a pie chart's slices (Unit 2) always fill the full circle.

POPULATION vs SAMPLE

Statistics deals with aggregates, so it constantly distinguishes the whole group from a part of it. The convention — worth memorising now — is Roman letters for a sample, Greek for a population:

QuantityPopulation (parameter)Sample (statistic)
Size\(N\)\(n\)
Mean\(\mu\)\(\bar{x}\)
Variance\(\sigma^{2}\)\(s^{2}\)
Std. deviation\(\sigma\)\(s\)
Proportion\(P\)\(p\)

A parameter is a fixed (usually unknown) property of the whole population; a statistic is computed from a sample and used to estimate it. Descriptive statistics (this course) summarises whatever data are in hand; inferential statistics later uses the sample statistic to reason about the population parameter.

Key Take-aways from Unit 1