A survey is a method of collecting information from a subset (sample) of a population in order to draw inferences about the entire population. The fundamental logic is that a properly selected sample can represent the population, allowing us to estimate population parameters with quantifiable precision.
Statistical inference from surveys involves two main activities: estimation (producing point or interval estimates of population parameters such as means, totals, or proportions) and hypothesis testing (deciding whether observed patterns are statistically significant or attributable to chance). Both activities rest on probability sampling designs that give every element a known, non-zero chance of selection.
However, survey inference is never perfect. Deviations of survey estimates from the true population values are called survey errors. These are broadly classified into two categories:
\[ \text{TSE} = \text{Sampling Error} + \text{Non-sampling Error} \]
where Sampling Error arises because only a subset of the population is observed, and Non-sampling Error encompasses all other sources of deviation (coverage errors, measurement errors, non-response, processing errors, etc.).
Sampling error is the difference between the estimate derived from a sample and the true population value that would be obtained from a complete census. It is inherent in any sampling process and can be quantified probabilistically.
For a sample mean \(\bar{x}\) estimating population mean \(\mu\):
\[ \text{Sampling Error} = \bar{x} - \mu \]
The standard error of the sample mean under simple random sampling is:
\[ SE(\bar{x}) = \frac{s}{\sqrt{n}} \sqrt{1 - \frac{n}{N}} \]
where \(s\) is the sample standard deviation, \(n\) is the sample size, \(N\) is the population size, and \(\sqrt{1 - n/N}\) is the finite population correction (fpc) factor.
Non-sampling errors are far more insidious because they can occur in both sample surveys and complete censuses. Unlike sampling error, they do not diminish with increasing sample size and cannot be easily quantified. Major types include:
A labour bureau uses simple random sampling to estimate the average monthly wage of factory workers in a district. The population has \(N = 10{,}000\) workers; a sample of \(n = 400\) is drawn. The sample yields \(\bar{x} = \text{₹}18{,}500\) with \(s = \text{₹}3{,}200\). The standard error is:
\[ SE(\bar{x}) = \frac{3200}{\sqrt{400}} \sqrt{1 - \frac{400}{10000}} = \frac{3200}{20} \times \sqrt{0.96} = 160 \times 0.9798 = \text{₹}156.77 \]
The 95% confidence interval is \(\bar{x} \pm 1.96 \times SE = 18500 \pm 307.27\), i.e., \([\text{₹}18{,}192.73,\; \text{₹}18{,}807.27]\). The sampling error here is at most about ₹307 with 95% confidence.
A national health survey uses a household listing from the latest census as its sampling frame. However, rapid urbanisation in the past three years means several new slum settlements are not in the frame. Additionally, 28% of selected households refuse to participate. The survey estimate of the prevalence of diabetes is 7.2%. However, because (a) the slum dwellers (who have higher diabetes prevalence) are excluded from the frame, and (b) non-respondents are predominantly older adults (also at higher risk), the true prevalence is likely closer to 9.5%. The non-sampling error (coverage error + non-response error) here is approximately 2.3 percentage points — larger than the reported sampling error of 0.8 percentage points. This illustrates why non-sampling errors can be far more damaging than sampling errors.
The target population is the complete collection of all elements that the researcher wishes to study and about which inferences are to be drawn. It is defined by the research objectives and must be specified before any sampling or data collection begins.
Defining the target population requires specifying three dimensions:
A poorly defined target population leads to ambiguous results that cannot be generalised meaningfully. It is crucial to distinguish the target population from the survey population (or frame population), which is the set of units actually available for sampling through the sampling frame.
The target population is the group you want to study. The survey population (frame population) is the group you can study given the available sampling frame. The difference between them is the source of coverage error.
A researcher wants to study the impact of online learning on academic performance of undergraduate students in one state. The target population is defined as: "All students enrolled in 3-year or 4-year UG programmes in universities and affiliated colleges of that state during the academic year 2024–25." This definition clearly specifies the content (UG students), geography (the state), and time (2024–25). Note that students in other states, postgraduate students, and students from previous academic years are excluded.
In a consumer expenditure survey, the target population is "all households in Mumbai city." The sampling frame is the electoral roll. However, the electoral roll excludes: (a) households that moved to Mumbai after the rolls were updated, (b) households in newly developed areas not yet included in electoral wards, and (c) non-citizens. It also includes households that have since moved out or been dissolved. The survey population (electoral roll households) does not perfectly match the target population, creating an under-coverage of new migrants and an over-coverage of departed households. The researcher must acknowledge this gap and, if possible, use supplementary area frames to reduce the discrepancy.
A sampling frame is a list or mechanism that identifies the units of the survey population from which the sample is to be drawn. It serves as the bridge between the target population and the actual sample. An ideal frame lists every element of the target population once and only once, with no extraneous units.
Common types of sampling frames include:
Coverage error occurs when there is a discrepancy between the target population and the sampling frame. It has three components:
If \(\mu_F\) is the mean of the frame population and \(\mu_T\) is the mean of the target population, then:
\[ \text{Coverage Bias} = \mu_F - \mu_T \]
This bias depends on (a) the proportion of the target population omitted from the frame, and (b) how much the omitted units differ from those included.
A political polling agency uses a random digit dialling (RDD) frame of landline telephone numbers to survey voting intentions. In 2010, about 25% of households had only mobile phones and no landline. These mobile-only households tend to be younger, more urban, and more likely to support progressive candidates. The landline-only frame causes significant under-coverage, and the poll overestimates conservative support. By 2024, the landline-only proportion dropped to under 30%, making a dual-frame (landline + mobile) approach essential. The coverage bias is estimated as the difference between the estimate using the landline frame only and the estimate using the combined frame.
A company wants to survey job satisfaction among its 5,000 employees. The HR department provides the employee register as the sampling frame. However, some employees appear twice — once under their current department and once under a previous department that has not been updated. If 200 employees are duplicated, the effective frame has 5,200 entries. A simple random sample from this frame gives duplicated employees approximately twice the selection probability of others. If 50 duplicates are selected, and job satisfaction differs by department, the estimate will be biased. The solution is to deduplicate the frame using a unique employee ID before drawing the sample.
Data collection is the systematic process of gathering information from selected units in a survey or experiment. The choice of method affects the quality, cost, and timeliness of the data. The four principal methods are: personal interviews, telephone interviews, mail questionnaires, and web/online surveys. Each has distinct advantages and limitations.
An interviewer visits the respondent in person and records answers on a structured questionnaire. This is the most traditional and still widely used method, especially in developing countries.
Advantages: High response rates (typically 70–90%); ability to probe and clarify questions; visual aids can be used; suitable for complex questionnaires; literacy of respondent not required.
Disadvantages: Expensive (travel and training costs); time-consuming; interviewer bias (differences in how interviewers ask questions or record answers); requires extensive supervision and quality control; safety concerns in some areas.
The interviewer contacts respondents by telephone and records responses. This is faster and cheaper than face-to-face interviewing.
Advantages: Faster data collection; lower cost per interview; easier supervision (call centre monitoring); safer for interviewers; suitable for short questionnaires.
Disadvantages: Limited to households with telephones (coverage problem); shorter interviews possible; no visual aids; difficulty establishing rapport; call screening and refusals are common.
A printed questionnaire is mailed to respondents with a return envelope. The respondent fills it out independently and mails it back.
Advantages: Low cost per respondent; no interviewer bias; respondents can answer at their convenience; greater perceived anonymity; suitable for sensitive topics; can reach widely dispersed populations.
Disadvantages: Very low response rates (typically 20–40%); no control over who completes the questionnaire; cannot clarify misunderstood questions; slow turnaround; requires literacy; no probing possible.
Questionnaires are administered through the internet, typically via platforms like Google Forms, SurveyMonkey, or Qualtrics. Respondents complete the survey on a computer or mobile device.
Advantages: Very low marginal cost; rapid data collection; automatic data entry and validation; multimedia content possible; skip patterns and branching logic; real-time monitoring.
Disadvantages: Coverage limited to internet users (severe in developing countries and elderly populations); self-selection bias; multiple submissions; lack of control over respondent environment; impersonal nature.
| Criterion | Personal | Telephone | Web | |
|---|---|---|---|---|
| Response Rate | High | Moderate | Low | Low–Moderate |
| Cost per Interview | High | Moderate | Low | Very Low |
| Speed | Slow | Fast | Slow | Very Fast |
| Interviewer Bias | Yes | Yes | No | No |
| Complex Questions | Good | Moderate | Poor | Good |
| Coverage | Full | Limited | Limited | Very Limited |
| Sensitive Topics | Poor | Moderate | Good | Good |
The National Family Health Survey (NFHS) in India uses personal interviews as the primary data collection method. The reasons are: (a) many respondents in rural areas are illiterate, making self-administered questionnaires impossible; (b) the questionnaire is long (over 100 questions) and complex, with skip patterns that require an interviewer; (c) biomarker collection (height, weight, blood samples) requires in-person contact; (d) telephone and internet coverage is low in rural areas. Despite the high cost (approximately ₹3,000–5,000 per completed interview), the response rate exceeds 90%, and the data quality is considered the gold standard for health indicators in India.
A university wants to collect student feedback on 200 courses taught in the previous semester. The target population is all enrolled students (15,000), all of whom have university email addresses and internet access. A web survey is chosen because: (a) the population has full internet coverage; (b) the questionnaire is short (15 questions); (c) cost is minimal (using the university's Google Workspace account); (d) speed is important — results are needed before the next semester begins; (e) anonymity encourages honest feedback. The survey is emailed to all students, with two reminders sent at weekly intervals. The response rate is 45%, which is acceptable for course evaluations. Non-response bias is assessed by comparing early responders (who tend to be more engaged students) with late responders (who respond after reminders).
Non-response occurs when a sampled unit fails to provide the requested information. It is one of the most serious sources of non-sampling error because, if non-respondents differ systematically from respondents, the survey estimates will be biased — a problem known as non-response bias.
\[ \text{Response Rate} = \frac{\text{Number of completed interviews}}{\text{Number of eligible units in the sample}} \times 100\% \]
A response rate below 60% is generally considered a warning sign for potential non-response bias.
If \(\bar{y}_R\) is the mean of respondents and \(\bar{y}_{NR}\) is the (unknown) mean of non-respondents, and \(p_R\) is the proportion of respondents, then:
\[ E(\bar{y}_R) = \bar{y}_R, \quad \mu = p_R \bar{y}_R + (1 - p_R) \bar{y}_{NR} \]
\[ \text{Non-response Bias} = \bar{y}_R - \mu = (1 - p_R)(\bar{y}_R - \bar{y}_{NR}) \]
The bias depends on both the non-response rate \((1 - p_R)\) and the difference between respondents and non-respondents \((\bar{y}_R - \bar{y}_{NR})\).
When non-response cannot be eliminated, statistical adjustments are used:
A salary survey of 1,000 professionals in the IT sector achieves a response rate of 55% (550 respondents). The sample mean salary of respondents is \(\bar{y}_R = \text{₹}12.5\) lakh per annum. However, a follow-up study of 50 non-respondents reveals their mean salary is \(\bar{y}_{NR} = \text{₹}16.2\) lakh — significantly higher because high-earning professionals are more reluctant to disclose income. The non-response bias is:
\[ \text{Bias} = (1 - 0.55)(12.5 - 16.2) = 0.45 \times (-3.7) = -\text{₹}1.665 \text{ lakh} \]
The unadjusted estimate of ₹12.5 lakh underestimates the true population mean by approximately ₹1.67 lakh. A weighting adjustment that gives higher weight to the few high-earning respondents who did participate can partially correct this bias.
In a customer satisfaction survey of 500 respondents, 80 respondents (16%) leave the "annual income" question blank. Using regression imputation, the researcher fits a regression model: \(\text{Income} = \beta_0 + \beta_1 \times \text{Age} + \beta_2 \times \text{Education} + \beta_3 \times \text{Occupation} + \varepsilon\), using the 420 complete records. The model yields: \(\hat{\text{Income}} = 1.2 + 0.35 \times \text{Age} + 2.1 \times \text{Education}\) (in lakh ₹). For each of the 80 non-respondents, the predicted income is computed from their known age, education, and occupation, and substituted for the missing value. This preserves the sample size and reduces bias compared to listwise deletion, though it underestimates variability. Multiple imputation (generating several plausible values and combining results) is preferred as it properly accounts for imputation uncertainty.
The quality of survey data is fundamentally determined by the quality of the questions asked and the answers obtained. Poorly worded questions lead to measurement error — the difference between the true value of a variable and the value recorded by the survey. Designing good survey questions is both an art and a science, requiring attention to cognitive processes, linguistic clarity, and statistical considerations.
Open-ended questions: Respondents answer in their own words (e.g., "What is the main problem facing your community?"). These provide rich, detailed data but are difficult to code and analyse statistically.
Closed-ended questions: Respondents choose from a set of pre-specified options. These are easier to process and analyse but may not capture nuances. Common formats include:
The order of questions matters significantly. General principles include:
A municipal corporation surveys residents about a proposed waste management policy. The original question reads: "Do you support the government's excellent new waste management plan that will keep our city clean?" This is clearly leading — it associates the plan with "excellence" and "clean city," pressuring agreement. The improved version is: "Do you support or oppose the proposed waste management plan?" with response options: (1) Strongly support, (2) Somewhat support, (3) Neither support nor oppose, (4) Somewhat oppose, (5) Strongly oppose. The neutral wording and balanced response scale allow the respondent to express a genuine opinion without social pressure. In a pilot test, the leading version got 72% support, while the neutral version got 48% — a 24-percentage-point difference attributable solely to question wording.
An employee satisfaction survey includes the question: "Are you satisfied with your salary and working conditions?" with options Yes/No. This is double-barrelled — a respondent may be satisfied with salary but not with working conditions, or vice versa. The fix is to split it into two separate questions:
(1) "How satisfied are you with your salary?" — Very satisfied / Satisfied / Neutral / Dissatisfied / Very dissatisfied
(2) "How satisfied are you with your working conditions?" — Very satisfied / Satisfied / Neutral / Dissatisfied / Very dissatisfied
Additionally, the income categories in the demographics section must be mutually exclusive and exhaustive. Instead of: ₹0–5 lakh, ₹5–10 lakh, ₹10–20 lakh (where ₹5 lakh and ₹10 lakh appear in two categories), the correct formulation is: ₹0–4.99 lakh, ₹5–9.99 lakh, ₹10–19.99 lakh, ₹20 lakh and above. This ensures every respondent fits into exactly one category.