Skip to the content

Topics Covered

Data Processing Descriptive Statistics Review Inferential Statistics Review Regression and Correlation Techniques of Interpretation Precautions in Interpretation
On this page
  1. 1. Data Processing
  2. 2. Descriptive Statistics Review
  3. 3. Inferential Statistics Review
  4. 4. Regression and Correlation
  5. 5. Techniques of Interpretation
  6. 6. Precautions in Interpretation

1. Data Processing

DEFINITION

Data processing is the series of operations performed on raw data to convert it into a clean, organised, and analysable form. It encompasses editing, coding, classification, tabulation, and data entry — all prerequisite steps before any statistical analysis can be performed.

1.1 Editing

Editing is the process of examining collected data to detect and correct errors, omissions, and inconsistencies. It ensures that the data are accurate, complete, and consistent with the survey objectives.

1.2 Coding

Coding is the process of assigning numerical or alphanumeric codes to responses so that they can be processed by computers. For closed-ended questions, codes are pre-assigned (e.g., Male = 1, Female = 2). For open-ended questions, a coding scheme must be developed after reviewing the responses.

1.3 Classification

Classification organises data into homogeneous groups based on common characteristics. Data may be classified by:

1.4 Tabulation

Tabulation is the systematic arrangement of data in rows and columns. It summarises large volumes of data into a compact and interpretable form.

EXAMPLE 1 — Coding Open-ended Responses

In a survey of 500 consumers, the open-ended question "What is the main reason you prefer online shopping?" yields diverse responses. After reviewing all responses, the researcher creates the following coding scheme: Code 1 = Convenience (home delivery, 24/7 access), Code 2 = Price (discounts, comparison), Code 3 = Variety (wider selection), Code 4 = Time-saving, Code 5 = Others. Each response is assigned exactly one code. The resulting frequency distribution: Convenience (185), Price (142), Variety (98), Time-saving (52), Others (23). This coded data can now be analysed statistically — chi-square tests, proportions, and visualisations are all possible.

EXAMPLE 2 — Cross Tabulation

A researcher wants to examine the relationship between gender and preference for online versus offline shopping. The raw data from 400 respondents are cross-tabulated as follows:

GenderOnlineOfflineTotal
Male12080200
Female90110200
Total210190400

The proportion of males preferring online shopping is \(120/200 = 60\%\), while for females it is \(90/200 = 45\%\). A chi-square test of independence can determine whether this difference is statistically significant: \(\chi^2 = \frac{(120 - 105)^2}{105} + \frac{(80 - 95)^2}{95} + \frac{(90 - 105)^2}{105} + \frac{(110 - 95)^2}{95} = 2.143 + 2.368 + 2.143 + 2.368 = 9.022\). With 1 degree of freedom, the critical value at \(\alpha = 0.05\) is 3.841. Since \(9.022 > 3.841\), we reject the null hypothesis — gender and shopping preference are associated.

2. Descriptive Statistics Review

KEY CONCEPT

Descriptive statistics summarise and describe the main features of a data set without drawing conclusions beyond the data at hand. They provide the foundation for all further analysis and are essential for understanding the distribution, central tendency, and variability of research variables.

2.1 Measures of Central Tendency

KEY FORMULAS

Arithmetic Mean: \(\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i\)

Weighted Mean: \(\bar{x}_w = \frac{\sum w_i x_i}{\sum w_i}\)

Median: The middle value when data are arranged in order. For grouped data:

\[ \text{Median} = L + \frac{(n/2 - cf)}{f} \times h \]

Mode: The most frequently occurring value.

2.2 Measures of Dispersion

KEY FORMULAS

Variance: \(s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^2\)

Standard Deviation: \(s = \sqrt{s^2}\)

Coefficient of Variation: \(CV = \frac{s}{\bar{x}} \times 100\%\)

Interquartile Range: \(IQR = Q_3 - Q_1\)

2.3 Measures of Shape

SKEWNESS AND KURTOSIS

Skewness:

\[ g_1 = \frac{n}{(n-1)(n-2)} \sum_{i=1}^{n} \left(\frac{x_i - \bar{x}}{s}\right)^3 \]

Kurtosis (excess):

\[ g_2 = \frac{n(n+1)}{(n-1)(n-2)(n-3)} \sum_{i=1}^{n} \left(\frac{x_i - \bar{x}}{s}\right)^4 - \frac{3(n-1)^2}{(n-2)(n-3)} \]

EXAMPLE 1 — Descriptive Summary of Crop Yields

An agricultural researcher records the yield (in quintals per hectare) of a new paddy variety from 8 trial plots: 42, 38, 45, 50, 35, 48, 52, 40.

Mean: \(\bar{x} = \frac{42+38+45+50+35+48+52+40}{8} = \frac{350}{8} = 43.75\) q/ha

Variance: \(s^2 = \frac{1}{7}\sum(x_i - 43.75)^2 = \frac{1}{7}(3.0625 + 33.0625 + 1.5625 + 39.0625 + 76.5625 + 18.0625 + 68.0625 + 14.0625) = \frac{253.5}{7} = 36.21\)

Standard Deviation: \(s = \sqrt{36.21} = 6.018\) q/ha

CV: \(\frac{6.018}{43.75} \times 100 = 13.76\%\)

Median: Arranging in order — 35, 38, 40, 42, 45, 48, 50, 52. Median = \(\frac{42+45}{2} = 43.5\) q/ha.

The slight positive difference between mean (43.75) and median (43.5) suggests mild right skewness.

EXAMPLE 2 — Comparing Two Distributions Using CV

A pharmaceutical company measures the weight (mg) of tablets from two production lines.

Line A: \(\bar{x}_A = 500\) mg, \(s_A = 10\) mg → \(CV_A = \frac{10}{500} \times 100 = 2.0\%\)

Line B: \(\bar{x}_B = 50\) mg, \(s_B = 5\) mg → \(CV_B = \frac{5}{50} \times 100 = 10.0\%\)

Although Line A has a larger absolute standard deviation, Line B is more variable relative to its mean. The coefficient of variation reveals that Line B is five times more variable in relative terms. This is crucial for quality control: Line B tablets vary by ±10% of the target, while Line A varies by only ±2%. Comparing standard deviations alone would be misleading because the two lines produce tablets of vastly different sizes.

3. Inferential Statistics Review

KEY CONCEPT

Inferential statistics use sample data to draw conclusions about the population from which the sample was drawn. The two main branches are estimation (point and interval) and hypothesis testing. In research, inferential methods allow us to generalise findings beyond the observed sample with a quantifiable level of confidence.

3.1 Estimation

CONFIDENCE INTERVAL FOR THE MEAN

For a population mean \(\mu\) with unknown \(\sigma\) (using t-distribution):

\[ \bar{x} \pm t_{\alpha/2,\, n-1} \cdot \frac{s}{\sqrt{n}} \]

For a population proportion \(p\):

\[ \hat{p} \pm z_{\alpha/2} \cdot \sqrt{\frac{\hat{p}(1-\hat{p})}{n}} \]

3.2 Hypothesis Testing Framework

STEPS IN HYPOTHESIS TESTING
  1. State the null hypothesis \(H_0\) and the alternative hypothesis \(H_1\).
  2. Choose the significance level \(\alpha\) (commonly 0.05 or 0.01).
  3. Select the appropriate test statistic (z, t, \(\chi^2\), F).
  4. Compute the test statistic from the sample data.
  5. Determine the p-value or compare with the critical value.
  6. Make a decision: reject \(H_0\) if p-value < \(\alpha\).

3.3 Commonly Used Tests

EXAMPLE 1 — Two-Sample t-test for Teaching Methods

A researcher compares the effectiveness of two teaching methods on student test scores.

Method A: \(n_1 = 30\), \(\bar{x}_1 = 72.5\), \(s_1 = 8.3\)

Method B: \(n_2 = 28\), \(\bar{x}_2 = 67.8\), \(s_2 = 9.1\)

Pooled variance: \(s_p^2 = \frac{29 \times 68.89 + 27 \times 82.81}{56} = \frac{1997.81 + 2235.87}{56} = \frac{4233.68}{56} = 75.60\), so \(s_p = 8.695\)

Test statistic: \(t = \frac{72.5 - 67.8}{8.695\sqrt{\frac{1}{30}+\frac{1}{28}}} = \frac{4.7}{8.695 \times 0.2628} = \frac{4.7}{2.285} = 2.057\)

With 56 degrees of freedom, the critical value at \(\alpha = 0.05\) (two-tailed) is approximately \(\pm 2.003\). Since \(|t| = 2.057 > 2.003\), we reject \(H_0\). There is a statistically significant difference between the two teaching methods at the 5% level. The 95% confidence interval for the difference is \(4.7 \pm 2.003 \times 2.285 = [0.12,\; 9.28]\), suggesting Method A scores between 0.12 and 9.28 points higher than Method B.

EXAMPLE 2 — Chi-square Test for Brand Preference and Age Group

A market researcher investigates whether brand preference for a soft drink is independent of age group. Survey data from 300 consumers:

Age GroupBrand ABrand BBrand CTotal
18–25403030100
26–40254530100
41+152560100
Total80100120300

Expected frequency for each cell: \(E_{ij} = \frac{R_i \times C_j}{N}\). For example, \(E_{11} = \frac{100 \times 80}{300} = 26.67\).

The row totals are all 100 and the column totals are 80, 100, 120, so the expected values are 26.67, 33.33, 40 in every row. Summing \((O-E)^2/E\) over all nine cells:

\(\chi^2 = 6.67 + 0.33 + 2.50 + 0.10 + 4.08 + 2.50 + 5.10 + 2.08 + 10.00 \approx 33.38\)

With \((3-1)(3-1) = 4\) degrees of freedom, the critical value at \(\alpha = 0.05\) is 9.488. Since \(33.38 \gg 9.488\), we strongly reject \(H_0\). Brand preference is highly dependent on age group. Younger consumers prefer Brand A, middle-aged consumers prefer Brand B, and older consumers prefer Brand C.

4. Regression and Correlation

KEY CONCEPT

Regression analysis models the relationship between a dependent variable and one or more independent variables, enabling prediction. Correlation analysis measures the strength and direction of the linear association between two variables. Together, they are among the most widely used tools in research data analysis.

4.1 Simple Linear Regression

REGRESSION MODEL

The simple linear regression model:

\[ Y = \beta_0 + \beta_1 X + \varepsilon \]

where \(\beta_0\) is the intercept, \(\beta_1\) is the slope, and \(\varepsilon\) is the random error term with \(E(\varepsilon) = 0\) and \(\text{Var}(\varepsilon) = \sigma^2\).

Least squares estimates:

\[ \hat{\beta}_1 = \frac{\sum(x_i - \bar{x})(y_i - \bar{y})}{\sum(x_i - \bar{x})^2} = \frac{S_{xy}}{S_{xx}} \]

\[ \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} \]

Coefficient of determination:

\[ R^2 = \frac{SSR}{SST} = 1 - \frac{SSE}{SST} \]

4.2 Correlation

PEARSON CORRELATION COEFFICIENT

\[ r = \frac{\sum(x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum(x_i - \bar{x})^2 \sum(y_i - \bar{y})^2}} = \frac{S_{xy}}{\sqrt{S_{xx} \cdot S_{yy}}} \]

where \(-1 \leq r \leq 1\). Test of significance: \(t = r\sqrt{\frac{n-2}{1-r^2}}\) with \(n-2\) df.

4.3 Multiple Regression

MULTIPLE REGRESSION MODEL

\[ Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_k X_k + \varepsilon \]

Adjusted \(R^2\):

\[ R^2_{\text{adj}} = 1 - \frac{(1-R^2)(n-1)}{n-k-1} \]

F-test for overall significance: \(F = \frac{MSR}{MSE} = \frac{R^2/k}{(1-R^2)/(n-k-1)}\)

EXAMPLE 1 — Predicting House Prices with Simple Regression

A real estate researcher studies the relationship between house area (sq. ft., \(X\)) and price (₹ lakh, \(Y\)) for 6 houses:

Area (X)100012001500180020002500
Price (Y)253238455260

\(\bar{x} = 1666.67\), \(\bar{y} = 42.0\), \(S_{xy} = 35400\), \(S_{xx} = 1513333.33\)

\(\hat{\beta}_1 = \frac{35400}{1513333.33} = 0.02339\), \(\hat{\beta}_0 = 42 - 0.02339 \times 1666.67 = 3.01\)

Regression equation: \(\hat{Y} = 3.01 + 0.02339 X\)

Interpretation: For every additional square foot, the price increases by approximately ₹0.0234 lakh (₹2,339). The intercept (3.01) has no practical meaning on its own (it is the extrapolated price of a house with zero area). \(R^2 = 0.988\), meaning 98.8% of the variation in price is explained by area. The correlation coefficient is \(r = 0.994\), indicating a very strong positive linear relationship.

EXAMPLE 2 — Multiple Regression for Crop Yield

An agricultural researcher models crop yield (\(Y\), tonnes/ha) as a function of rainfall (\(X_1\), cm) and fertiliser (\(X_2\), kg/ha) using data from 30 farms. The fitted model is:

\[ \hat{Y} = 0.85 + 0.032 X_1 + 0.018 X_2 \]

with \(R^2 = 0.82\) and adjusted \(R^2 = 0.807\).

Interpretation: (a) Holding fertiliser constant, each additional cm of rainfall increases yield by 0.032 tonnes/ha. (b) Holding rainfall constant, each additional kg/ha of fertiliser increases yield by 0.018 tonnes/ha. (c) The adjusted \(R^2 = 0.807\) means about 81% of the variation in yield is explained by the two predictors jointly. (d) The F-statistic is \(F = \frac{0.82/2}{0.18/27} = 61.5\) which is highly significant (p < 0.001). Both predictors have significant t-statistics (p < 0.05). The researcher should check residual diagnostics (normality, homoscedasticity, independence) before trusting these results.

5. Techniques of Interpretation

DEFINITION

Interpretation is the process of giving meaning to the results of statistical analysis — relating the numbers back to the research questions, theory, and real-world context. It is the bridge between statistical output and substantive conclusions.

Effective interpretation requires the researcher to move through several levels:

5.1 Descriptive Interpretation

What do the numbers say? Describe patterns, central tendencies, and variability in plain language. For example, "The average income of urban households is 2.3 times that of rural households."

5.2 Inferential Interpretation

What can we generalise? Translate test statistics and p-values into substantive conclusions. For example, "The difference in mean income between urban and rural households is statistically significant (t = 5.42, p < 0.001), confirming that urban households earn significantly more."

5.3 Causal Interpretation

Why does this relationship exist? This is the most challenging level and requires careful distinction between association and causation. Causal claims are only justified with experimental designs (randomised controlled trials) or strong quasi-experimental methods (instrumental variables, regression discontinuity, difference-in-differences).

5.4 Practical Significance vs. Statistical Significance

A statistically significant result may be practically trivial, and a non-significant result may be practically important. Always consider the effect size and the confidence interval, not just the p-value.

COMMON EFFECT SIZE MEASURES
EXAMPLE 1 — Statistical Significance Without Practical Significance

A pharmaceutical company tests a new drug to reduce blood pressure. The trial involves 10,000 patients. The treatment group has a mean reduction of 2.1 mm Hg and the placebo group has a mean reduction of 1.8 mm Hg. The t-test yields \(t = 3.45\), \(p = 0.0006\) — highly statistically significant. However, the effect size is Cohen's \(d = \frac{2.1 - 1.8}{12} = 0.025\), which is negligible. A blood pressure reduction of 0.3 mm Hg has no clinical relevance whatsoever. The statistical significance is entirely due to the enormous sample size, which makes even a trivially small detectable difference significant. The correct interpretation is: "The drug produces a statistically significant but clinically meaningless reduction in blood pressure."

EXAMPLE 2 — Interpreting a Regression in Context

An education researcher fits the model: \(\widehat{\text{Score}} = 45.2 + 3.8 \times \text{StudyHours} + 2.1 \times \text{TutorAttendance}\) with \(R^2 = 0.64\). The coefficient of StudyHours is 3.8 with \(p = 0.002\) and a 95% CI of [1.5, 6.1]. The coefficient of TutorAttendance is 2.1 with \(p = 0.12\) and a 95% CI of [-0.6, 4.8].

Interpretation: (a) Each additional hour of study per week is associated with a 3.8-point increase in test score, and we are 95% confident the true effect is between 1.5 and 6.1 points. (b) Having a tutor is associated with a 2.1-point score increase, but this is not statistically significant at \(\alpha = 0.05\) — the CI includes zero, so we cannot rule out that tutoring has no effect. (c) The model explains 64% of score variation. (d) This is an observational study, so the coefficients represent associations, not causal effects. Students who study more may also be more motivated, which could confound the relationship.

6. Precautions in Interpretation

KEY PRINCIPLE

Interpretation is where research conclusions are drawn, and it is also where the most serious errors occur. A researcher must exercise constant vigilance against common fallacies, misinterpretations, and overstatements. The following precautions are essential.

6.1 Correlation Does Not Imply Causation

This is the single most important precaution. Just because two variables are correlated does not mean one causes the other. The correlation may be due to:

6.2 Avoid Overgeneralisation

Conclusions apply only to the population from which the sample was drawn. A study of college students in one city cannot be generalised to all Indians. A survey of internet users cannot represent the general population. Always specify the limits of generalisability.

6.3 Distinguish Between Practical and Statistical Significance

As discussed earlier, always report effect sizes and confidence intervals alongside p-values. A large sample can make trivial effects significant; a small sample can miss important effects.

6.4 Beware of Multiple Comparisons

If you perform many statistical tests, some will be significant by chance alone. With 20 tests at \(\alpha = 0.05\), you expect one false positive on average. Use corrections such as Bonferroni (divide \(\alpha\) by the number of tests) or Benjamini-Hochberg (control the false discovery rate).

BONFERRONI CORRECTION

If \(k\) tests are performed, use significance level \(\alpha^* = \alpha/k\).

For example, with \(k = 10\) tests and \(\alpha = 0.05\): \(\alpha^* = 0.005\).

6.5 Consider the Research Design

The strength of conclusions depends on the design. Randomised experiments support causal claims; observational studies do not. Cross-sectional studies cannot establish temporal ordering; longitudinal studies can. The interpretation must be consistent with the design's capacity for drawing conclusions.

6.6 Check Assumptions

Every statistical test rests on assumptions (normality, independence, homoscedasticity, linearity). If assumptions are violated, the results may be invalid. Always check assumptions before interpreting results. Use diagnostic plots (residual plots, Q-Q plots) and formal tests (Shapiro-Wilk, Levene's test).

6.7 Acknowledge Limitations

Every study has limitations — in design, sampling, measurement, or analysis. Honest reporting of limitations strengthens, not weakens, the credibility of the research.

EXAMPLE 1 — Spurious Correlation Due to Confounding

A researcher finds a strong positive correlation (\(r = 0.85\)) between the number of churches in a city and the crime rate. Naively interpreting this as "churches cause crime" is absurd. The confounding variable is population size — larger cities have more churches AND more crime simply because they have more people. When the researcher controls for population by computing the partial correlation between churches and crime rate holding population constant, the correlation drops to \(r = 0.03\) — essentially zero. This illustrates how a third variable can create a completely spurious association between two unrelated variables.

EXAMPLE 2 — Multiple Comparisons Fallacy

A nutrition researcher tests the effect of a new supplement on 20 different health biomarkers (cholesterol, blood pressure, glucose, iron, etc.) using a sample of 50 subjects. Using \(\alpha = 0.05\) for each test, the researcher finds that the supplement significantly reduces C-reactive protein (p = 0.03) and significantly increases vitamin D (p = 0.04). The researcher announces that the supplement "has significant health benefits." However, with 20 independent tests at \(\alpha = 0.05\), the expected number of false positives is \(20 \times 0.05 = 1.0\). Finding 2 significant results out of 20 is entirely consistent with chance. Using the Bonferroni correction, the adjusted significance level is \(\alpha^* = 0.05/20 = 0.0025\). Neither result meets this threshold. The correct interpretation is: "No significant effects were found after correcting for multiple comparisons. The apparent effects on CRP and vitamin D could easily be due to chance."