Skip to the content

Topics Covered

Variable View Syntax Charts Descriptives Parametric Tests Non-Parametric Tests UNIANOVA Model Selection Logistic & Probit Multivariate Control Charts
On this page
  1. The Data Set
  2. Practical 1: Data Entry, Import, Export and File Handling
  3. Practical 2: Data Visualization
  4. Practical 3: Descriptive Statistics
  5. Practical 4: Parametric Tests
  6. Practical 5: Non-Parametric Tests
  7. Practical 6: Design and Analysis of Experiments
  8. Practical 7: Regression Analysis
  9. Practical 8: Multivariate Data Analysis
  10. Practical 9: Statistical Quality Control
  11. How Marks Are Lost
  12. What the Practical Record Should Contain
About this course. STS-207 is a practical whose stated outcome is “Able to carry out the Statistical Analysis and writing statistical Report using SPSS for any dataset.” Its nine headings are nine bodies of statistics that are taught elsewhere on this site. None of that theory is written again here. What this page supplies is the part that belongs to this course and nowhere else: the SPSS procedure, the syntax, the names of the output tables, and how to read them.
About the numbers. SPSS is proprietary and is not run on this site. Every figure below was computed independently, in double precision, on the data set printed in full in the next section, and is quoted so that an SPSS run can be checked against it; each was rechecked in R 4.3.3. Where SPSS's own convention matters — which sum of squares UNIANOVA uses by default, which correlation FACTOR starts from, which \(F\) REGRESSION uses to enter a variable — it is named.

The data set is the one used in STS-108, Data Handling using R, with one column added. Running both papers on the same twenty records is deliberate: the output of the two packages can then be compared line for line, and a difference is either a convention or a mistake.

The Data Set

TWENTY CASES, SEVEN VARIABLES
VariableMeasure in SPSSValues
idNominal1 to 20
ageScale 23 45 31 58 27 39 52 34 29 48 36 41 25 55 33 44 30 50 28 38
incomeScale 40 57 37 69 36 38 63 32 35 57 42 49 26 39 48 49 23 57 27 33
visitsScale 7 5 8 10 6 5 7 9 3 5 6 7 4 2 6 5 10 5 9 5
scoreScale 54 55 64 77 69 60 85 57 65 59 64 60 51 57 60 60 55 67 40 56
genderNominalM F alternating, 10 each
buyNominal (0 = No, 1 = Yes) 0 1 0 1 1 0 1 0 0 1 0 0 0 0 1 1 0 1 0 0

grade is derived from score as Low below 57, Med from 57 to 64, High from 65 — and is Ordinal.

Practical 1: Data Entry, Import, Export and File Handling

1. Question

Read the twenty cases of the data set into SPSS from customers.csv; define every variable in Variable View — its label, value labels, missing-value code and Measure; derive the ordinal grade from score; check the coding with frequency tables; and save the file as .sav and as an Excel workbook.

2. Aim

Set up a data file that every later procedure can trust: the right Measure for each variable, labelled codes and declared missing values.

3. Steps

  1. File → Import Data → CSV Data (GET DATA) reads the file; name the data set.
  2. In Variable View give each variable its label, its value labels, its missing-value code and its Measure: Nominal, Ordinal or Scale.
  3. Transform → Recode into Different Variables makes grade from score, each class closed on the left; label it, set it Ordinal, and run EXECUTE.
  4. Analyze → Descriptive Statistics → Frequencies on grade, gender and buy checks the coding against the data.
  5. File → Save As writes the .sav; SAVE TRANSLATE (File → Export) writes the Excel copy.
  6. Press Paste, not OK, in every dialog, so that each step is written into the Syntax Editor.
MENUS, COMMANDS AND CONVENTIONS
WindowHoldsSaved as
Data Editor — Data Viewthe cases, one per row .sav
Data Editor — Variable View the definitions: name, type, width, decimals, label, value labels, missing, measurethe same .sav
Output Viewerevery table and chart produced.spv
Syntax Editorthe commands.sps

The Measure column is not cosmetic. SPSS uses it to decide which procedures a variable may enter, which charts the Chart Builder will offer, and how a variable is treated inside a model. Its three values map onto the four measurement scales:

SPSS MeasureScale of measurementConsequence
Nominalnominal counts and \(\chi^{2}\) only; enters models as a set of dummy variables
Ordinalordinal median and rank methods; ordinal regression accepts it as the outcome
Scaleinterval or ratio means, standard deviations, correlations, everything

SPSS does not distinguish interval from ratio, so the one restriction it cannot enforce for you is the one on ratios: a score of 80 is not “twice” a score of 40 unless the zero is real. That judgement stays with the analyst.

Value labels are entered once and used everywhere. Coding buy as 0 and 1 and labelling them No and Yes means every later table reads “No” and “Yes” rather than 0 and 1 — and the coding, not the label, is what the model uses. In a binary logistic regression SPSS models the higher code, so labelling them the wrong way round reverses the sign of every coefficient.

Every menu path ends in a Paste button, which writes the command into the Syntax Editor instead of running it. Use it always. Menus leave no record of what was done; syntax is the record, it can be re-run, and it is what a practical record must contain.

4. Programme

PRACTICAL 1 — SPSS SYNTAX
* Read a CSV file.  SPSS syntax comments begin with * and end with a full stop.
GET DATA /TYPE=TXT /FILE='C:\data\customers.csv'
  /DELIMITERS=',' /FIRSTCASE=2 /VARIABLES=
  id F3.0 age F3.0 income F4.0 visits F3.0 score F4.0 gender A1 buy F1.0.
DATASET NAME customers WINDOW=FRONT.

* Define the variables properly -- this is the step most often skipped.
VARIABLE LABELS income 'Monthly income, thousands' score 'Assessment score'.
VALUE LABELS buy 0 'No' 1 'Yes'.
VARIABLE LEVEL id gender buy (NOMINAL) / age income visits score (SCALE).
MISSING VALUES income (-99).

* Derive the ordinal grade, closing each class on the LEFT.
RECODE score (LO THRU 56.999=1) (57 THRU 64.999=2) (65 THRU HI=3) INTO grade.
VALUE LABELS grade 1 'Low' 2 'Med' 3 'High'.
VARIABLE LEVEL grade (ORDINAL).
EXECUTE.

FREQUENCIES VARIABLES=grade gender buy.
* grade: Low 6, Med 9, High 5.   gender: F 10, M 10.   buy: No 12, Yes 8.

* Export.
SAVE OUTFILE='C:\data\customers.sav' /COMPRESSED.
SAVE TRANSLATE OUTFILE='C:\data\customers.xlsx' /TYPE=XLS /VERSION=12
  /FIELDNAMES /CELLS=VALUES.

5. Execution and Results

FREQUENCIES — THE COUNTS SPSS PRINTS (RECOMPUTED IN R)
VariableCategoryFrequency
gradeLow6
Med9
High5
genderF10
M10
buyNo12
Yes8

Three file-handling traps. EXECUTE is needed after transformations such as RECODE and COMPUTE before the new values exist — SPSS defers them until a procedure reads the data. /CELLS=VALUES exports the codes and /CELLS=LABELS exports the labels; choosing the second and then re-importing turns a numeric variable into a string. And a declared missing value (MISSING VALUES income (-99)) is excluded from analysis, while an undeclared \(-99\) is treated as income of minus ninety-nine thousand and will quietly ruin every mean.

RESULT

The file holds the twenty cases with every variable defined. The frequency tables show grade Low 6, Med 9, High 5; gender 10 and 10; buy No 12 and Yes 8, which agree with the data as entered, so the coding and the recode are right. The .sav keeps the definitions; the Excel copy keeps only the values.

Practical 2: Data Visualization

1. Question

For the data set, draw a pie chart of grade, a bar chart of gender, a histogram of score with a normal curve, a scatter plot of score against income, boxplots of score by gender, and an ogive of score.

2. Aim

Produce each prescribed chart from syntax, so that every one can be drawn again exactly.

3. Steps

  1. Graphs → Legacy Dialogs → Pie and Bar for the categorical variables, counting cases.
  2. Graphs → Legacy Dialogs → Histogram for score, with Display normal curve ticked.
  3. Graphs → Legacy Dialogs → Scatter/Dot, score on the vertical axis and income on the horizontal.
  4. Analyze → Descriptive Statistics → Explore for the boxplots of score by gender.
  5. The ogive, which has no button: the cumulative count of score drawn as a line.
  6. Paste each dialog into the Syntax Editor instead of pressing OK.
MENUS, COMMANDS AND CONVENTIONS

The theory of which diagram suits which variable is in STS-108, section 4 and is not repeated. SPSS offers three routes to the same charts, and the syntax route is the only reproducible one.

ChartMenuSyntax command
PieGraphs → Legacy Dialogs → PieGRAPH /PIE
BarGraphs → Legacy Dialogs → BarGRAPH /BAR
HistogramGraphs → Legacy Dialogs → Histogram GRAPH /HISTOGRAM
LineGraphs → Legacy Dialogs → LineGRAPH /LINE
ScatterGraphs → Legacy Dialogs → Scatter/Dot GRAPH /SCATTERPLOT
BoxplotGraphs → Legacy Dialogs → Boxplot EXAMINE ... /PLOT BOXPLOT
Frequency polygon, ogiveChart Builder, or a line over binned data GGRAPH
Ganttnot a built-in chart a clustered bar of durations with a transparent first segment

4. Programme

PRACTICAL 2 — SPSS SYNTAX
GRAPH /PIE=COUNT BY grade.
GRAPH /BAR(SIMPLE)=COUNT BY gender.
GRAPH /HISTOGRAM(NORMAL)=score.
GRAPH /SCATTERPLOT(BIVAR)=income WITH score /MISSING=LISTWISE.
EXAMINE VARIABLES=score BY gender /PLOT=BOXPLOT /STATISTICS=NONE /NOTOTAL.

* the ogive, which SPSS has no button for: cumulative percentage against the class
FREQUENCIES VARIABLES=score /FORMAT=NOTABLE /NTILES=10.
GRAPH /LINE(SIMPLE)=CUM(COUNT) BY score.

5. Execution and Results

SPSS draws each chart in the Output Viewer. The charts are not reproduced here, because SPSS is not run on this site, but the figures they are drawn from can be checked: the normal curve on the histogram has the mean 60.75 and standard deviation 9.4361 of score (Practical 3), the pie has the slices Low 6, Med 9 and High 5 (Practical 1), and the two bars are 10 and 10.

GRAPH /HISTOGRAM(NORMAL) superimposes a normal curve with the sample's own mean and standard deviation. It is a picture, not a test: the Shapiro–Wilk statistic in EXAMINE ... /PLOT NPPLOT is the test, and on a sample of twenty it has very little power, so a histogram that “looks normal” and a Shapiro–Wilk that does not reject are both weak evidence.

RESULT

Six charts, each from one line of syntax that redraws it exactly. The scatter plot shows the upward trend that the correlation of 0.6536 between income and score measures (Practical 7). The histogram with its normal curve is a picture of the shape of score, not a test of normality.

Practical 3: Descriptive Statistics

1. Question

Find the mean, median, standard deviation, minimum, maximum and skewness of age, income and score; the median and mode of grade; and the summaries of score for each gender.

2. Aim

Choose the right one of SPSS's three descriptive procedures for each summary, and know which conventions its output uses.

3. Steps

  1. Analyze → Descriptive Statistics → Descriptives for the scale variables, with skewness and kurtosis ticked under Options.
  2. Frequencies for grade, with the median and the mode under Statistics: it is the only one of the three that gives the mode.
  3. Explore for score by gender, with boxplots and normality plots.
  4. Before quoting the skewness, convert between SPSS's coefficient and the moment coefficient, by the factor under the table.
MENUS, COMMANDS AND CONVENTIONS
ProcedureGivesUse for
FREQUENCIES the frequency table, plus any statistic you ask for, including the mode and percentilescategorical variables; and it is the only one that gives the mode
DESCRIPTIVES mean, standard deviation, minimum, maximum, skewness, kurtosis — no mediana compact table of scale variables; and \(z\) scores, via /SAVE
EXAMINE everything, by group: the five-number summary, trimmed mean, confidence interval, normality tests, boxplots and stem-and-leaf anything that has to be broken down by a factor

4. Programme

PRACTICAL 3 — SPSS SYNTAX
DESCRIPTIVES VARIABLES=age income score
  /STATISTICS=MEAN STDDEV MIN MAX SKEWNESS KURTOSIS.
FREQUENCIES VARIABLES=grade /STATISTICS=MEDIAN MODE.
EXAMINE VARIABLES=score BY gender /PLOT BOXPLOT NPPLOT /STATISTICS DESCRIPTIVES.

5. Execution and Results

DESCRIPTIVE STATISTICS (RECOMPUTED IN R; SKEWNESS AS THE MOMENT COEFFICIENT)
VariableMeanMedianStd. DeviationMinimum MaximumSkewness
age38.300037.000010.45342358 0.3507
income42.850039.500012.82792369 0.3874
score60.750060.00009.43614085 0.5613

Two conventions to state in the report. SPSS's Std. Deviation divides by \(n-1\); and the skewness SPSS prints is not the moment coefficient \(g_1 = m_3/m_2^{3/2}\) but the sample-adjusted version

\[ G_1 = \frac{n}{(n-1)(n-2)}\sum\left(\frac{x-\bar x}{s}\right)^{3} = g_1\,\frac{\sqrt{n(n-1)}}{n-2}, \]

so at \(n = 20\) it is larger by a factor \(\sqrt{380}/18 = 1.082977\), or \(8.3\%\). The table above gives the moment coefficients, so that they match the R and Python pages; SPSS will print \(0.379812\), \(0.419544\) and \(0.607924\) for the three variables, and the difference is the convention and not an error.

RESULT

Mean score 60.75 with standard deviation 9.44; age 38.3 (10.45); income 42.85 (12.83). The medians come from FREQUENCIES or EXAMINE, since DESCRIPTIVES gives none. All three variables are mildly skewed to the right: moment skewness 0.35, 0.39 and 0.56, which SPSS prints as 0.380, 0.420 and 0.608. The median of grade is Med (the 10th and 11th of the twenty ordered grades are both Med) and so is the mode, with 9 cases.

Practical 4: Parametric Tests

1. Question

Test whether the mean score differs between men and women, and run a one-way analysis of variance of score by grade. Say what each result means.

2. Aim

Run the two-sample \(t\) test in the right order, the variances first, and recognise an analysis that SPSS will run but that answers nothing.

3. Steps

  1. Analyze → Compare Means → Independent-Samples T Test: score as the test variable, gender as the grouping variable with the groups M and F.
  2. Read Levene's test first, then the \(t\) row it points to.
  3. Analyze → Compare Means → One-Way ANOVA: score by grade, with the descriptives, the homogeneity test and Tukey's and the LSD post hoc tests.
  4. Before reading the ANOVA, ask whether the groups could differ by construction.
MENUS, COMMANDS AND CONVENTIONS

The theory is in Inferential Statistics and Estimation Theory (STS-201).

TestMenuSyntax
One meanAnalyze → Compare Means → One-Sample T Test T-TEST /TESTVAL=
Two independent means Analyze → Compare Means → Independent-Samples T Test T-TEST GROUPS=
Paired meansAnalyze → Compare Means → Paired-Samples T Test T-TEST PAIRS=
One proportionAnalyze → Nonparametric Tests → Binomial NPAR TESTS /BINOMIAL
One-way ANOVAAnalyze → Compare Means → One-Way ANOVA ONEWAY
Two-way and factorial ANOVA Analyze → General Linear Model → Univariate UNIANOVA

The independent-samples table has two rows, and choosing between them is the examinable step. SPSS prints Levene's test for equality of variances first, then a \(t\) assuming equal variances and a \(t\) not assuming it (Welch, with fractional degrees of freedom). Read Levene, then read the row it points to.

4. Programme

PRACTICAL 4 — SPSS SYNTAX
T-TEST GROUPS=gender('M' 'F') /VARIABLES=score /CRITERIA=CI(.95).
ONEWAY score BY grade /STATISTICS DESCRIPTIVES HOMOGENEITY /POSTHOC=TUKEY LSD ALPHA(.05).

5. Execution and Results

INDEPENDENT-SAMPLES T TEST (RECOMPUTED IN R)
Output lineValue
Mean, gender = M60.7000
Mean, gender = F60.8000
Variance, M144.4556
Variance, F43.5111
Levene's Test for Equality of Variances, \(F\)2.1833
Levene's Test, Sig.0.1568
\(t\), equal variances assumed, df = 18−0.0231
\(t\), equal variances not assumed, df = 13.971−0.0231
Sig. (2-tailed), both rows0.9819
Variance ratio \(s_M^{2}/s_F^{2}\) (not printed by SPSS)3.3200

The two means differ by \(0.1\) and the variances by a factor of \(3.32\). A \(t\) test of means answers a question this data does not raise; the difference here is entirely one of spread, which is why the variance line is read first and reported whatever it says.

The ANOVA of score by grade is on the page as a trap. It gives \(F = 19.756\) on \((2, 17)\) with Sig. \(= 0.000\), and it is worthless: grade was recoded out of score, so the groups differ by construction. SPSS will print the table without complaint. A procedure that runs is not a result.

Corrected. The table listed “\(F\) for equality of variances (variance ratio) 3.3200”. SPSS does not print the variance ratio there: the \(F\) in that column is Levene's, an analysis of variance of the absolute deviations from each group's mean. On this data it is 2.1833 with Sig. 0.1568, rechecked in R. So the larger spread of the men's scores is not significant on twenty cases, and the “Equal variances assumed” row is the one to read. With ten in each group the Welch \(t\) is the same \(-0.0231\), on 13.97 df.
RESULT

The means are 60.7 and 60.8: \(t = -0.023\) on 18 df, Sig. \(= 0.982\), so there is no evidence that the mean score differs by gender. Levene's test (Sig. 0.157) does not reject equal variances, although the men's variance is 3.3 times the women's. The ANOVA of score by grade (\(F = 19.76\)) is not a finding, because grade is made from score.

Practical 5: Non-Parametric Tests

1. Question

Run the sign and Wilcoxon signed-rank tests on a pair of before-and-after variables, the Mann–Whitney test of score by gender, the runs test and the one-sample Kolmogorov–Smirnov test on score, and the \(\chi^{2}\) test of independence of gender and grade.

2. Aim

Find each prescribed non-parametric test in SPSS, and read the warnings its output prints.

3. Steps

  1. Analyze → Nonparametric Tests → Legacy Dialogs → 2 Related Samples, with Sign and Wilcoxon ticked, for the paired variables.
  2. 2 Independent Samples → Mann-Whitney U, score by gender.
  3. Runs, cut at the median; and 1-Sample K-S against the normal, for score.
  4. Analyze → Descriptive Statistics → Crosstabs, gender by grade, with Chi-square under Statistics and the expected counts under Cells.
  5. Read the footnote under the Chi-Square Tests table before the Sig. column.
MENUS, COMMANDS AND CONVENTIONS

What each test uses of the data, and what it throws away, is worked in STS-108, Practical 8 — including the comparison where the sign test gives \(p = 0.146\) on data the Wilcoxon signed rank rejects at \(p = 0.0086\). Here: which dialog, and what the output is called.

TestMenu (Legacy Dialogs)Syntax
Sign2 Related Samples → Sign/SIGN=
Wilcoxon signed rank2 Related Samples → Wilcoxon /WILCOXON=
Mann–Whitney \(U\)2 Independent Samples → Mann-Whitney U /M-W=
RunsRuns/RUNS(MEDIAN)=
Kolmogorov–Smirnov, one sample1-Sample K-S /K-S(NORMAL)=
\(\chi^{2}\) goodness of fitChi-Square/CHISQUARE=
\(\chi^{2}\) independenceDescriptive Statistics → Crosstabs CROSSTABS /STATISTICS=CHISQ

4. Programme

PRACTICAL 5 — SPSS SYNTAX
NPAR TESTS /SIGN=before WITH after (PAIRED) /WILCOXON=before WITH after (PAIRED).
NPAR TESTS /M-W=score BY gender('M' 'F').
NPAR TESTS /RUNS(MEDIAN)=score.
NPAR TESTS /K-S(NORMAL)=score.
CROSSTABS TABLES=gender BY grade /STATISTICS=CHISQ /CELLS=COUNT EXPECTED ROW.

5. Execution and Results

CROSSTABS, GENDER BY GRADE (RECOMPUTED IN R)
genderLowMedHighTotal
F, count26210
F, expected3.04.52.510.0
M, count43310
M, expected3.04.52.510.0
Total69520

Pearson Chi-Square \(1.867\), df 2, Asymptotic Sig. (2-sided) \(0.393\); footnote: 6 cells (100.0%) have expected count less than 5, the minimum expected count being 2.50.

Read the footnote under the Chi-Square Tests table. On this data it says that 6 cells (100.0%) have expected count less than 5, and it is right: the six expected counts are \(3.0\), \(4.5\), \(2.5\), \(3.0\), \(4.5\), \(2.5\). The Pearson \(\chi^{2}\) is \(1.867\) on 2 df with Sig. \(= 0.393\), and that significance cannot be quoted as it stands. Pool Med and High, or read Fisher's Exact Test, which SPSS prints in the same table for a \(2\times2\) and offers under /METHOD=EXACT otherwise.

SPSS's /RUNS defaults to cutting at the median, not at zero. To run the test on the signs of regression residuals — the use it is really for — save the residuals from REGRESSION, then use /RUNS(0).

Corrected. The Wilcoxon \(p\) quoted from STS-108 was 0.0096, the value from the untied variance. With the variance corrected for ties, as SPSS and R both compute it, it is 0.0086. The conclusion is unchanged.
RESULT

For gender by grade, \(\chi^{2} = 1.867\) on 2 df, Sig. \(= 0.393\), but every expected count is below 5, so that significance is not quoted. Pooling Med and High still leaves expected counts of 3, so Fisher's exact test is the one to use: it gives \(p = 0.523\) (rechecked in R). There is no evidence that grade depends on gender.

Practical 6: Design and Analysis of Experiments

1. Question

Write the UNIANOVA syntax for a response \(y\) in a completely randomised design, a randomised block design, a Latin square, and the \(2^{2}\) and \(2^{3}\) factorials.

2. Aim

Express each design through the /DESIGN subcommand of one procedure, and know what leaving a term out does.

3. Steps

  1. Analyze → General Linear Model → Univariate: \(y\) as the dependent variable, the treatment and blocking factors as fixed factors.
  2. Under Model choose Build terms and enter exactly the terms of the design, from the table below.
  3. Ask for the treatment means compared with a Bonferroni adjustment, the effect sizes and the homogeneity test.
  4. Paste, and check the /DESIGN line before running.
MENUS, COMMANDS AND CONVENTIONS

Every design named in this course's list is analysed in Design and Analysis of Experiments and in STS-203, with the sums of squares worked by hand. In SPSS all five are UNIANOVA, and the design is expressed entirely in the /DESIGN subcommand.

Design/DESIGNWhat it says
Completely randomised/DESIGN=treat one factor, no blocking
Randomised block/DESIGN=block treat main effects only — blocks are assumed not to interact
Latin square/DESIGN=row col treat two blocking factors, no interaction
\(2^{2}\) factorial/DESIGN=A B A*B both main effects and the interaction
\(2^{3}\) factorial/DESIGN=A B C A*B A*C B*C A*B*C the full model — seven effects on seven degrees of freedom

4. Programme

PRACTICAL 6 — SPSS SYNTAX
UNIANOVA y BY block treat
  /METHOD=SSTYPE(3) /INTERCEPT=INCLUDE
  /EMMEANS=TABLES(treat) COMPARE ADJ(BONFERRONI)
  /PRINT=ETASQ HOMOGENEITY DESCRIPTIVE
  /DESIGN=block treat.

UNIANOVA y BY A B C
  /METHOD=SSTYPE(3)
  /DESIGN=A B C A*B A*C B*C A*B*C.

5. Execution and Results

The Tests of Between-Subjects Effects table has one row for each term named in /DESIGN, and an Error row that holds everything left out. The sums of squares for each design are worked by hand on the pages linked above, and the SPSS table can be checked against them.

Leaving an interaction out of /DESIGN does not delete it — it pools it into the error. That is exactly right for a randomised block design, where the block-by-treatment interaction is the error term, and exactly wrong for a factorial, where it is the effect of interest. The default /DESIGN that the menus write includes every interaction; the randomised block analysis requires you to take them out deliberately.

Type III sums of squares are SPSS's default and are the right choice for a balanced design, where all four types coincide. For an unbalanced design they do not, and the difference is the subject of STS-203, Unit 1, where the naive sums of squares are shown producing an impossible negative interaction. Say which type you used.

RESULT

One procedure, five /DESIGN lines. In the randomised block design and the Latin square the interactions are left out on purpose and become the error; in the factorials they are the effects being tested and must be listed. Type III sums of squares, SPSS's default, are right for these balanced designs.

Practical 7: Regression Analysis

1. Question

(a) Regress score on age, income and visits, and choose the model by all possible subsets and by forward, backward and stepwise selection. (b) Model the probability that buy is Yes from income by logistic and by probit regression, and compare the two fits.

2. Aim

Select a regression model by five methods and show that they agree; then fit the two binary-response models and say why one is usually preferred.

3. Steps

  1. (a) Analyze → Regression → Linear: score as dependent, the three predictors entered together, with the descriptives, the collinearity diagnostics and the \(R^{2}\) change.
  2. (a) Fit every subset of the three predictors, and tabulate \(R^{2}\), adjusted \(R^{2}\), Mallows's \(C_p\), AIC and BIC for each.
  3. (a) Run the regression again with Method = Forward, Backward and Stepwise, stating the entry and removal thresholds.
  4. (b) Analyze → Regression → Binary Logistic: buy as dependent, income as covariate.
  5. (b) The probit model on the same 0/1 data, through Generalized Linear Models with a binomial distribution and a probit link.
  6. (b) Compare the coefficients, the likelihoods, the fitted probabilities and the classification.
MENUS, COMMANDS AND CONVENTIONS

Both model \(P(\text{buy} = \text{Yes})\) as a monotone function of income, and differ only in which function:

\[ \text{logit}: \; P = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x)}}, \qquad \text{probit}: \; P = \Phi\!\left(\beta_0 + \beta_1 x\right). \]

Multinomial logistic (NOMREG) extends the same model to an outcome with more than two unordered categories, fitting \(k-1\) logits against a reference category; PLUM is the ordinal version, which fits one slope and \(k-1\) thresholds and is the right choice for grade.

4. Programme

PRACTICAL 7 — SPSS SYNTAX
REGRESSION /DESCRIPTIVES MEAN STDDEV CORR
  /STATISTICS COEFF OUTS R ANOVA COLLIN TOL CHANGE
  /DEPENDENT score
  /METHOD=ENTER age income visits
  /RESIDUALS DURBIN
  /SAVE RESID.
REGRESSION /STATISTICS COEFF OUTS R ANOVA CHANGE
  /DEPENDENT score
  /METHOD=FORWARD age income visits.       /* PIN = .05 by default   */

REGRESSION /DEPENDENT score
  /METHOD=BACKWARD age income visits.      /* POUT = .10 by default  */

REGRESSION /DEPENDENT score
  /METHOD=STEPWISE age income visits.      /* both, in alternation   */

* the two thresholds are settable, and the values used must be reported
REGRESSION /CRITERIA=PIN(.05) POUT(.10) /DEPENDENT score
  /METHOD=STEPWISE age income visits.
LOGISTIC REGRESSION VARIABLES buy
  /METHOD=ENTER income
  /PRINT=GOODFIT CI(95)
  /CRITERIA=PIN(0.05) POUT(0.10) ITERATE(20) CUT(0.5).

PROBIT r OF n WITH income
  /MODEL=PROBIT /PRINT=ALL.
* PROBIT wants grouped data -- r successes out of n trials at each level of the
* covariate.  For ungrouped 0/1 data use GENLIN with a probit link instead:
GENLIN buy (REFERENCE=FIRST) WITH income
  /MODEL income DISTRIBUTION=BINOMIAL LINK=PROBIT
  /CRITERIA METHOD=FISHER(1) SCALE=1
  /PRINT SOLUTION FIT.
NOMREG outcome (BASE=FIRST ORDER=ASCENDING) WITH income age
  /MODEL /PRINT=CLASSTABLE FIT PARAMETER LRT.
PLUM grade WITH income /LINK=LOGIT /PRINT=FIT PARAMETER TPARALLEL.
* TPARALLEL tests the proportional-odds assumption PLUM rests on.
* If it is rejected, PLUM is the wrong model and NOMREG is the fallback.

5. Execution and Results

(a) Model selection

This is the part of the course that is not taught elsewhere on the site, so it is worked in full. Regress score on age, income and visits. The three predictors are correlated with the outcome at \(0.5082\), \(0.6536\) and \(0.0535\), and age and income are correlated with each other at \(0.7795\).

All possible subsets — with three predictors there are eight models, and every criterion can be computed for each:

ALL POSSIBLE SUBSETS (RECOMPUTED IN R)
Model\(p\)\(R^{2}\)Adjusted \(R^{2}\) Mallows \(C_p\)AICBIC
intercept only10.00000.000010.109 90.75691.751
age20.25830.21714.84986.780 88.771
income20.42720.3954 0.10081.610 83.602
visits20.0029−0.052512.029 92.69894.690
age + income30.42720.35982.100 83.61086.597
age + visits30.27330.18786.427 88.37091.358
income + visits30.43070.36372.003 83.48986.476
age + income + visits40.43080.32414.000 85.48589.468

Read the two \(R^{2}\) columns against each other. Plain \(R^{2}\) rises with every variable added — from \(0.4272\) to \(0.4307\) to \(0.4308\) — because it must; adding a column can never increase the residual sum of squares. Adjusted \(R^{2}\) falls, from \(0.3954\) to \(0.3637\) to \(0.3241\), because it charges for each one. The third decimal of \(R^{2}\) is not evidence of anything.

Mallows's \(C_p\) is read against \(p\), not against zero: a model with negligible bias has \(C_p \approx p\). Here income alone has \(C_p = 0.100\) against \(p = 2\), the full model has \(C_p = 4.000\) against \(p = 4\) — which it must, identically — and the intercept-only model has \(10.109\) against \(1\), which is the signal that it is badly biased.

All four criteria choose the same model. Highest adjusted \(R^{2}\), lowest \(C_p\) relative to \(p\), lowest AIC and lowest BIC all point at income alone.

FORWARD AND BACKWARD SELECTION (RECOMPUTED IN R)
MethodStepStatisticSig.Action
Forwardenter income \(F = 13.4262\) on \((1,18)\)0.001775accepted
try visits\(F = 0.1034\) on \((1,17)\)0.751683 rejected, \(p > 0.05\) — stop
Backwardfull model, smallest \(|t|\) is age \(t = 0.0541\)0.957561removed
income + visits, smallest is visits \(t = 0.3216\)0.751683removed
income alone\(t = 3.6642\)0.001775 kept, \(p < 0.10\) — stop

Stepwise alternates the two and reaches the same place. The final model is

\[ \widehat{\text{score}} = 40.147655 + 0.480802\,\text{income}, \]

with standard errors \(5.857109\) and \(0.131217\), so \(t = 6.854517\) and \(3.664175\).

Why age is dropped although it correlates with score at \(0.5082\). In the model age + income its coefficient is \(-0.002852\) with a standard error of \(0.264487\), so \(t = -0.0108\) and the sign has even flipped. The reason is the correlation of \(0.7795\) between age and income: once income is in the model, age has nothing left to explain. A predictor's simple correlation with the outcome tells you nothing about whether it belongs in a model that already contains its neighbours, and this is the whole reason selection procedures exist.

Two warnings that belong in the report. Stepwise selection with three candidates on twenty cases is already at the edge: the rule of thumb is ten or more cases per candidate predictor. And the \(p\) values printed for the final model are not honest, because the model was chosen by looking at them; they are conditional on a selection that the same data performed.

(b) Logistic and probit

LOGIT AGAINST PROBIT (RECOMPUTED IN R)
QuantityLogitProbit
Constant−11.401132−6.462962
income0.2501240.142527
\(-2\) Log likelihood11.641511.6193
\(-2\) Log likelihood, null model26.920526.9205
Likelihood-ratio \(\chi^{2}\), 1 df15.279015.3012
Sig.0.0000930.000092
McFadden \(R^{2}\)0.5675590.568385
Income at which \(P = 0.5\)45.58201045.345683
Classification accuracy at cut 0.50.90000.9000

The two fits are all but identical, and that is the normal situation: the logistic and normal distribution functions differ only in the tails. The coefficients differ by a scale factor and nothing else — \(0.250124/0.142527 = 1.755\), and the textbook rule of thumb is \(\beta_{\text{logit}} \approx 1.6\) to \(1.8\) times \(\beta_{\text{probit}}\). The fitted probabilities are so close that both classify all twenty cases identically.

So why choose one? Because \(\exp(\beta_1)\) is an odds ratio under the logit and means nothing under the probit. Here \(e^{0.250124} = 1.284184\): each extra thousand of income multiplies the odds of buying by \(1.28\). Nothing in the probit output has an interpretation that simple, which is why logistic regression is the default everywhere and probit survives mainly in bioassay, where the normal tolerance distribution has a physical meaning.

The classification table, and the five ratios read off it, are worked in STS-108, Practical 6 — accuracy \(0.900\) against a no-information rate of \(0.600\), precision and recall both \(0.875\), specificity \(0.9167\), AUC \(0.921875\). SPSS prints the classification table automatically and the ROC curve under Analyze → ROC Curve.

Checked. Every figure in the three tables was recomputed in R 4.3.3 and agrees. The probit figures need the convergence tolerance tightened, with glm.control(epsilon = 1e-14): R's default stops one iteration early and differs in the sixth decimal.
RESULT

(a) Every method chooses \(\widehat{\text{score}} = 40.15 + 0.481\,\text{income}\) (\(R^{2} = 0.427\), adjusted 0.395). age adds nothing once income is in, because the two are correlated at 0.78. (b) The logit and probit fits are almost identical: \(-2\) log likelihood 11.64 and 11.62, and both classify 18 of the 20 cases correctly. The logit is preferred because \(e^{0.2501} = 1.28\) is an odds ratio: each extra thousand of income multiplies the odds of buying by 1.28.

Practical 8: Multivariate Data Analysis

1. Question

Write the SPSS syntax for a linear discriminant analysis of buy on age, income and visits; a principal components analysis, with varimax rotation, of age, income, visits and score; and a hierarchical and a \(k\)-means cluster analysis of age, income and visits.

2. Aim

Know where SPSS keeps each multivariate method, and which option changes its answer.

3. Steps

  1. Analyze → Classify → Discriminant: buy as the grouping variable (0, 1), the three predictors entered together, prior probabilities from the group sizes, and Box's \(M\).
  2. Analyze → Dimension Reduction → Factor: principal components from the correlation matrix, eigenvalues over 1 kept, varimax rotation.
  3. Analyze → Classify → Hierarchical Cluster (average linkage between groups, squared Euclidean distance, the dendrogram), then K-Means Cluster with three clusters.
  4. For each method, decide and record the option in the last column of the table below.
MENUS, COMMANDS AND CONVENTIONS

Discriminant analysis, principal components, factor analysis, multidimensional scaling and cluster analysis are built from the definitions, with every eigenvalue computed, in Multivariate Analysis (STS-202), Unit 3 and Unit 4. Nothing about the methods is repeated here — only where SPSS keeps them and which options change the answer.

MethodMenuSyntaxThe option that changes the answer
Linear discriminantClassify → Discriminant DISCRIMINANT prior probabilities: equal, or proportional to group size
Principal componentsDimension Reduction → Factor FACTOR /EXTRACTION=PC correlation matrix (default) or covariance matrix — different components
Factor analysisthe same dialog FACTOR /EXTRACTION=PAF or ML the rotation, and whether it is orthogonal or oblique
Multidimensional scalingScale → Multidimensional Scaling ALSCAL or PROXSCAL the measurement level of the proximities: ordinal or ratio
Cluster analysisClassify → Hierarchical or K-Means CLUSTER, QUICK CLUSTER the linkage, and whether the variables were standardized first

4. Programme

PRACTICAL 8 — SPSS SYNTAX
DISCRIMINANT GROUPS=buy(0 1) /VARIABLES=age income visits
  /METHOD=DIRECT /PRIORS=SIZE /STATISTICS=MEAN STDDEV BOXM TABLE.

FACTOR /VARIABLES age income visits score
  /MISSING LISTWISE /ANALYSIS age income visits score
  /PRINT INITIAL KMO EXTRACTION ROTATION
  /CRITERIA MINEIGEN(1) ITERATE(25)
  /EXTRACTION PC /ROTATION VARIMAX /METHOD=CORRELATION.

CLUSTER age income visits /METHOD BAVERAGE /MEASURE=SEUCLID
  /PLOT DENDROGRAM /PRINT SCHEDULE.
QUICK CLUSTER age income visits /CRITERIA=CLUSTER(3) MXITER(20) /PRINT ANOVA.

5. Execution and Results

The procedures run on this data and print their tables: the Classification Results for DISCRIMINANT, Total Variance Explained and the Rotated Component Matrix for FACTOR, the Agglomeration Schedule and the dendrogram for CLUSTER, and the Final Cluster Centers for QUICK CLUSTER. They are not reproduced here, since SPSS is not run on this site, and on twenty cases they should not be interpreted.

Three warnings, all of which SPSS will let you walk past.

RESULT

The syntax for all four analyses, with the option that changes each answer named. The output is shown in the record but reported as illustrative only: twenty cases cannot support any of these methods, and saying so is part of the result.

Practical 9: Statistical Quality Control

1. Question

Write the SPSS syntax for an \(\bar X\) and R chart, with \(C_p\) and \(C_{pk}\), for measurements taken in rational subgroups; and for a \(p\) chart of the defectives found in samples of varying size, with the run rules applied to both.

2. Aim

Draw the control charts from SPCHART, and read them with the run rules.

3. Steps

  1. Choose the chart from the kind of data, by the table below.
  2. Analyze → Quality Control → Control Charts, the chart type, and the variable that identifies the subgroups.
  3. Under Control Rules tick all the rules; under Statistics give the specification limits, which \(C_p\) and \(C_{pk}\) need.
  4. Read the chart for the flagged rule violations first, and only then for capability.
MENUS, COMMANDS AND CONVENTIONS

The control limits, the constants \(A_2\), \(D_3\), \(D_4\), \(B_3\), \(B_4\), \(d_2\), and the choice between charts are all in Statistical Quality Control — Unit 2 for the variable charts and process capability, Unit 3 for the attribute charts. SPSS builds all of them from SPCHART.

ChartDataSPCHART subcommand
\(\bar X\) and Rvariables, subgroup size below 10 /XR=
\(\bar X\) and svariables, subgroup size 10 or more /XS=
Individuals and moving rangevariables, subgroup size 1 /IR=
\(p\)fraction defective, subgroup size may vary/P=
\(np\)number defective, constant subgroup size/NP=
\(c\)number of defects, constant area of opportunity /C=
\(u\)defects per unit, varying area/U=

4. Programme

PRACTICAL 9 — SPSS SYNTAX
SPCHART /XR=measure BY subgroup
  /STATISTICS=CP CPK /RULES=ALL.

SPCHART /P=defectives BY sample (inspected)
  /SIGMAS=3 /RULES=ALL.

5. Execution and Results

SPSS draws each chart with its centre line and its \(3\sigma\) limits, and lists under it any points that break a rule.

/RULES=ALL applies the Western Electric run rules — a point beyond \(3\sigma\), two of three beyond \(2\sigma\), four of five beyond \(1\sigma\), eight in a row on one side — and flags them on the chart. A chart read only for points outside the limits misses every slow drift, which is the failure mode these rules exist to catch.

\(C_p\) and \(C_{pk}\) require the specification limits, which are an engineering input and not in the data. /STATISTICS=CP CPK without them produces nothing. And process capability means nothing until the process is in control: computing \(C_{pk}\) from a chart with points outside the limits describes a process that has no single standard deviation to speak of.

RESULT

The chart is chosen by the data, not by habit: \(\bar X\) and R for small subgroups of measurements, a \(p\) chart for defectives in samples of varying size. A process is judged by the run rules as well as the limits, and its capability is quoted only once it is in control.

How Marks Are Lost

THE RECURRING ERRORS

What the Practical Record Should Contain

FOR EACH ANALYSIS
  1. Question — the problem in the terms it was set, and the variables used with their SPSS Measure settings.
  2. Aim — in one line.
  3. Steps — the menu path; the Variable View definitions used (labels, value labels, missing values); and the assumptions the procedure requires.
  4. Programme — the pasted syntax. Give both the menu path and the syntax, because the examination may ask for either.
  5. Execution and Results — the output tables by name, with the statistic, its degrees of freedom and its Sig. value; the output that checks each assumption (Levene, Shapiro–Wilk, Box's \(M\), the expected-count footnote, the collinearity diagnostics); the conclusion in words, in the terms of the original question; and what the analysis does not establish — on twenty cases that is usually the longer list.