UNIANOVA uses by default,
which correlation FACTOR starts from, which \(F\) REGRESSION uses to
enter a variable — it is named.
The data set is the one used in STS-108, Data Handling using R, with one column added. Running both papers on the same twenty records is deliberate: the output of the two packages can then be compared line for line, and a difference is either a convention or a mistake.
| Variable | Measure in SPSS | Values |
|---|---|---|
id | Nominal | 1 to 20 |
age | Scale | 23 45 31 58 27 39 52 34 29 48 36 41 25 55 33 44 30 50 28 38 |
income | Scale | 40 57 37 69 36 38 63 32 35 57 42 49 26 39 48 49 23 57 27 33 |
visits | Scale | 7 5 8 10 6 5 7 9 3 5 6 7 4 2 6 5 10 5 9 5 |
score | Scale | 54 55 64 77 69 60 85 57 65 59 64 60 51 57 60 60 55 67 40 56 |
gender | Nominal | M F alternating, 10 each |
buy | Nominal (0 = No, 1 = Yes) | 0 1 0 1 1 0 1 0 0 1 0 0 0 0 1 1 0 1 0 0 |
grade is derived from score as Low below 57, Med from 57 to 64,
High from 65 — and is Ordinal.
Read the twenty cases of the data set into SPSS from customers.csv; define every
variable in Variable View — its label, value labels, missing-value code and Measure; derive the
ordinal grade from score; check the coding with frequency tables; and save
the file as .sav and as an Excel workbook.
Set up a data file that every later procedure can trust: the right Measure for each variable, labelled codes and declared missing values.
GET DATA) reads the file; name the data set.grade from score, each class closed on the left; label it, set it Ordinal, and run EXECUTE.grade, gender and buy checks the coding against the data..sav; SAVE TRANSLATE (File → Export) writes the Excel copy.| Window | Holds | Saved as |
|---|---|---|
| Data Editor — Data View | the cases, one per row | .sav |
| Data Editor — Variable View | the definitions: name, type, width, decimals, label, value labels, missing, measure | the same .sav |
| Output Viewer | every table and chart produced | .spv |
| Syntax Editor | the commands | .sps |
The Measure column is not cosmetic. SPSS uses it to decide which procedures a variable may enter, which charts the Chart Builder will offer, and how a variable is treated inside a model. Its three values map onto the four measurement scales:
| SPSS Measure | Scale of measurement | Consequence |
|---|---|---|
| Nominal | nominal | counts and \(\chi^{2}\) only; enters models as a set of dummy variables |
| Ordinal | ordinal | median and rank methods; ordinal regression accepts it as the outcome |
| Scale | interval or ratio | means, standard deviations, correlations, everything |
SPSS does not distinguish interval from ratio, so the one restriction it cannot enforce for you is the one on ratios: a score of 80 is not “twice” a score of 40 unless the zero is real. That judgement stays with the analyst.
Value labels are entered once and used everywhere. Coding
buy as 0 and 1 and labelling them No and Yes means every later table reads
“No” and “Yes” rather than 0 and 1 — and the coding, not the
label, is what the model uses. In a binary logistic regression SPSS models the
higher code, so labelling them the wrong way round reverses the sign of every
coefficient.
Every menu path ends in a Paste button, which writes the command into the Syntax Editor instead of running it. Use it always. Menus leave no record of what was done; syntax is the record, it can be re-run, and it is what a practical record must contain.
* Read a CSV file. SPSS syntax comments begin with * and end with a full stop.
GET DATA /TYPE=TXT /FILE='C:\data\customers.csv'
/DELIMITERS=',' /FIRSTCASE=2 /VARIABLES=
id F3.0 age F3.0 income F4.0 visits F3.0 score F4.0 gender A1 buy F1.0.
DATASET NAME customers WINDOW=FRONT.
* Define the variables properly -- this is the step most often skipped.
VARIABLE LABELS income 'Monthly income, thousands' score 'Assessment score'.
VALUE LABELS buy 0 'No' 1 'Yes'.
VARIABLE LEVEL id gender buy (NOMINAL) / age income visits score (SCALE).
MISSING VALUES income (-99).
* Derive the ordinal grade, closing each class on the LEFT.
RECODE score (LO THRU 56.999=1) (57 THRU 64.999=2) (65 THRU HI=3) INTO grade.
VALUE LABELS grade 1 'Low' 2 'Med' 3 'High'.
VARIABLE LEVEL grade (ORDINAL).
EXECUTE.
FREQUENCIES VARIABLES=grade gender buy.
* grade: Low 6, Med 9, High 5. gender: F 10, M 10. buy: No 12, Yes 8.
* Export.
SAVE OUTFILE='C:\data\customers.sav' /COMPRESSED.
SAVE TRANSLATE OUTFILE='C:\data\customers.xlsx' /TYPE=XLS /VERSION=12
/FIELDNAMES /CELLS=VALUES.
| Variable | Category | Frequency |
|---|---|---|
grade | Low | 6 |
| Med | 9 | |
| High | 5 | |
gender | F | 10 |
| M | 10 | |
buy | No | 12 |
| Yes | 8 |
Three file-handling traps. EXECUTE is needed after
transformations such as RECODE and COMPUTE before the new values
exist — SPSS defers them until a procedure reads the data. /CELLS=VALUES
exports the codes and /CELLS=LABELS exports the labels; choosing the second and
then re-importing turns a numeric variable into a string. And a declared missing value
(MISSING VALUES income (-99)) is excluded from analysis, while an undeclared
\(-99\) is treated as income of minus ninety-nine thousand and will quietly ruin every
mean.
The file holds the twenty cases with every variable defined. The frequency tables show
grade Low 6, Med 9, High 5; gender 10 and 10; buy No 12 and
Yes 8, which agree with the data as entered, so the coding and the recode are right. The
.sav keeps the definitions; the Excel copy keeps only the values.
For the data set, draw a pie chart of grade, a bar chart of gender, a
histogram of score with a normal curve, a scatter plot of score against
income, boxplots of score by gender, and an ogive of
score.
Produce each prescribed chart from syntax, so that every one can be drawn again exactly.
score, with Display normal curve ticked.score on the vertical axis and income on the horizontal.score by gender.score drawn as a line.The theory of which diagram suits which variable is in STS-108, section 4 and is not repeated. SPSS offers three routes to the same charts, and the syntax route is the only reproducible one.
| Chart | Menu | Syntax command |
|---|---|---|
| Pie | Graphs → Legacy Dialogs → Pie | GRAPH /PIE |
| Bar | Graphs → Legacy Dialogs → Bar | GRAPH /BAR |
| Histogram | Graphs → Legacy Dialogs → Histogram | GRAPH /HISTOGRAM |
| Line | Graphs → Legacy Dialogs → Line | GRAPH /LINE |
| Scatter | Graphs → Legacy Dialogs → Scatter/Dot | GRAPH /SCATTERPLOT |
| Boxplot | Graphs → Legacy Dialogs → Boxplot | EXAMINE ... /PLOT BOXPLOT |
| Frequency polygon, ogive | Chart Builder, or a line over binned data | GGRAPH |
| Gantt | not a built-in chart | a clustered bar of durations with a transparent first segment |
GRAPH /PIE=COUNT BY grade.
GRAPH /BAR(SIMPLE)=COUNT BY gender.
GRAPH /HISTOGRAM(NORMAL)=score.
GRAPH /SCATTERPLOT(BIVAR)=income WITH score /MISSING=LISTWISE.
EXAMINE VARIABLES=score BY gender /PLOT=BOXPLOT /STATISTICS=NONE /NOTOTAL.
* the ogive, which SPSS has no button for: cumulative percentage against the class
FREQUENCIES VARIABLES=score /FORMAT=NOTABLE /NTILES=10.
GRAPH /LINE(SIMPLE)=CUM(COUNT) BY score.
SPSS draws each chart in the Output Viewer. The charts are not reproduced here, because SPSS
is not run on this site, but the figures they are drawn from can be checked: the normal curve on the
histogram has the mean 60.75 and standard deviation 9.4361 of score
(Practical 3), the pie has the slices Low 6, Med 9 and High 5 (Practical 1), and the two bars are 10
and 10.
GRAPH /HISTOGRAM(NORMAL) superimposes a normal curve with the
sample's own mean and standard deviation. It is a picture, not a test: the Shapiro–Wilk
statistic in EXAMINE ... /PLOT NPPLOT is the test, and on a sample of twenty it
has very little power, so a histogram that “looks normal” and a
Shapiro–Wilk that does not reject are both weak evidence.
Six charts, each from one line of syntax that redraws it exactly. The scatter plot shows the
upward trend that the correlation of 0.6536 between income and score
measures (Practical 7). The histogram with its normal curve is a picture of the shape of
score, not a test of normality.
Find the mean, median, standard deviation, minimum, maximum and skewness of age,
income and score; the median and mode of grade; and the
summaries of score for each gender.
Choose the right one of SPSS's three descriptive procedures for each summary, and know which conventions its output uses.
grade, with the median and the mode under Statistics: it is the only one of the three that gives the mode.score by gender, with boxplots and normality plots.| Procedure | Gives | Use for |
|---|---|---|
FREQUENCIES |
the frequency table, plus any statistic you ask for, including the mode and percentiles | categorical variables; and it is the only one that gives the mode |
DESCRIPTIVES |
mean, standard deviation, minimum, maximum, skewness, kurtosis — no median | a compact table of scale variables; and \(z\) scores, via
/SAVE |
EXAMINE |
everything, by group: the five-number summary, trimmed mean, confidence interval, normality tests, boxplots and stem-and-leaf | anything that has to be broken down by a factor |
DESCRIPTIVES VARIABLES=age income score
/STATISTICS=MEAN STDDEV MIN MAX SKEWNESS KURTOSIS.
FREQUENCIES VARIABLES=grade /STATISTICS=MEDIAN MODE.
EXAMINE VARIABLES=score BY gender /PLOT BOXPLOT NPPLOT /STATISTICS DESCRIPTIVES.
| Variable | Mean | Median | Std. Deviation | Minimum | Maximum | Skewness |
|---|---|---|---|---|---|---|
| age | 38.3000 | 37.0000 | 10.4534 | 23 | 58 | 0.3507 |
| income | 42.8500 | 39.5000 | 12.8279 | 23 | 69 | 0.3874 |
| score | 60.7500 | 60.0000 | 9.4361 | 40 | 85 | 0.5613 |
Two conventions to state in the report. SPSS's Std. Deviation divides by \(n-1\); and the skewness SPSS prints is not the moment coefficient \(g_1 = m_3/m_2^{3/2}\) but the sample-adjusted version
\[ G_1 = \frac{n}{(n-1)(n-2)}\sum\left(\frac{x-\bar x}{s}\right)^{3} = g_1\,\frac{\sqrt{n(n-1)}}{n-2}, \]so at \(n = 20\) it is larger by a factor \(\sqrt{380}/18 = 1.082977\), or \(8.3\%\). The table above gives the moment coefficients, so that they match the R and Python pages; SPSS will print \(0.379812\), \(0.419544\) and \(0.607924\) for the three variables, and the difference is the convention and not an error.
Mean score 60.75 with standard deviation 9.44; age 38.3 (10.45); income 42.85 (12.83). The
medians come from FREQUENCIES or EXAMINE, since
DESCRIPTIVES gives none. All three variables are mildly skewed to the right: moment
skewness 0.35, 0.39 and 0.56, which SPSS prints as 0.380, 0.420 and 0.608. The median of
grade is Med (the 10th and 11th of the twenty ordered grades are both Med) and so is
the mode, with 9 cases.
Test whether the mean score differs between men and women, and run a one-way
analysis of variance of score by grade. Say what each result means.
Run the two-sample \(t\) test in the right order, the variances first, and recognise an analysis that SPSS will run but that answers nothing.
score as the test variable, gender as the grouping variable with the groups M and F.score by grade, with the descriptives, the homogeneity test and Tukey's and the LSD post hoc tests.The theory is in Inferential Statistics and Estimation Theory (STS-201).
| Test | Menu | Syntax |
|---|---|---|
| One mean | Analyze → Compare Means → One-Sample T Test | T-TEST /TESTVAL= |
| Two independent means | Analyze → Compare Means → Independent-Samples T Test | T-TEST GROUPS= |
| Paired means | Analyze → Compare Means → Paired-Samples T Test | T-TEST PAIRS= |
| One proportion | Analyze → Nonparametric Tests → Binomial | NPAR TESTS /BINOMIAL |
| One-way ANOVA | Analyze → Compare Means → One-Way ANOVA | ONEWAY |
| Two-way and factorial ANOVA | Analyze → General Linear Model → Univariate | UNIANOVA |
The independent-samples table has two rows, and choosing between them is the examinable step. SPSS prints Levene's test for equality of variances first, then a \(t\) assuming equal variances and a \(t\) not assuming it (Welch, with fractional degrees of freedom). Read Levene, then read the row it points to.
T-TEST GROUPS=gender('M' 'F') /VARIABLES=score /CRITERIA=CI(.95).
ONEWAY score BY grade /STATISTICS DESCRIPTIVES HOMOGENEITY /POSTHOC=TUKEY LSD ALPHA(.05).
| Output line | Value |
|---|---|
| Mean, gender = M | 60.7000 |
| Mean, gender = F | 60.8000 |
| Variance, M | 144.4556 |
| Variance, F | 43.5111 |
| Levene's Test for Equality of Variances, \(F\) | 2.1833 |
| Levene's Test, Sig. | 0.1568 |
| \(t\), equal variances assumed, df = 18 | −0.0231 |
| \(t\), equal variances not assumed, df = 13.971 | −0.0231 |
| Sig. (2-tailed), both rows | 0.9819 |
| Variance ratio \(s_M^{2}/s_F^{2}\) (not printed by SPSS) | 3.3200 |
The two means differ by \(0.1\) and the variances by a factor of \(3.32\). A \(t\) test of means answers a question this data does not raise; the difference here is entirely one of spread, which is why the variance line is read first and reported whatever it says.
The ANOVA of score by grade is on the page as a trap.
It gives \(F = 19.756\) on \((2, 17)\) with Sig. \(= 0.000\), and it is worthless:
grade was recoded out of score, so the groups differ by construction.
SPSS will print the table without complaint. A procedure that runs is not a result.
The means are 60.7 and 60.8: \(t = -0.023\) on 18 df, Sig. \(= 0.982\), so there is no
evidence that the mean score differs by gender. Levene's test (Sig. 0.157) does not reject equal
variances, although the men's variance is 3.3 times the women's. The ANOVA of score by
grade (\(F = 19.76\)) is not a finding, because grade is made from
score.
Run the sign and Wilcoxon signed-rank tests on a pair of before-and-after variables, the
Mann–Whitney test of score by gender, the runs test and the one-sample
Kolmogorov–Smirnov test on score, and the \(\chi^{2}\) test of independence of
gender and grade.
Find each prescribed non-parametric test in SPSS, and read the warnings its output prints.
score by gender.score.gender by grade, with Chi-square under Statistics and the expected counts under Cells.What each test uses of the data, and what it throws away, is worked in STS-108, Practical 8 — including the comparison where the sign test gives \(p = 0.146\) on data the Wilcoxon signed rank rejects at \(p = 0.0086\). Here: which dialog, and what the output is called.
| Test | Menu (Legacy Dialogs) | Syntax |
|---|---|---|
| Sign | 2 Related Samples → Sign | /SIGN= |
| Wilcoxon signed rank | 2 Related Samples → Wilcoxon | /WILCOXON= |
| Mann–Whitney \(U\) | 2 Independent Samples → Mann-Whitney U | /M-W= |
| Runs | Runs | /RUNS(MEDIAN)= |
| Kolmogorov–Smirnov, one sample | 1-Sample K-S | /K-S(NORMAL)= |
| \(\chi^{2}\) goodness of fit | Chi-Square | /CHISQUARE= |
| \(\chi^{2}\) independence | Descriptive Statistics → Crosstabs | CROSSTABS /STATISTICS=CHISQ |
NPAR TESTS /SIGN=before WITH after (PAIRED) /WILCOXON=before WITH after (PAIRED).
NPAR TESTS /M-W=score BY gender('M' 'F').
NPAR TESTS /RUNS(MEDIAN)=score.
NPAR TESTS /K-S(NORMAL)=score.
CROSSTABS TABLES=gender BY grade /STATISTICS=CHISQ /CELLS=COUNT EXPECTED ROW.
| gender | Low | Med | High | Total |
|---|---|---|---|---|
| F, count | 2 | 6 | 2 | 10 |
| F, expected | 3.0 | 4.5 | 2.5 | 10.0 |
| M, count | 4 | 3 | 3 | 10 |
| M, expected | 3.0 | 4.5 | 2.5 | 10.0 |
| Total | 6 | 9 | 5 | 20 |
Pearson Chi-Square \(1.867\), df 2, Asymptotic Sig. (2-sided) \(0.393\); footnote: 6 cells (100.0%) have expected count less than 5, the minimum expected count being 2.50.
Read the footnote under the Chi-Square Tests table. On this data it says
that 6 cells (100.0%) have expected count less than 5, and it is right: the six
expected counts are \(3.0\), \(4.5\), \(2.5\), \(3.0\), \(4.5\), \(2.5\). The Pearson
\(\chi^{2}\) is \(1.867\) on 2 df with Sig. \(= 0.393\), and that significance cannot
be quoted as it stands. Pool Med and High, or read Fisher's Exact Test, which SPSS
prints in the same table for a \(2\times2\) and offers under /METHOD=EXACT
otherwise.
SPSS's /RUNS defaults to cutting at the median, not at zero.
To run the test on the signs of regression residuals — the use it is really for —
save the residuals from REGRESSION, then use /RUNS(0).
For gender by grade, \(\chi^{2} = 1.867\) on 2 df, Sig.
\(= 0.393\), but every expected count is below 5, so that significance is not quoted. Pooling Med
and High still leaves expected counts of 3, so Fisher's exact test is the one to use: it gives
\(p = 0.523\) (rechecked in R). There is no evidence that grade depends on gender.
Write the UNIANOVA syntax for a response \(y\) in a completely randomised design,
a randomised block design, a Latin square, and the \(2^{2}\) and \(2^{3}\) factorials.
Express each design through the /DESIGN subcommand of one procedure, and know what leaving a term out does.
/DESIGN line before running.Every design named in this course's list is analysed in
Design
and Analysis of Experiments and in
STS-203, with the sums of
squares worked by hand. In SPSS all five are UNIANOVA, and the design is expressed
entirely in the /DESIGN subcommand.
| Design | /DESIGN | What it says |
|---|---|---|
| Completely randomised | /DESIGN=treat |
one factor, no blocking |
| Randomised block | /DESIGN=block treat |
main effects only — blocks are assumed not to interact |
| Latin square | /DESIGN=row col treat |
two blocking factors, no interaction |
| \(2^{2}\) factorial | /DESIGN=A B A*B |
both main effects and the interaction |
| \(2^{3}\) factorial | /DESIGN=A B C A*B A*C B*C A*B*C |
the full model — seven effects on seven degrees of freedom |
UNIANOVA y BY block treat
/METHOD=SSTYPE(3) /INTERCEPT=INCLUDE
/EMMEANS=TABLES(treat) COMPARE ADJ(BONFERRONI)
/PRINT=ETASQ HOMOGENEITY DESCRIPTIVE
/DESIGN=block treat.
UNIANOVA y BY A B C
/METHOD=SSTYPE(3)
/DESIGN=A B C A*B A*C B*C A*B*C.
The Tests of Between-Subjects Effects table has one row for each term named in
/DESIGN, and an Error row that holds everything left out. The sums of squares for each
design are worked by hand on the pages linked above, and the SPSS table can be checked against
them.
Leaving an interaction out of /DESIGN does not delete it — it
pools it into the error. That is exactly right for a randomised block design, where
the block-by-treatment interaction is the error term, and exactly wrong for a
factorial, where it is the effect of interest. The default /DESIGN that the menus
write includes every interaction; the randomised block analysis requires you to take them out
deliberately.
Type III sums of squares are SPSS's default and are the right choice for a balanced design, where all four types coincide. For an unbalanced design they do not, and the difference is the subject of STS-203, Unit 1, where the naive sums of squares are shown producing an impossible negative interaction. Say which type you used.
One procedure, five /DESIGN lines. In the randomised block design and the Latin
square the interactions are left out on purpose and become the error; in the factorials they are the
effects being tested and must be listed. Type III sums of squares, SPSS's default, are right for these
balanced designs.
(a) Regress score on age, income and visits,
and choose the model by all possible subsets and by forward, backward and stepwise selection.
(b) Model the probability that buy is Yes from income by logistic and by
probit regression, and compare the two fits.
Select a regression model by five methods and show that they agree; then fit the two binary-response models and say why one is usually preferred.
score as dependent, the three predictors entered together, with the descriptives, the collinearity diagnostics and the \(R^{2}\) change.buy as dependent, income as covariate.Both model \(P(\text{buy} = \text{Yes})\) as a monotone function of income, and differ only in which function:
\[ \text{logit}: \; P = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x)}}, \qquad \text{probit}: \; P = \Phi\!\left(\beta_0 + \beta_1 x\right). \]Multinomial logistic (NOMREG) extends the same model to an
outcome with more than two unordered categories, fitting \(k-1\) logits against a reference
category; PLUM is the ordinal version, which fits one slope and \(k-1\) thresholds
and is the right choice for grade.
REGRESSION /DESCRIPTIVES MEAN STDDEV CORR
/STATISTICS COEFF OUTS R ANOVA COLLIN TOL CHANGE
/DEPENDENT score
/METHOD=ENTER age income visits
/RESIDUALS DURBIN
/SAVE RESID.
REGRESSION /STATISTICS COEFF OUTS R ANOVA CHANGE
/DEPENDENT score
/METHOD=FORWARD age income visits. /* PIN = .05 by default */
REGRESSION /DEPENDENT score
/METHOD=BACKWARD age income visits. /* POUT = .10 by default */
REGRESSION /DEPENDENT score
/METHOD=STEPWISE age income visits. /* both, in alternation */
* the two thresholds are settable, and the values used must be reported
REGRESSION /CRITERIA=PIN(.05) POUT(.10) /DEPENDENT score
/METHOD=STEPWISE age income visits.
LOGISTIC REGRESSION VARIABLES buy
/METHOD=ENTER income
/PRINT=GOODFIT CI(95)
/CRITERIA=PIN(0.05) POUT(0.10) ITERATE(20) CUT(0.5).
PROBIT r OF n WITH income
/MODEL=PROBIT /PRINT=ALL.
* PROBIT wants grouped data -- r successes out of n trials at each level of the
* covariate. For ungrouped 0/1 data use GENLIN with a probit link instead:
GENLIN buy (REFERENCE=FIRST) WITH income
/MODEL income DISTRIBUTION=BINOMIAL LINK=PROBIT
/CRITERIA METHOD=FISHER(1) SCALE=1
/PRINT SOLUTION FIT.
NOMREG outcome (BASE=FIRST ORDER=ASCENDING) WITH income age
/MODEL /PRINT=CLASSTABLE FIT PARAMETER LRT.
PLUM grade WITH income /LINK=LOGIT /PRINT=FIT PARAMETER TPARALLEL.
* TPARALLEL tests the proportional-odds assumption PLUM rests on.
* If it is rejected, PLUM is the wrong model and NOMREG is the fallback.
This is the part of the course that is not taught elsewhere on the site, so it is
worked in full. Regress score on age, income and
visits. The three predictors are correlated with the outcome at
\(0.5082\), \(0.6536\) and \(0.0535\), and age and income are
correlated with each other at \(0.7795\).
All possible subsets — with three predictors there are eight models, and every criterion can be computed for each:
| Model | \(p\) | \(R^{2}\) | Adjusted \(R^{2}\) | Mallows \(C_p\) | AIC | BIC |
|---|---|---|---|---|---|---|
| intercept only | 1 | 0.0000 | 0.0000 | 10.109 | 90.756 | 91.751 |
| age | 2 | 0.2583 | 0.2171 | 4.849 | 86.780 | 88.771 |
| income | 2 | 0.4272 | 0.3954 | 0.100 | 81.610 | 83.602 |
| visits | 2 | 0.0029 | −0.0525 | 12.029 | 92.698 | 94.690 |
| age + income | 3 | 0.4272 | 0.3598 | 2.100 | 83.610 | 86.597 |
| age + visits | 3 | 0.2733 | 0.1878 | 6.427 | 88.370 | 91.358 |
| income + visits | 3 | 0.4307 | 0.3637 | 2.003 | 83.489 | 86.476 |
| age + income + visits | 4 | 0.4308 | 0.3241 | 4.000 | 85.485 | 89.468 |
Read the two \(R^{2}\) columns against each other. Plain \(R^{2}\) rises with every variable added — from \(0.4272\) to \(0.4307\) to \(0.4308\) — because it must; adding a column can never increase the residual sum of squares. Adjusted \(R^{2}\) falls, from \(0.3954\) to \(0.3637\) to \(0.3241\), because it charges for each one. The third decimal of \(R^{2}\) is not evidence of anything.
Mallows's \(C_p\) is read against \(p\), not against zero: a model with
negligible bias has \(C_p \approx p\). Here income alone has \(C_p = 0.100\)
against \(p = 2\), the full model has \(C_p = 4.000\) against \(p = 4\) — which it must,
identically — and the intercept-only model has \(10.109\) against \(1\), which is the
signal that it is badly biased.
All four criteria choose the same model. Highest adjusted \(R^{2}\),
lowest \(C_p\) relative to \(p\), lowest AIC and lowest BIC all point at income
alone.
| Method | Step | Statistic | Sig. | Action |
|---|---|---|---|---|
| Forward | enter income |
\(F = 13.4262\) on \((1,18)\) | 0.001775 | accepted |
try visits | \(F = 0.1034\) on \((1,17)\) | 0.751683 | rejected, \(p > 0.05\) — stop | |
| Backward | full model, smallest \(|t|\) is age |
\(t = 0.0541\) | 0.957561 | removed |
income + visits, smallest is visits |
\(t = 0.3216\) | 0.751683 | removed | |
income alone | \(t = 3.6642\) | 0.001775 | kept, \(p < 0.10\) — stop |
Stepwise alternates the two and reaches the same place. The final model is
\[ \widehat{\text{score}} = 40.147655 + 0.480802\,\text{income}, \]with standard errors \(5.857109\) and \(0.131217\), so \(t = 6.854517\) and \(3.664175\).
Why age is dropped although it correlates with score at
\(0.5082\). In the model age + income its coefficient is
\(-0.002852\) with a standard error of \(0.264487\), so \(t = -0.0108\) and the sign has even
flipped. The reason is the correlation of \(0.7795\) between age and
income: once income is in the model, age has nothing left to explain.
A predictor's simple correlation with the outcome tells you nothing about whether it
belongs in a model that already contains its neighbours, and this is the whole reason
selection procedures exist.
Two warnings that belong in the report. Stepwise selection with three candidates on twenty cases is already at the edge: the rule of thumb is ten or more cases per candidate predictor. And the \(p\) values printed for the final model are not honest, because the model was chosen by looking at them; they are conditional on a selection that the same data performed.
| Quantity | Logit | Probit |
|---|---|---|
| Constant | −11.401132 | −6.462962 |
| income | 0.250124 | 0.142527 |
| \(-2\) Log likelihood | 11.6415 | 11.6193 |
| \(-2\) Log likelihood, null model | 26.9205 | 26.9205 |
| Likelihood-ratio \(\chi^{2}\), 1 df | 15.2790 | 15.3012 |
| Sig. | 0.000093 | 0.000092 |
| McFadden \(R^{2}\) | 0.567559 | 0.568385 |
| Income at which \(P = 0.5\) | 45.582010 | 45.345683 |
| Classification accuracy at cut 0.5 | 0.9000 | 0.9000 |
The two fits are all but identical, and that is the normal situation: the logistic and normal distribution functions differ only in the tails. The coefficients differ by a scale factor and nothing else — \(0.250124/0.142527 = 1.755\), and the textbook rule of thumb is \(\beta_{\text{logit}} \approx 1.6\) to \(1.8\) times \(\beta_{\text{probit}}\). The fitted probabilities are so close that both classify all twenty cases identically.
So why choose one? Because \(\exp(\beta_1)\) is an odds ratio under the logit and means nothing under the probit. Here \(e^{0.250124} = 1.284184\): each extra thousand of income multiplies the odds of buying by \(1.28\). Nothing in the probit output has an interpretation that simple, which is why logistic regression is the default everywhere and probit survives mainly in bioassay, where the normal tolerance distribution has a physical meaning.
The classification table, and the five ratios read off it, are worked in STS-108, Practical 6 — accuracy \(0.900\) against a no-information rate of \(0.600\), precision and recall both \(0.875\), specificity \(0.9167\), AUC \(0.921875\). SPSS prints the classification table automatically and the ROC curve under Analyze → ROC Curve.
glm.control(epsilon = 1e-14): R's default stops one iteration early and differs in the
sixth decimal.
(a) Every method chooses \(\widehat{\text{score}} = 40.15 + 0.481\,\text{income}\)
(\(R^{2} = 0.427\), adjusted 0.395). age adds nothing once income is
in, because the two are correlated at 0.78. (b) The logit and probit fits are almost identical:
\(-2\) log likelihood 11.64 and 11.62, and both classify 18 of the 20 cases correctly. The logit
is preferred because \(e^{0.2501} = 1.28\) is an odds ratio: each extra thousand of income
multiplies the odds of buying by 1.28.
Write the SPSS syntax for a linear discriminant analysis of buy on
age, income and visits; a principal components analysis,
with varimax rotation, of age, income, visits and
score; and a hierarchical and a \(k\)-means cluster analysis of age,
income and visits.
Know where SPSS keeps each multivariate method, and which option changes its answer.
buy as the grouping variable (0, 1), the three predictors entered together, prior probabilities from the group sizes, and Box's \(M\).Discriminant analysis, principal components, factor analysis, multidimensional scaling and cluster analysis are built from the definitions, with every eigenvalue computed, in Multivariate Analysis (STS-202), Unit 3 and Unit 4. Nothing about the methods is repeated here — only where SPSS keeps them and which options change the answer.
| Method | Menu | Syntax | The option that changes the answer |
|---|---|---|---|
| Linear discriminant | Classify → Discriminant | DISCRIMINANT |
prior probabilities: equal, or proportional to group size |
| Principal components | Dimension Reduction → Factor | FACTOR /EXTRACTION=PC |
correlation matrix (default) or covariance matrix — different components |
| Factor analysis | the same dialog | FACTOR /EXTRACTION=PAF or ML |
the rotation, and whether it is orthogonal or oblique |
| Multidimensional scaling | Scale → Multidimensional Scaling | ALSCAL or PROXSCAL |
the measurement level of the proximities: ordinal or ratio |
| Cluster analysis | Classify → Hierarchical or K-Means | CLUSTER, QUICK CLUSTER |
the linkage, and whether the variables were standardized first |
DISCRIMINANT GROUPS=buy(0 1) /VARIABLES=age income visits
/METHOD=DIRECT /PRIORS=SIZE /STATISTICS=MEAN STDDEV BOXM TABLE.
FACTOR /VARIABLES age income visits score
/MISSING LISTWISE /ANALYSIS age income visits score
/PRINT INITIAL KMO EXTRACTION ROTATION
/CRITERIA MINEIGEN(1) ITERATE(25)
/EXTRACTION PC /ROTATION VARIMAX /METHOD=CORRELATION.
CLUSTER age income visits /METHOD BAVERAGE /MEASURE=SEUCLID
/PLOT DENDROGRAM /PRINT SCHEDULE.
QUICK CLUSTER age income visits /CRITERIA=CLUSTER(3) MXITER(20) /PRINT ANOVA.
The procedures run on this data and print their tables: the Classification Results for
DISCRIMINANT, Total Variance Explained and the Rotated Component Matrix for
FACTOR, the Agglomeration Schedule and the dendrogram for CLUSTER, and the
Final Cluster Centers for QUICK CLUSTER. They are not reproduced here, since SPSS is not
run on this site, and on twenty cases they should not be interpreted.
Three warnings, all of which SPSS will let you walk past.
FACTOR /EXTRACTION=PC is not factor analysis. SPSS's
default under the Factor menu is principal components, which is a different model with a
different meaning — components absorb the specific variance that a common factor sets
aside. The comparison on one correlation matrix is worked in
STS-202, Unit 4.age (range 35) and visits (range 8) is dominated by
age; the clusters are then a partition of age alone.The syntax for all four analyses, with the option that changes each answer named. The output is shown in the record but reported as illustrative only: twenty cases cannot support any of these methods, and saying so is part of the result.
Write the SPSS syntax for an \(\bar X\) and R chart, with \(C_p\) and \(C_{pk}\), for measurements taken in rational subgroups; and for a \(p\) chart of the defectives found in samples of varying size, with the run rules applied to both.
Draw the control charts from SPCHART, and read them with the run rules.
The control limits, the constants \(A_2\), \(D_3\), \(D_4\), \(B_3\), \(B_4\), \(d_2\), and
the choice between charts are all in
Statistical Quality Control —
Unit 2 for the variable charts and
process capability, Unit 3 for the
attribute charts. SPSS builds all of them from SPCHART.
| Chart | Data | SPCHART subcommand |
|---|---|---|
| \(\bar X\) and R | variables, subgroup size below 10 | /XR= |
| \(\bar X\) and s | variables, subgroup size 10 or more | /XS= |
| Individuals and moving range | variables, subgroup size 1 | /IR= |
| \(p\) | fraction defective, subgroup size may vary | /P= |
| \(np\) | number defective, constant subgroup size | /NP= |
| \(c\) | number of defects, constant area of opportunity | /C= |
| \(u\) | defects per unit, varying area | /U= |
SPCHART /XR=measure BY subgroup
/STATISTICS=CP CPK /RULES=ALL.
SPCHART /P=defectives BY sample (inspected)
/SIGMAS=3 /RULES=ALL.
SPSS draws each chart with its centre line and its \(3\sigma\) limits, and lists under it any points that break a rule.
/RULES=ALL applies the Western Electric run rules — a
point beyond \(3\sigma\), two of three beyond \(2\sigma\), four of five beyond \(1\sigma\),
eight in a row on one side — and flags them on the chart. A chart read only for points
outside the limits misses every slow drift, which is the failure mode these rules exist to
catch.
\(C_p\) and \(C_{pk}\) require the specification limits, which are an
engineering input and not in the data. /STATISTICS=CP CPK without them produces
nothing. And process capability means nothing until the process is in control:
computing \(C_{pk}\) from a chart with points outside the limits describes a process that has
no single standard deviation to speak of.
The chart is chosen by the data, not by habit: \(\bar X\) and R for small subgroups of measurements, a \(p\) chart for defectives in samples of varying size. A process is judged by the run rules as well as the limits, and its capability is quoted only once it is in control.
EXECUTE after RECODE or
COMPUTE, then wondering why the new variable is empty.FACTOR /EXTRACTION=PC a factor analysis. It is
principal components, which is a different model.