Skip to the content

Welcome

This is the complete study package for Statistical Methods using Python Programming (STS-105). It is a practical paper with no theory units of its own: the statistics it implements belongs to the other papers, and what it adds is the discipline of building each method from arithmetic rather than calling it.

The course's own constraint. The prescribed note says the programs must be written “Without usage of Python Packages for statistical tools”, using functions, loops, object-oriented constructs, methods and built-ins wherever possible. That is not an obstacle, it is the syllabus: a student who has written the incomplete beta function once knows what a \(t\) table is, and a student who has called scipy.stats.ttest_1samp does not.
The Python is assumed, and it is on this site. Everything in the course's “theoretical and practical concepts” list — input and output, variables and type conversion, decision and repetition structures, functions and their argument forms, modules, exception handling, files, lists, tuples, dictionaries, sets and strings — is taught with worked programs in Python Programming and Data Structures. None of it is repeated in these pages. Use that course for the language and this one for the statistics.

For the same methods done with packages — which is what STS-208 and most employment ask for — see Python for Data Analysis.

Course Objectives and Outcomes

  1. Use various data types, loop statements, object-oriented concepts, exceptions and string operations for a specified problem.
  2. Design, implement and debug a given problem using Python.
  3. Execute programs using derived and user-defined data types.
  4. Implement programs using a modular approach and file input and output.
  5. Write Python code for any statistical method for a given data set.

What is in this Course

PRACTICAL

The Twelve Programs

Every prescribed program, written out and run: matrix arithmetic and exact inverses; four sorts and two searches with their comparison counts; median and mode including the cases a library hides; grouped frequency tables and the grouping error they carry; four moments converted and checked a second way; five distributions generated from a linear congruential stream; binomial, Poisson and negative binomial fitted to one over-dispersed data set; normal, exponential and Cauchy fitted to one grouped one; correlation with both regression lines; seven tests of hypotheses; and both analyses of variance.

REFERENCE

Official Syllabus

The prescribed objectives, the list of concepts to be covered, and the twelve-item practical list, as printed.

Where the Statistics Comes From

ProgramThe theory behind it
1, 2 — matrices, determinant, inverse Linear Algebra and Linear Models, Unit 1, and the by-hand methods of STS-106
4, 5, 6 — summaries, moments, shape Descriptive Statistics
7 — random number generation Distribution Theory, Unit 1 for the distributions themselves; the inverse transform is the probability integral transform
8, 9 — fitting distributions Distribution Theory and the goodness-of-fit test of Inferential Statistics
10 — correlation and regression Statistical Methods, Unit 2 for the correlation coefficient and Unit 4 for the two regression lines, their properties and the angle between them
11 — tests of hypotheses Inferential Statistics, Units 2 and 3, and Estimation Theory (STS-201)
12 — analysis of variance Design and Analysis of Experiments, and STS-203 for the two-way case with several observations per cell

How to Use This Paper

  1. Write tails.py first. Four of the twelve programs need a \(t\), \(F\) or \(\chi^{2}\) tail area, and no package may supply one. It is written once on the practical page and imported thereafter.
  2. Type the programs, do not read them. The course is examined by writing code on paper and running it; reading a listing does not build that.
  3. Check every answer a second way. Each program on the practical page ends with a check — an identity that must hold, or the same quantity by another route. That habit is what the examination is really testing.

Next course in learning order: Data Science using Python Statistical computing