Skip to the content
On this page
  1. Why this is the most important course in the degree
  2. Where it sits
  3. Course objectives (verbatim)
  4. Units in this Course

Why this is the most important course in the degree

Every other course teaches you about data science. This one teaches the tools you will actually use, every day, in any data job you take.

NumPy and Pandas are not one option among several. They are the foundation: scikit-learn takes NumPy arrays, matplotlib plots them, every deep learning framework mirrors their API, and essentially every Python data pipeline written in the last decade passes through a DataFrame. Data Mining's mlxtend and sklearn labs are Pandas code underneath.

This is also where finding D8 finally closes. Computer Fundamentals and Office Automation taught spreadsheet analysis and Statistical Foundations for Data Science taught statistics by hand; neither connected to the programming in Problem Solving Using C and Python Programming and Data Structures. Python for Data Analysis and Visualization is the join: everything you computed with a formula in Statistical Foundations for Data Science becomes one method call here, and §5.7 recomputes Statistical Foundations for Data Science's worked examples in Pandas to prove the two agree.

Where it sits

From You have Used here for
Python Programming and Data Structures Python, lists, dicts, comprehensions, files The language itself; Unit 1 contrasts lists with arrays
Statistical Foundations for Data Science Mean, variance, correlation, distributions Unit 1's statistical functions; Unit 5's groupby
Database Management Systems SQL SELECT, WHERE, GROUP BY, JOIN Unit 5's merge and groupby are the same operations
Data Science with R dplyr's five verbs, ggplot2 The direct R counterpart — §5.8 maps them
Data Mining Preprocessing theory Unit 3 is that theory, executed

The three languages of one idea

The single most useful thing to notice in this course is that SQL, dplyr and Pandas express the same operations:

Operation SQL dplyr (Data Science with R) Pandas (here)
Filter rows WHERE filter() df[df.x > 5]
Pick columns SELECT select() df[["a", "b"]]
New column AS mutate() df.assign(...)
Sort ORDER BY arrange() df.sort_values()
Aggregate GROUP BY group_by() + summarise() df.groupby().agg()
Join JOIN left_join() pd.merge()

Learn one column and you have learned three.

Course objectives (verbatim)

  1. Introduce foundational concepts of NumPy arrays and array operations for efficient numerical computing.

  2. Teach key data structures and manipulation techniques using Pandas.

  3. Enable students to perform data input/output operations and implement basic data cleaning workflows.

  4. Explore string processing methods and feature engineering strategies in Pandas.

  5. Guide learners in advanced data wrangling tasks including merging, reshaping, hierarchical indexing and visualization.

Units in this Course

UNIT 1

NumPy Essentials

The ndarray against the Python list; creating arrays and dtypes; arithmetic and broadcasting; basic, boolean and fancy indexing, and which return views; transposing and swapping axes; universal functions; statistical functions and the axis parameter; random number generation.

UNIT 2

Pandas Basics and Data Structures

Series and DataFrame; Index objects; the three accessors and why ‘loc’ is inclusive; filtering and boolean indexing; arithmetic and data alignment; sorting and the five ranking methods; dropping entries; duplicate indexes.

UNIT 3

Data Input, Output and Cleaning

read_csv and the parameters that matter; JSON and json_normalize; Excel; detecting, dropping and filling missing data; replacing sentinel values; renaming axes; removing duplicates; filtering outliers by z-score and IQR; transforming with map, apply and transform.

UNIT 4

String Operations and Feature Engineering

The .str accessor; regular expressions with extract, contains and replace; engineering features from dates, numbers and categories; dummy and indicator variables and the dummy variable trap; permutation, stratified sampling and the bootstrap.

UNIT 5

Wrangling, Reshaping and Visualization

Merging and the four join types; concatenation; combining with overlap; hierarchical indexing; pivot, melt, stack and unstack; split–apply–combine; recomputing Course 4 in Pandas; matplotlib, Seaborn and Plotly.

PRACTICE

Practice

Exam-style questions with fully worked solutions.

LAB

Lab

Every prescribed lab experiment, with code and expected output.

Next course in learning order: Time Series Analysis and Forecasting Statistics & analysis