Skip to the content

Source document. This page reproduces the syllabus this course was written to, as published — its semesters, credits and paper numbers are that document’s, not this site’s. The course itself is studied on its own, in any order.

This page reproduces the prescribed outline so that the teaching pages can be checked against it line by line. It is the syllabus, not a summary of it. Nothing about marks, duration or examination pattern appears on this site.

Course Objectives and Outcomes

  1. Able to apply to the data set to Extract, Transform, Load pipeline which will extract raw data (text files, CSV files, XML files, JSON, HTML files, SQL databases, NoSQL databases etc.), clean the data, perform transformations on data, load data and visualize the data.
  2. Able to familiar with expert in usage of Python libraries/modules: Pandas, Numpy, Beautiful Soup, PyMysql, PyMongo, Nltk, MatplotLib etc.

Datasets

AS PRESCRIBED

Publicly available appropriate datasets can be used for statistical data analysis using Python. Few of them are: MNIST, UCI Machine Learning Repository, Twitter Data, MOSPI, IIPS, Kaggle etc.

The practical page uses a small data set written out in full instead, so that every figure on the page can be reproduced without downloading anything.

List of Practicals

AS PRESCRIBED
  1. Write programs to parse text files, CSV, HTML, XML and JSON documents and extract relevant data. After retrieving data check any anomalies in the data, missing values etc.
  2. Write programs for reading and writing binary files.
  3. Write programs for searching, splitting, and replacing strings based on pattern matching using regular expressions.
  4. Design a relational database for a small application and populate the database. Using SQL do the CRUD (create, read, update and delete) operations.
  5. Create a Python MongoDB client using the Python module pymongo. Using a collection object practice functions for inserting, searching, removing, updating, replacing, and aggregating documents, as well as for creating indexes.
  6. Write programs to create Numpy arrays of different shapes and from different sources, reshape and slice arrays, add array indexes, and apply arithmetic, logic, and aggregation functions to some or all array elements.
  7. Write programs to use the Pandas data structures: Frames and series as storage containers and for a variety of data-wrangling operations, such as:
    • Single-level and hierarchical indexing
    • Handling missing data
    • Arithmetic and Boolean operations on entire columns and tables
    • Database-type operations (such as merging and aggregation)
    • Plotting individual columns and whole tables
    • Reading data from files and writing data to files.

Note, as Printed

THE PRACTICAL RECORD

Practical Record should contain all practical's with their implementation and is Mandatory and it carries 5 Marks. The Semester end practical exam contains answer any two with their implementation out of four Questions. (Answer any two out of the four).

Where the Prerequisites Are

ASSUMED BEFORE THIS PAPER, AND ON THIS SITE

The Python language is covered by Python Programming and Data Structures; NumPy and Pandas by Python for Data Analysis; SQL and relational design by Database Management Systems; the document model by Document Databases; and NLTK by Natural Language Processing. The course home page maps each of the seven prescribed items to the pages that teach it and names what is written out on the practical page instead.

→ The pipeline, end to end