Source document. This page reproduces the official programme syllabus the courses were written to — its semesters, elective tracks and course numbers are that document’s, not this site’s. The courses themselves are studied on their own, in any order.
Source: docs/Data-Science-Major-Sem3-4.pdf — 25 pages.
Extracted verbatim so every claim in the notes is traceable to a page.
Regenerate with python3 tools/extract_syllabus.py <pdf>.
SEMESTER-III COURSE 6: DATA SCIENCE WITH R Theory Credits: 3 3 hrs/week Course Objectives 1. Introduce the data science process, lifecycle, and applications in real-world domains. 2. Build proficiency in R programming for data manipulation, exploration, and visualization. 3. Train students in handling structured, unstructured, and time-based data effectively. 4. Familiarize with basic machine learning and statistical modeling using R. 5. Develop awareness of ethical, interpretability, and responsible use of data science. Course Outcomes At the end of the course, students will be able to: 1. Explain the Data Science process and perform EDA (Exploratory Data Analysis). 2. Write R programs using variables, functions, loops, and packages for basic analytics. 3. Perform data wrangling, cleaning, and visualization with R libraries (dplyr, tidyr, ggplot2). 4. Build and evaluate basic machine learning models such as regression and clustering. 5. Apply data science techniques to practical case studies. Unit 1. Introduction to Data Science Process: Introduction- Definition - Data Science in various fields - Examples - Impact of Data Science - Data Analytics Life Cycle - Data Science Toolkit - Data Scientist - Data Science Team, Exploratory Data Analysis (EDA), Feature Engineering & Data Transformation Unit 2. Basics of R Programming: Introduction to R and RStudio, Data Types, Variables, Operators, Control Structures (if, loops, apply), Functions and Packages, Data Input/Output (CSV, Excel, XML, JSON). Unit 3. Data Handling & Visualization in R: Data Frames, Lists, Matrices, Data Wrangling with dplyr and tidyr, Handling Missing Data, Working with Date/Time in R. Visualization with ggplot2: grammar of graphics, aesthetics, geometries, scales.
Faceting and layering techniques,Visualizing categorical and numerical data, Customizing and exporting plots Unit 4. Applications & Case Studies in Data Science: Simple Linear Regression, Multiple Regression Model Evaluation Method: Accuracy, Confusion Matrix, ROC. K-Means Clustering, Text Mining & Word Clouds, Recommender Systems Basics, Ethical Issues in Data Science Unit 5. Advanced Topics in Data Science with R : Introduction to Time Series Analysis in R (ARIMA basics)- Concept of time series (trend, seasonality, noise), Time series objects in R (ts, zoo, xts), Plotting and decomposing time series, Stationarity and differencing, Autocorrelation & Partial Autocorrelation (ACF/PACF), AR, MA, ARIMA model basics, Forecasting using forecast package Creating interactive visualizations with plotly packages-Converting ggplot2 plots to interactive plots Animations and sliders in plotly R Shiny: Building interactive web applications-Introduction to Shiny framework, UI and server functions, Reactive expressions and reactivity in Shiny, Input and output widgets (sliders, dropdowns, text), Layouts and dashboard design Textbooks 1. An Introduction to Statistical Learning with Applications in R, Gareth James, Daniela Witten, Trevor Hastie, Robert Tibshirani, Springer, 2nd Edition, 2021 2. Reference Books 1. The Art of R Programming, Norman Matloff,, No Starch Press, 2011. 2. Modern Applied Statistics with S, W.N. Venables & B.D. Ripley, Springer, 2002. 3. Introduction to Data Science: Data Analysis and Prediction Algorithms with R, Rafael A. Irizarry, CRC Press, 2020. 4. Data Science from Scratch: First Principles with Python (for conceptual clarity only), Joel Grus,
Activities: Outcome: Explain the Data Science process and perform EDA (Exploratory Data Analysis). Activity: Use a real-world dataset (e.g., Titanic or COVID data) to: Outline the steps of the Data Science workflow Perform EDA using summary statistics and visualizations (histograms, boxplots, scatterplots) Evaluation Method: Presentation and checklist (10-point scale): Clear explanation of workflow stages Quality of EDA insights Use of appropriate plots and summaries Outcome: Write R programs using variables, functions, loops, and packages for basic analytics. Activity: Write an R script that: Reads a CSV file Uses if, for, and while loops Defines and calls custom functions with arguments and return values Evaluation Method: Code review and execution test to verify (10-point scale): Correctness of the syntax and logic Functionality of control structures Output accuracy and modularity Outcome: Perform data wrangling, cleaning, and visualization with R libraries (dplyr, tidyr, ggplot2). Activity: Clean a messy dataset using: dplyr for filtering, selecting, and mutating tidyr for reshaping and handling missing values Time-based operations (e.g., filling gaps, formatting dates) Evaluation Method: Before-and-after comparison (10 point score): Completeness of cleaning steps Use of appropriate functions Handling of missing/time data
Outcome: Implement basic machine learning models and evaluate performance using appropriate metrics and visual tools. Activity: Build a simple classification model (e.g., logistic regression or decision tree) using R: Train/test split Predict outcomes Evaluate using confusion matrix, accuracy, precision, recall Evaluation Method: Model report and demo (10 point scale): Correct implementation of model Use of evaluation metrics
SEMESTER-III COURSE 6: DATA SCIENCE WITH R Practical Credits: 1 2 hrs/week List of Practicals: 1. Compute Mean, Median, Mode, Variance, and Standard Deviation 2. Visualize Binomial, Normal, and Poisson Distributions 3. Perform t-test and Chi-Square Test in R 4. Calculate Correlation and Build a Simple Linear Regression Model 5. Conduct Exploratory Data Analysis (EDA) on a Real-World Dataset 6. Apply Feature Engineering: Scaling, Normalization, and Encoding 7. Practice R Programming: Variables, Control Structures, and Functions 8. Read and Write Data from CSV, Excel, JSON, and XML Files 9. Use dplyr and tidyr for Data Wrangling Tasks 10. Handle Missing Data and Detect Outliers 11. Work with Dates and Times in R 12. Visualize Data Using ggplot2 (Bar, Scatter, Histogram, Boxplot) 13. Perform K-Means Clustering and Visualize Clusters 14. Evaluate Models Using Confusion Matrix, Accuracy, and ROC Curve 15. Perform Text Mining and Create a Word Cloud 16. Time Series Forecasting with ARIMA on a real dateset (e.g., monthly airline passengers, stock prices, or temperature data). 17. Create interactive bar, line, and scatter plots using plotly. On a real dataset (e.g., COVID-19 cases, sales data, or student marks). 18. Develop a Shiny app that lets users upload a CSV file.
SEMESTER-III COURSE 7: WEB TECHNOLOGIES Theory Credits: 3 3 hrs/week Course Objectives 1. Understand the principles of web design and distinguish between web and desktop application architectures. 2. Develop static web pages using HTML elements, attributes, and multimedia integration techniques. 3. Style web pages effectively using CSS, including layout control, responsive design, and UI enhancements. 4. Implement dynamic behaviors and form validations using JavaScript and the Document Object Model (DOM). 5. Explore JSON and jQuery for handling structured data and simplifying client-side scripting in web development. Course Outcomes At the end of the course, students will be able to: 1. Design and structure HTML-based webpages incorporating text, images, tables, forms, and multimedia content. 2. Apply CSS styling rules to manage layout aesthetics, interactivity, and responsiveness across devices. 3. Use JavaScript for string manipulation, event handling, arrays, object operations, and basic validation. 4. Employ client-side scripting to enhance form functionality, create dialog interactions, and add animations via events. 5. Parse JSON data and use jQuery to simplify DOM manipulation, AJAX calls, and build dynamic, data-driven web applications. Unit 1.HTML: Introduction to web designing, difference between web applications and desktop applications, introduction to HTML, HTML structure, elements, attributes, headings, paragraphs, images, tables, lists, blocks, symbols, embedding multi-media components in HTML, HTML forms
Unit 2.CSS: CSS home, introduction, syntax, CSS combinators, colors, background, borders, margins, padding, height/width, text, fonts, tables, lists, position, overflow, float, pseudo class, pseudo elements, opacity, tool tips, image gallery, CSS forms, CSS counters. Unit 3.Java Script: What is DHTML, JavaScript, basics, variables, operators, statements, string manipulations, mathematical functions, arrays, functions. objects, regular expressions, exception handling. Unit 4. Client-Side Scripting: Accessing HTML form elements using Java Script object model, basic data validations, data format validations, generating responsive messages, opening windows using java script, different kinds of dialog boxes, accessing status bar using java script, embedding basic animative features using different keyboard and mouse events. Unit 5. JSON and jQuery Introduction to JSON: Need for data exchange formats, JSON syntax, JSON vs XML, parsing JSON, creating JSON objects and arrays, accessing nested JSON data, reading/writing JSON in JavaScript. Working with jQuery: Introduction to jQuery, selectors, filters, DOM manipulation, event handling, animations, effects, and chaining. Text Book(s) 1. Web Programming: Building Internet Applications, Chris Bates, Wiley, Second Edition. 2. An Introduction to Web Design plus Programming, Paul S. Wang, Sanda S. Katila, Thomson. 3. Learning jQuery Jonathan Chaffer, Karl Swedberg, Packt Publishing. 4. JSON at Work Media. Reference Books 1. 2. An Introduction to HTML and JavaScript: for Scientists and Engineers, David R. Brooks, Springer.
SEMESTER-III COURSE 7: WEB TECHNOLOGIES Practical Credits: 1 2 hrs/week List of Experiments: 1. Create an HTML document with the following formatting options: (a) Bold, (b) Italics, (c) Underline, (d) Headings (Using H1 to H6 heading styles), (e) Font (Type, Size and Color), (f) Background (Colored background/Image in background), (g) Paragraph, (h) Line Break, (i) Horizontal Rule, (j) Pre tag 2. Create an HTML document which consists of: (a) Ordered List (b) Unordered List (c) Nested List (d) Image 3. Collect any ten images of your choice. Using table tag, align the images as follows: 4. Create a form using HTML which has the following types of controls: (a) Text Box (b) Option/radio buttons (c) Check boxes (d) Reset and Submit buttons 5. Embed a calendar object in your web page. 6. Create a form that accepts the information from the subscriber of a mailing system. 7. Apply CSS to design a student registration form (use different selectors, colors, borders, spacing). 8. Create a responsive webpage using CSS Flexbox/Grid. 9. Add hover effects and transitions on images and buttons using CSS. 10. Write a JavaScript program to perform string operations (reverse, substring, count vowels). 11. Create a JavaScript form validation program (check email format, password length, required fields).
SEMESTER-IV COURSE 8: DATA MINING Theory Credits: 3 3 hrs/week Course Objectives: Provide an understanding of data warehousing concepts, architecture, and OLAP operations for effective storage, modeling, and analysis. Develop knowledge of data mining fundamentals, tasks, and preprocessing techniques to prepare data for mining. Introduce students to association rule mining algorithms for discovering hidden patterns and relationships in large datasets. Enable learners to apply classification techniques (decision trees, Bayesian, nearest neighbor, rule-based) for predictive modeling. Equip students with knowledge of clustering paradigms and algorithms (partitioning, hierarchical, density-based, categorical) for data grouping and pattern discovery. Course Outcomes: Upon successful completion of the course, the student will be able to: 1. Explain the architecture, schemas, and OLAP operations of data warehousing and distinguish it from traditional database systems. 2. Apply preprocessing techniques (data cleaning, dimensionality reduction, feature selection, transformation, and similarity measures) to prepare raw data for analysis. 3. Implement various association rule mining algorithms (Apriori, Partition, FP-Growth, etc.) to uncover meaningful relationships within large datasets. 4. Build and evaluate classification models using decision tree algorithms (ID3, C4.5, CART), rule-based classifiers, Bayesian classifiers, and nearest-neighbor methods. 5. Analyze and implement clustering techniques such as K-Means, K-Medoid, DBSCAN, BIRCH, and categorical clustering methods (STIRR, ROCK, CACTUS) for grouping and pattern discovery in different types of datasets. Unit-1: Data Warehousing: Introduction to Data Ware House, Differences between Database systems and Data Ware House, Data Ware House characteristics, Data Ware House Architecture and its components,
Data Modeling, Schema Design, star and snow-Flake Schema, Fact Constellation, Fact Table, OLAP cube, OLAP Operations. Unit-2: Data Mining: What is Data Mining? Data Mining: Definitions, KDD vs Data Mining, Data Mining Tasks, Data Preprocessing- Data Cleaning, Missing Data, Dimensionality Reduction, Feature Subset Selection, Discretization and Binarization, Data Transformation; Measures of similarity and Dissimilarity-Basics. Issues and Challenges in DM, DM Applications- Case Studies Unit-3: Association Analysis: Association Rules: What is an Association Rule?, Methods to Discover Association Rules, A Priori Algorithm, Partition Algorithm, Pincer-Search Algorithm, Dynamic Itemset Counting Algorithms, FP-Tree Growth Algorithm, Generalized Association Rule, Association Rules with Item Constraints Unit-4: Classification: Definition, What is Decision Tree?, Tree Construction Principle, Best Split, Splitting Indices, Splitting Criteria, Decision Tree Construction Algorithms: CART, ID3, C4.5, Method for Comparing Classifiers, Rule Based Classifiers, Nearest Neighbor Classifiers, Bayesian Classifiers. Unit-5: Clustering Techniques: Clustering Paradigms, Partitioning Algorithms (K-Means), k-Medoid Algorithms, Hierarchical Clustering: DBSCAN, BIRCH, Categorical Clustering Algorithms: STIRR, ROCK, CACTUS Textbooks: 1. Data Mining Techniques, Arun K Pujari, 3rd Edition 2. Data Mining: Concepts and Techniques, Jiawei Han, Micheline Kamber, Jian Pei, 3rd Edition, Morgan Kaufmann Publishers Reference Books: 1. K.P. Soman, Shyam Diwakar, V.Ajay,2006, Insight into Data Mining Theory and Practice, Prentice Hall of India Pvt. Ltd.
Evaluation Method: Submissions will be graded on correctness of generated rules, clarity of code/parameters, quality of interpretation, and ability to link discovered rules to practical business insights. Outcome 4: Build and evaluate classification models using decision tree algorithms (ID3, C4.5, CART), rule-based classifiers, Bayesian classifiers, and nearest-neighbor methods. Activity: Students will implement at least two classification algorithms (e.g., Decision Tree + Naïve Bayes or KNN) on a dataset like Iris, Titanic, or Student Performance. They will evaluate models using confusion matrix, accuracy, precision, recall, and F1-score and compare results. Evaluation Method: Assessment will consider correctness of implementation, clarity in performance comparison, and visualization of decision trees/rules. Students must explain why certain models performed better. Outcome: Analyze and implement clustering techniques such as K-Means, K-Medoid, DBSCAN, BIRCH, and categorical clustering methods (STIRR, ROCK, CACTUS). Activity: Students will implement at least two clustering algorithms (e.g., K-Means and DBSCAN or BIRCH) on real datasets (e.g., customer segmentation, text data, student groups). They will compare clusters using Silhouette Score, Davies-Bouldin Index, and visualize results using scatter plots/dendrograms. Evaluation Method: Evaluation will be based on accuracy of clustering implementation, choice of parameters (e.g., k in K-Means, eps in DBSCAN), visualization quality, and clarity in interpretation of clusters.
SEMESTER-IV COURSE 8: DATA MINING Practical Credits: 1 2 hrs/week List of Experiments: Recommended datasets: weather.arff, iris.arff, supermarket.arff, vote.arff, contact-lenses.arff, or custom CSV datasets. 1. Load datasets in WEKA and explore data formats (ARFF/CSV) 2. Perform data cleaning and handle missing values using filters 3. Apply normalization and discretization on numeric attributes 4. Reduce data using attribute selection and PCA 5. Summarize and visualize data using statistical tools and class-wise comparison in WEKA. 6. Generate association rules using the Apriori algorithm 7. Apply multilevel association rule mining using hierarchical attributes 8. Apply K-means clustering and interpret the cluster outputs. 9. Perform hierarchical clustering and visualize results using dendrograms. 10. Apply Expectation-Maximization (EM) clustering and analyze cluster summaries. 11. Build a decision tree classifier using J48 and evaluate its performance. 12. Perform Naive Bayes classification and compare with decision tree results. 13. Apply rule-based classification using PART or JRip algorithms. 14. Compare classifiers using confusion matrix, accuracy, and ROC curves. 15. Perform basic text preprocessing and clustering using TF-IDF and K-means.
SEMESTER-IV COURSE 9: PYTHON FOR DATA ANALYSIS AND VISUALIZATION Theory Credits: 3 3 hrs/week Course Objectives: 1. Introduce foundational concepts of NumPy arrays and array operations for efficient numerical computing. 2. Teach key data structures and manipulation techniques using Pandas. 3. Enable students to perform data input/output operations and implement basic data cleaning workflows. 4. Explore string processing methods and feature engineering strategies in Pandas. 5. Guide learners in advanced data wrangling tasks including merging, reshaping, hierarchical indexing and visualization. Course Outcomes: 1. Demonstrate proficiency in creating and manipulating NumPy arrays for mathematical operations and simulations. 2. Apply Pandas Series and DataFrame operations for structured data handling and analysis. 3. Read, write, and clean diverse data formats using Python tools, addressing missing values and outliers. 4. Implement vectorized string operations and create derived features for enhanced model readiness. 5. Perform complex data wrangling tasks such as merging datasets, reshaping data structures, generating group-level statistics, Visualize the data. Unit 1. NumPy Essentials: NumPy ndarray: A Multidimensional Array Object, Creating ndarrays, Data Types for ndarrays, Arithmetic with Arrays, Basic Indexing and Slicing, Boolean Indexing, Fancy Indexing, Transposing Arrays, Swapping Axes, Universal Functions: Element-wise Operations, Basic Mathematical and Statistical Functions, Random Number Generation (basic use)
Unit 2. Pandas Basics and Data Structures: Series, DataFrame, Index objects, Indexing and Selection, Filtering and Boolean Indexing, Arithmetic and Data Alignment, Sorting and Ranking, Dropping Entries, Handling Duplicate Indexes Unit 3. Data Input, Output, and Cleaning: Reading and Writing Data in Text Format (CSV, TXT), Working with JSON, Reading Microsoft Excel Files, Handling Missing Data, Dropping and Filling Missing Values, Replacing Values, Renaming Axis Indexes, Removing Duplicates, Filtering Outliers, Transforming Data Using Mapping or Functions Unit 4. String Operations and Feature Engineering: String Methods in pandas, Basic Regular Expressions, Vectorized String Functions, Creating Dummy/Indicator Variables, Permutation and Random Sampling. Unit 5. Data Wrangling, Reshaping & Visualization: Merging and Joining Datasets, Concatenating Along an Axis, Combining Data with Overlap, Reshaping with Pivot, Stack, and Unstack, Basic Hierarchical Indexing, Summary Statistics by Group or Level Introduction to matplotlib: plots, customization, styling, Seaborn for statistical data, visualization, Plotly for interactive charts and dashboards. Textbooks 1. Python for Data Analysis: Data Wrangling with pandas, NumPy, and Jupyter, Wes 2. Python Programming-An Object Oriented Approach, Anita Goel 3. Python for Data Science For Dummies, Yuli Vasiliev, 2nd Edition, Wiley, 2022. Reference Books 1. Python Data Science Handbook: Essential Tools for Working with Data, Jake VanderPlas, Media, Reprint Edition, 2023.
Cleaning completeness Outlier detection logic Outcome: Implement vectorized string operations and create derived features for enhanced model readiness. Activity: Prepare text data for modelling to: Use str methods to clean and standardize strings Extract features (e.g., domain from email, length of name) Encode categorical variables (e.g., get_dummies, LabelEncoder) Evaluation Method: Feature report to check (10-point scale): Efficiency of vectorized operations Relevance of derived features Readiness for ML input Outcome: Perform complex data wrangling tasks such as merging datasets, reshaping data structures, generating group-level statistics and Visualization. Activity: Integrate and reshape datasets: Merge two datasets on a common key Reshape using pivot, melt, stack, unstack Generate group-level stats (e.g., mean sales per region) Visualize the Datasets Evaluation Method: Before-and-after comparison to validate: Accuracy of merge and reshape Correct use of aggregation Final structure suitability for analysis
SEMESTER-IV COURSE 9: PYTHON FOR DATA ANALYSIS AND VISUALIZATION Practical Credits: 1 2 hrs/week C9P: Python for Data Analysis and Visualization Lab List of Practicals: 1. Create and Manipulate NumPy ndarrays; Explore Data Types 2. Perform Arithmetic Operations and Element-wise Calculations on Arrays 3. Practice Indexing, Slicing, Boolean, and Fancy Indexing on ndarrays 4. Use Universal Functions and Compute Basic Mathematical/Statistical Functions with NumPy 5. Create and Manipulate Pandas Series and DataFrames 6. Perform Indexing, Selection, Filtering, and Boolean Indexing in Pandas 7. Conduct Arithmetic Operations and Data Alignment in DataFrames 8. Sort, Rank, Drop Entries and Handle Duplicate Indexes in Pandas 9. Read and Write Data in CSV, TXT, JSON, and Excel Formats 10. Handle Missing Data: Detect, Drop, Fill, and Replace Missing Values 11. Rename Axis Indexes, Remove Duplicates, and Filter Outliers 12. Transform Data Using Mapping Functions and Apply String Operations 13. Perform String Operations and Use Regular Expressions on DataFrames 14. Create Dummy Variables and Perform Permutations and Random Sampling 15. Merge, Join, and Concatenate Datasets Using Pandas 16. Reshape Data Using Pivot, Stack, Unstack, and Perform Hierarchical Indexing 17. Compute Summary Statistics Grouped by Levels or Categories 18. Basic Visualizations using Matplotlib
SEMESTER-IV COURSE 10: DOCUMENT ORIENTED DATABASE Theory Credits: 3 3 hrs/week Course Objectives 1. To introduce students to the concepts of NoSQL databases and their significance compared to traditional relational databases. 2. To provide hands-on experience with MongoDB for performing CRUD operations, querying, and advanced data handling. 3. To develop skills in schema design, data modeling, and working with embedded and referenced documents. 4. replication, and transactions. 5. To prepare students for real-world applications of MongoDB in scalable, high-performance data-driven applications. Course Outcomes On successful completion of this course, students will be able to: 1. Differentiate between SQL and NoSQL databases, and explain the architecture and features of MongoDB. 2. Perform CRUD operations and construct queries using MongoDB Query Language (MQL). 3. Apply schema design strategies and use appropriate data modeling techniques for different application scenarios. 4. Utilize advanced features like indexing, aggregation, GridFS, and transactions to optimize data handling. 5. Implement replication concepts and ensure high availability, fault tolerance, and scalability in MongoDB-based applications. Unit 1. Introduction to NoSQL & Fundamentals of MongoDB: What is NoSQL DB? History & evolution of NoSQL, Features of NoSQL databases, CAP theorem & BASE properties, Types of NoSQL (Key-Value, Document, Column, Graph), Difference between RDBMS & NoSQL, Why and when to use NoSQL?, NoSQL Database misconceptions, Benefits & real-world use cases of NoSQL, Comparison of popular NoSQL systems (Redis, Cassandra, CouchDB, Neo4j), Introduction to JSON & BSON
Installation & Setupservice), connecting via Mongo shell or GUI. Unit 2. MongoDB Architecture, Data Modeling and Basics: MongoDB Architecture: Database, Collection, Document concepts, BSON format, Advantages of MongoDB over RDBMS, MongoDB Datatypes (String, Number, Date, Boolean, Array, ObjectId, Embedded Documents, Null) Data Modeling in MongoDB: Schema design strategies, Embedded vs Referenced documents Database & Collection Management: Create & Drop Database, Create & Drop Collection Unit 3. CRUD Operations and Querying in MongoDB: CRUD Operations: Insert Documents (insertOne, insertMany), Query Documents (find, operators, conditions), Update Documents (updateOne, updateMany, replaceOne), Delete Documents (deleteOne, deleteMany) Query operators ($gt, $lt, $in, $nin, $and, $or, $not), Regular expression queries, Bulk operations Working with Arrays. Unit 4. Data Modelling and Aggregation: Data Modelling and Aggregation: Data Models: Introduction to embedded vs normalized models, advantages and trade-offs Embedded Data Models: Use cases, benefits, and limitations Normalized Data Models: References between documents, when to normalize data Relationships Between Documents, Data Model Using an Embedded Document, Data Model Using Document References Aggregation Basics: Introduction to MongoDB Aggregation Framework, simple pipelines and operators Unit 5. Advanced Query Processing and Optimization in MongoDB: Query Optimization & Operations: Projection, Limiting & Skipping Records, Sorting Records Indexing in MongoDB (single field, compound, multikey, text index), Aggregation Framework (pipelines, stages, operators), Replication Concepts: Replica sets, failover, consistency
Textbooks: 1. MongoDB: The Definitive Guide, Shannon Bradshaw, Eoin Brazil, Kristina Chodorow, 2. MongoDB Recipes: With Data Modeling and Query Building Strategies, Subhashini Chellappan, Dharanitharan Ganesan , Apress Reference Books: 1. MongoDB in Action, Kyle Banker, 2nd Edition, Manning Publications, 2016. 2. 3. M Web Resources: 1. Official MongoDB Documentation: https://www.mongodb.com/docs/ 2. MongoDB University Free Courses: https://learn.mongodb.com/ 3. W3Schools, Tutorialspoint, geeksforgeeks Activities: Outcome: Differentiate between SQL and NoSQL databases, and explain the architecture and features of MongoDB Activity: Students create a comparative chart (SQL vs NoSQL) and draw MongoDB architecture diagram. Evaluation Method: Assess based on accuracy, clarity of explanation, and presentation. Outcome: Perform CRUD operations and construct queries using MongoDB Query Language (MQL) Activity: Implement CRUD operations and write 5 different queries on a sample dataset (e.g., Library or E-commerce). Evaluation Method: Marks for correct execution, query correctness, and output validation. Outcome: Apply schema design strategies and use appropriate data modeling techniques for different application scenarios Activity: Design schema for a university management system (students, courses, faculty) using MongoDB data modeling techniques. Evaluation Method: Evaluate on schema correctness, use of embedding/referencing, and justification of design.
Outcome: Utilize advanced features like indexing, aggregation, GridFS, and transactions to optimize data handling Activity: Create indexes and implement an aggregation pipeline to generate sales report from a dataset. Evaluation Method: Marks for index implementation, aggregation correctness, and efficiency of results. Outcome: Implement replication concepts and ensure high availability, fault tolerance, and scalability in MongoDB-based applications Activity: Configure a replica set with primary and secondary nodes, and demonstrate failover. Evaluation Method: Assess on setup correctness, successful failover demonstration, and report/documentation.
SEMESTER-IV COURSE 10: DOCUMENT ORIENTED DATABASE Practical Credits: 1 2 hrs/week 1. Installation and setup of MongoDB, connecting to Mongo Shell and Compass. 2. Creating and using databases, creating collections, inserting documents. 3. Basic queries using find(), filtering with comparison operators. 4. Using logical operators ($and, $or, $not, $nor) for complex queries. 5. Updating documents with $set, $unset, $inc, $rename. 6. Deleting documents using deleteOne() and deleteMany(). 7. Using projection to display selective fields. 8. Sorting documents, limiting output, skipping records. 9. Designing an Embedded Data Model for a student-course enrollment system. 10. Designing a Normalized Data Model using document references. 11. Modeling relationships: One-to-One, One-to-Many, Many-to-Many in MongoDB. 12. Implementing schema validation using JSON Schema in MongoDB. 13. Creating and testing single-field and compound indexes. 14. Using text search and multikey indexes. 15. Building aggregation pipelines with $match, $group, $project, $sort. 16. Advanced aggregation operators: $lookup, $unwind, $bucket. 17. Configuring and testing replication with a replica set (minimum 3 nodes). 18. Storing and retrieving large files using GridFS. 19. Using MongoDB Transactions for multi-document consistency. 20. Case Study: Developing a mini-application (e.g., Library Management / E-commerce Cart) using MongoDB CRUD, Aggregation, Indexing, and Replication.