Course Description
This course explores the principles, algorithms, and systems needed to extract transformative insights from massive datasets. It provides an integrated view of modern data analytics across three core pillars: foundational big data algorithms, data management systems, and advanced analytical methods for real-world challenges.
Moving beyond traditional "big data" paradigms, the course emphasizes the disruptive opportunities created by Generative AI. Topics range from core distributed storage and computing frameworks (HDFS, Hadoop, Spark, and data lakes) to streaming and large-scale machine learning. Students will also explore next-generation architectures, including AI and semantic databases (Palimpzest, Lotus, DocETL), agentic systems, and cloud-native data analytics.
Pre-requisites: CS 242, CS 251, and CS 373
Personnel
Instructor
Chunwei LiuEmail: chunwei@purdue.edu
(Note: You MUST include "[CS440]" prefix in the email subject line)
Logistics & Materials
Meeting Times
- Lectures: Mon & Wed 11:30 AM - 12:20 PM (MSEE B012)
- Office Hours: After class or by appointment
Labs & PSOs (Starting Week 4)
- Lab 1: Wed 9:30 AM - 11:20 AM (DSAI B039)
- Lab 2: Wed 3:30 PM - 5:20 PM (DSAI B069)
- Lab 3: Fri 11:30 AM - 1:20 PM (DSAI B033)
Online Communications
- Piazza: For announcements, discussions, and Q&A.
- Brightspace: For occasional email announcements.
- Gradescope: For submitting and grading homework.
Optional Textbooks
Note that textbooks are optional and the lecture slides are self-contained.
- Spark: The Definitive Guide. By Bill Chambers and Matei Zaharia.
- Database System Concepts (7th edition). By Avi Silberschatz, Henry F. Korth, and S. Sudarshan.
- Mining of Massive Datasets. By Jure Leskovec, Anand Rajaraman, and Jeff Ullman.
Grading Breakdown
- Homeworks (3): 15% (5% each)
- Projects (3): 30% (10% each) - Related to and explained in labs
- Midterm Exam: 25% (PC-based, multiple choice, closed book)
- Final Exam: 30% (PC-based, multiple choice + open ended *TBD, closed book)
Homework Schedule
- HW 1: Relational DB basics (Leading TA: Yaoxu Song)
- HW 2: SQL & Optimization (Leading TA: Yaoxu Song)
- HW 3: Hadoop (Leading TA: Xinzhi Wang)
Project Schedule
- Project 1: MongoDB (Leading TA: Yaoxu Song)
- Project 2: Hadoop (Leading TA: Xinzhi Wang)
- Project 3: PySpark (Leading TA: Yaoxu Song)
Academic Integrity & Policies
Please review the full Integrity Policy and the Purdue AI Policy.
Academic Integrity: Each student should write up their own solutions independently. While you may discuss and obtain help with basic concepts covered in lectures or the textbook, homework specifications (but not solutions), and program design (but not implementation), any student found not following these guidelines is subject to an automatic F (final grade).
Generative AI (ChatGPT) Usage: You can ask ChatGPT for help on your homeworks and projects, but you cannot directly copy answers, and you are entirely responsible for the correctness of the generated content (as ChatGPT may return wrong answers). You cannot use ChatGPT or any electronic devices during the mid-term and final exams.
Class Schedule
* Schedule is tentative and subject to change.
| Week | Date | Lecture / Activity | Assignment | Project | Note |
|---|---|---|---|---|---|
| 1 | 08/24 | Course Introduction | |||
| 08/26 | Relational DB & Big Data | ||||
| 2 | 08/31 | Skipped | VLDB | ||
| 09/02 | Skipped | VLDB | |||
| 3 | 09/07 | No Class | HW1 Start | Labor Day | |
| 09/09 | AI Databases | ||||
| 4 | 09/14 | SQL | PSO session Start W4 | ||
| 09/16 | Database Storage | ||||
| 5 | 09/21 | Compression and Encodings | Project 1 Start | ||
| 09/23 | Index | ||||
| 6 | 09/28 | Query Processing | |||
| 09/30 | Query Processing 2 | ||||
| 7 | 10/05 | Transaction | HW2 Start | ||
| 10/07 | Concurrency Control | ||||
| 8 | 10/12 | No Class | Fall Break | ||
| 10/14 | Crash Recovery | ||||
| 9 | 10/19 | Crash Recovery 2 | Project 2 Start | ||
| 10/21 | Distributed Databases | ||||
| 10 | 10/26 | Midterm Exam (In-class) | |||
| 10/28 | Hadoop | ||||
| 11 | 11/02 | SQL-on-Hadoop | |||
| 11/04 | Big Data File Formats | HW3 Start | |||
| 12 | 11/09 | Big Data Storage | |||
| 11/11 | Spark Core | ||||
| 13 | 11/16 | Spark SQL | |||
| 11/18 | Spark ML | Project 3 Start | |||
| 14 | 11/23 | Spark Streaming | |||
| 11/25 | No Class | Thanksgiving Break | |||
| 15 | 11/30 | Spark Graph | |||
| 12/02 | Vector Data Analytics | ||||
| 16 | 12/07 | Cloud-Native Data Analytics | |||
| 12/09 | Review |