Big Data Processing with Apache Spark
Why colleges run this
When data outgrows a single machine, Spark is the industry answer; hands-on distributed processing over a genuinely large dataset is a scarce, valuable skill.
Course outcomes
CO1
Explain the distributed-processing model that Spark implements
K2CO2
Transform data using the Spark DataFrame API
K3CO3
Query large datasets with Spark SQL
K3CO4
Optimise a Spark job through partitioning and caching
K4CO5
Build a Spark pipeline over a large dataset
K3Modules tap a module for its theory & lab
Hours shown are the recommended 5-day format — module time scales to the duration you pick.
01Big data foundations5 h
Theory
The distributed-processing problem, the Spark model
Lab
Reason about why single-machine processing fails
Output
Foundations note
02Spark core5 h
Theory
RDDs, transformations, actions, lazy evaluation
Lab
Transform data with RDDs
Output
Core notebook
03DataFrames5 h
Theory
Structured data, the DataFrame API
Lab
Process structured data with DataFrames
Output
DataFrame notebook
04Spark SQL5 h
Theory
Querying at scale, query optimisation
Lab
Query a large dataset with Spark SQL
Output
Query notebook
05Data sources and formats5 h
Theory
Partitioning, columnar formats
Lab
Read and write partitioned formats
Output
Format exercise
06Performance5 h
Theory
Partitioning, caching, shuffles
Lab
Optimise a slow Spark job
Output
Optimised job
07Streaming basics5 h
Theory
Structured streaming
Lab
Build a simple streaming job
Output
Streaming job
08Capstone5 h
Theory
A pipeline over a genuinely large dataset
Lab
Build the capstone pipeline
Output
Large-scale pipeline
Every participant receives
Certificate of completion Course material LMS access Interview question bank Mock interview & viva practice Optional nasscom NSQF assessment