Introduction to Apache Spark and SparkSQL gives attendees a hands-on introduction to Spark and SparkSQL across two days. The first day covers the fundamentals of the Spark distributed computing engine, including installing and running Spark. Attendees also work with Resilient Distributed Datasets (RDDs) and ingest and process file-based data. The second day shifts focus to SparkSQL, covering the Hive Metastore, DataFrames, querying external databases over JDBC/ODBC, and writing user-defined functions. Hands-on labs accompany every module, giving attendees practical, immediately usable skills.
RDD labs move from basic data analysis into key/value operations. Attendees transform, aggregate, group, join, and sort keyed datasets, core skills for any distributed data pipeline. Day one closes with file-based I/O labs that ingest and process data directly from disk. This gives attendees a feel for how Spark partitions and reads large datasets in parallel.
The second day opens by querying unstructured Hadoop data through the Hive Metastore. It then moves into manipulating DataFrames alongside RDDs to compare the two programming models directly. A dedicated module covers querying external databases over JDBC and ODBC sources, followed by writing and applying user-defined functions to extend built-in SQL capabilities. Each lab builds on the previous one, so attendees leave with a working, end-to-end pipeline rather than isolated exercises.
Who Should Attend
Application developers, analysts and data scientists
What Attendees Will Learn
Upon completing Introduction to Apache Spark and SparkSQL, attendees will be able to:
- Install and configure Apache Spark and run Spark applications
- Work with Resilient Distributed Datasets (RDDs) to transform, aggregate, group, join, and sort data
- Ingest and process file-based data with Spark I/O
- Query structured and unstructured data using SparkSQL and the Hive Metastore
- Manipulate data using Spark DataFrames
- Query external databases through JDBC/ODBC sources from Spark
- Write and use Spark user-defined functions (UDFs)
Prerequisites
Basic Linux command line skills are helpful. The coding examples use Python and PySpark so some experience with Python is important.