Introduction to Apache Spark and SparkSQL

}
2 Days

Available On-Site

Available Virtually

Contact Us for Open Enrollment
f

Customizable

Introduction to Apache Spark and SparkSQL gives attendees a hands-on introduction to Spark and SparkSQL across two days. The first day covers the fundamentals of the Spark distributed computing engine, including installing and running Spark. Attendees also work with Resilient Distributed Datasets (RDDs) and ingest and process file-based data. The second day shifts focus to SparkSQL, covering the Hive Metastore, DataFrames, querying external databases over JDBC/ODBC, and writing user-defined functions. Hands-on labs accompany every module, giving attendees practical, immediately usable skills.

RDD labs move from basic data analysis into key/value operations. Attendees transform, aggregate, group, join, and sort keyed datasets, core skills for any distributed data pipeline. Day one closes with file-based I/O labs that ingest and process data directly from disk. This gives attendees a feel for how Spark partitions and reads large datasets in parallel.

The second day opens by querying unstructured Hadoop data through the Hive Metastore. It then moves into manipulating DataFrames alongside RDDs to compare the two programming models directly. A dedicated module covers querying external databases over JDBC and ODBC sources, followed by writing and applying user-defined functions to extend built-in SQL capabilities. Each lab builds on the previous one, so attendees leave with a working, end-to-end pipeline rather than isolated exercises.

Who Should Attend

Application developers, analysts and data scientists

What Attendees Will Learn

Upon completing Introduction to Apache Spark and SparkSQL, attendees will be able to:

  • Install and configure Apache Spark and run Spark applications
  • Work with Resilient Distributed Datasets (RDDs) to transform, aggregate, group, join, and sort data
  • Ingest and process file-based data with Spark I/O
  • Query structured and unstructured data using SparkSQL and the Hive Metastore
  • Manipulate data using Spark DataFrames
  • Query external databases through JDBC/ODBC sources from Spark
  • Write and use Spark user-defined functions (UDFs)

Prerequisites

Basic Linux command line skills are helpful. The coding examples use Python and PySpark so some experience with Python is important.

Delivery

Available for Instructor-Led (ILT) in-person/onsite training or Virtual Instructor-Led training (VILT) delivery.

Each attendee will require the ability to ssh into a cloud hosted virtual machine (provided with the course). In environments where SSH is not possible, local lab VMs or browser accessible lab systems can be provided. For web-based delivery, participants require an Internet-connected computer capable of teleconferencing.

Related Instructor-Led (ILT & VILT) Training Courses

If you are interested in other Cloud Native, AI, programming, or other courses, check out the full course list or search our entire catalog:

Secret Link