Apache Spark on Amazon EMR

}
3 Days

Available On-Site

Available Virtually

Contact Us for Open Enrollment
f

Customizable

Apache Spark on Amazon EMR teaches attendees how to use Apache Spark within Amazon EMR to solve real-world data science problems. Students learn to produce valuable insights across a wide range of scenarios. Data science draws insights from both structured and unstructured data. Apache Spark and AWS rank among the most important, widely used tools in that field.

Day one focuses on Amazon EMR basics, including an overview of related services and data acquisition, scrubbing, and manipulation. Attendees then move into basic application design and deployment. An operations module follows, covering day-to-day system management and cost control. The day closes with a module on EMR’s development tools, where students work directly in EMR Studio and notebooks. It also gives a general overview of data science applications and the analytics and machine learning processes typically used. Attendees examine a number of practical use cases during class and lab sessions.

Day two focuses on Spark applications and how they best operate in the Amazon EMR environment. Students review high-level Spark features such as Stream Processing, focusing on how these features operate in the cloud. A machine learning module follows, giving attendees hands-on practice training models on EMR. An interactive SQL module then shows how to query large datasets directly against a Spark cluster. Day two concludes with a focus on securing Spark pipelines, along with related cost information and tracking.

Day three moves further into the cloud, integrating with purpose-driven services such as AWS Glue to improve complex pipelines. The day then turns to EMR Serverless so attendees can run Spark jobs without managing cluster infrastructure directly. AWS SageMaker also gets significant focus across two modules, starting with basic ML pipelines and progressing to intermediate SageMaker techniques. This helps data scientists build and maintain complex applications and pipelines involving Spark.

Who Should Attend

Application developers, analysts and data scientists

What Attendees Will Learn

Upon completing Apache Spark on Amazon EMR, attendees will be able to:

  • Explain Amazon EMR fundamentals and design basic Spark applications
  • Manage EMR system operations and cost
  • Use EMR development tools, including Studio and notebooks
  • Apply Spark stream processing in the cloud
  • Apply Spark machine learning and interactive SQL
  • Secure Spark pipelines and track operational costs
  • Integrate Spark with AWS Glue and EMR Serverless
  • Build Spark ML pipelines using AWS SageMaker

Prerequisites

Basic Linux command line skills are valuable but not required.

Delivery

Available for Instructor-Led (ILT) in-person/onsite training or Virtual Instructor-Led training (VILT) delivery.

Each attendee will require the ability to ssh into a cloud hosted virtual machine (provided with the course). In environments where SSH is not possible, local lab VMs or browser accessible lab systems can be provided. For web-based delivery, participants require an Internet-connected computer capable of teleconferencing.

Related Instructor-Led (ILT & VILT) Training Courses

If you are interested in other Cloud Native, AI, programming, or other courses, check out the full course list or search our entire catalog:

Secret Link