Apache Spark on Amazon EMR teaches attendees how to use Apache Spark within Amazon EMR to solve real-world data science problems. Students learn to produce valuable insights across a wide range of scenarios. Data science draws insights from both structured and unstructured data. Apache Spark and AWS rank among the most important, widely used tools in that field.
Day one focuses on Amazon EMR basics, including an overview of related services and data acquisition, scrubbing, and manipulation. Attendees then move into basic application design and deployment. An operations module follows, covering day-to-day system management and cost control. The day closes with a module on EMR’s development tools, where students work directly in EMR Studio and notebooks. It also gives a general overview of data science applications and the analytics and machine learning processes typically used. Attendees examine a number of practical use cases during class and lab sessions.
Day two focuses on Spark applications and how they best operate in the Amazon EMR environment. Students review high-level Spark features such as Stream Processing, focusing on how these features operate in the cloud. A machine learning module follows, giving attendees hands-on practice training models on EMR. An interactive SQL module then shows how to query large datasets directly against a Spark cluster. Day two concludes with a focus on securing Spark pipelines, along with related cost information and tracking.
Day three moves further into the cloud, integrating with purpose-driven services such as AWS Glue to improve complex pipelines. The day then turns to EMR Serverless so attendees can run Spark jobs without managing cluster infrastructure directly. AWS SageMaker also gets significant focus across two modules, starting with basic ML pipelines and progressing to intermediate SageMaker techniques. This helps data scientists build and maintain complex applications and pipelines involving Spark.
Who Should Attend
Application developers, analysts and data scientists
What Attendees Will Learn
Upon completing Apache Spark on Amazon EMR, attendees will be able to:
- Explain Amazon EMR fundamentals and design basic Spark applications
- Manage EMR system operations and cost
- Use EMR development tools, including Studio and notebooks
- Apply Spark stream processing in the cloud
- Apply Spark machine learning and interactive SQL
- Secure Spark pipelines and track operational costs
- Integrate Spark with AWS Glue and EMR Serverless
- Build Spark ML pipelines using AWS SageMaker
Prerequisites
Basic Linux command line skills are valuable but not required.