Data Engineering and DataOps Foundation gives participants a solid, two-day foundation in data engineering and DataOps principles. Attendees gain hands-on experience with essential data engineering tools, then learn how to apply DataOps methods that streamline pipelines and protect data quality. The course opens by grounding attendees in the role of data engineering, differentiating it from data science, and exploring key concepts such as push versus pull ingestion patterns, before a first lab sets up a working development environment for the rest of the week.
From there, attendees dig into data ingestion and storage, comparing batch and real-time patterns and working hands-on with Apache Kafka to build a real ingestion pipeline. The next module covers data transformation, applying filtering, aggregating, and joining techniques with Apache Spark and tuning pipelines for performance. Day one closes with data storage and warehousing, comparing relational and NoSQL options and designing a warehouse with Amazon Redshift.
Day two shifts to DevOps principles for data management. Attendees start with the DataOps lifecycle and the metrics teams use to measure pipeline health, then build CI/CD pipelines for data using Jenkins. Data quality and monitoring follow, implementing automated checks with Apache Airflow, and the course closes with data governance and compliance, applying policy, auditing, and lineage tracking hands-on with Apache Atlas. By the end, attendees are ready to put this expertise to work in data-driven organizations.
Who Should Attend
Developers, Data Engineers, IT Professionals, Analysts, Data Scientists, DevOps Professionals, MLOps Professionals
What Attendees Will Learn
Upon completing Data Engineering and DataOps Foundation, attendees will be able to:
- Explain the role of data engineering and how it differs from data science
- Build data ingestion pipelines using batch and real-time patterns with Apache Kafka
- Transform and process data at scale using Apache Spark
- Design data warehouses using Amazon Redshift
- Apply DataOps principles, metrics, and KPIs to data pipelines
- Build CI/CD pipelines for data using Jenkins
- Implement data quality checks and monitoring using Apache Airflow
- Apply data governance and compliance practices using Apache Atlas
Prerequisites
Basic understanding of data concepts. Familiarity with at least one programming language (e.g. C, C++, JavaScript, Python). Comfortable with using a command-line interface.
These prerequisites will help participants engage effectively with the course material and hands-on labs, making the learning experience more rewarding.