
Would you like to schedule this training on a specific date? Contact us by email or via the contact form.
The AWS Data Engineer training course tackles what happens between the raw data and the dashboard: ingestion, transformation, storage — and above all the part nobody talks about: making sure the pipeline still runs in six months, that you know where every figure came from, and what it cost.
Over 3 days, you build a complete chain on AWS: batch and streaming ingestion with Kinesis and DMS, data lake storage on S3 with the columnar formats that change everything (Parquet, partitioning, Iceberg), transformation with Glue and EMR, warehousing with Redshift, ad hoc querying with Athena. Orchestration, data quality and access governance with Lake Formation get full treatment, because that is where data projects fail — rarely on the transformation logic itself.
The programme follows the syllabus of the AWS Certified Data Engineer – Associate (DEA-C01) certification, which replaces the retired Data Analytics Specialty.
If AWS is new to you, our AWS Cloud Practitioner Training Course lays the groundwork, and our AWS Essential Services Training Course covers the S3 and compute services used here. On the streaming side, our Apache Kafka Training Course illuminates the concepts you will meet again in Kinesis. To compare ecosystems, see our Azure Data Lake, Synapse, Data Factory and Spark Training Course and our Azure Databricks Training Course. Finally, running pipelines in production extends our AWS DevOps Engineer Training Course, and keeping the bill under control our AWS FinOps Training Course.
Sessions are guaranteed from a single registrant (except in cases of force majeure). A preliminary discussion takes place between the participant and/or a company representative to fully take into account the participant’s profile. Assessment: quizzes, role-play and practical exercises. A certificate of completion is issued at the end of the training. This training prepares you for the AWS Certified Data Engineer – Associate (DEA-C01) certification (exam not included). This training is part of our Cloud Computing Training Courses catalogue. Discover our other cloud training courses to master architectures, services and best practices on AWS, Azure and GCP. To get started on AWS, we recommend our AWS Cloud Practitioner Training Course as a prerequisite, which lays the foundations of the Amazon Cloud.
By the end of the training, participants will be able to:
It is not blocking, but it helps a great deal in the Glue and EMR modules, where Spark jobs are written in Python. The genuinely indispensable prerequisite is SQL. If Python is unfamiliar, say so during the preliminary discussion: we adapt the labs to stay with the visual approach of Glue Studio.
AWS has retired the Data Analytics Specialty and replaced it with Data Engineer – Associate (DEA-C01), which focuses on building and operating pipelines rather than on analysis. This course follows the new syllabus — which is also why a title such as “Data Warehousing”, still found in some catalogues, corresponds to a dated offering.
Yes, and the answer depends mostly on your query profile. Athena suits ad hoc work and irregular volume, billed by data scanned. Redshift makes sense for repeated, predictable analytical workloads. Both are implemented during the course, on the same data, so the comparison is concrete.
It is possible and often desirable, within the limits of what you can take out of your environment. The preliminary discussion is there to frame this; failing that, we provide realistic datasets.