
edXModern data systems are built on a distributed stack: storage (Hadoop), processing (Spark), cloud data warehouses (Snowflake), and ML lifecycle management (MLflow). This course teaches you how they fit together and when to use each.
You'll code in PySpark, master the platforms' optimization techniques, and learn to think like a data engineer: building reliable pipelines, applying DevOps discipline to data workflows, and treating the ML lifecycle as end-to-end.
For aspiring data engineers, data scientists expanding into engineering, and software developers transitioning into data roles with hands-on, practical skill-building.
In this course, you will:
This course is designed for learners who want to pursue or advance their career in data science or data engineering, or for software developers or engineers who want to grow their data management skill set. With quizzes to test your knowledge throughout, this comprehensive course will help guide your learning journey to become a proficient data engineer, ready to tackle the challenges of today's data-driven world.
Price
Advertisement
No active coupon right now
View Course →No coupon right now. We'll tell you when there is one.
More in Software Engineering




Tracking since 1 Aug— not enough history yet to tell you whether today's price is any good. Watch the course and we'll tell you when it drops.
This is what we recorded in US pricing — not every price this course has ever had, and prices differ by country.