
This advanced course addresses the challenges of preparing massive, distributed datasets for machine learning using PySpark. It covers the gritty, practical details of data wrangling and scalable feature engineering when working with Big Data.
Data engineers, Big Data practitioners, and data scientists using PySpark to prepare large-scale datasets for machine learning models.
The real world is messy and your job is to make sense of it. Toy datasets like MTCars and Iris are the result of careful curation and cleaning, even so the data needs to be transformed for it to be useful for powerful machine learning algorithms to extract meaning, forecast, classify or cluster. This course will cover the gritty details that data scientists are spending 70-80% of their time on; data wrangling and feature engineering. With size of datasets now becoming ever larger, let's use PySpark to cut this Big Data problem down to size!
Price
This course is free to enrol.
Enrol free on DataCamp →Advertisement

More in Data Science & AI