The Korshub Take
This advanced course teaches you how to scale machine learning algorithms to massive datasets using Apache Spark and PySpark. It covers Spark's distributed compute capabilities, ML pipelines, and key classification and regression algorithms.
Who It's For
Data engineers and data scientists looking to build big data machine learning models and pipelines across distributed clusters.
Key Takeaways
- Distributed ML Pipelines: Ingest big data into Spark and build end-to-end Machine Learning pipelines.
- Decision Trees & Recursive Partitioning: Implement decision trees using recursive partitioning to split data and evaluate node split information.
- Spark Regressors & Classifiers: Master PySpark implementations of Linear Regression and Logistic Regression to refine predictive modeling.