As Big Data transitions from a buzzword to a corporate standard, understanding the tools that power it is essential. This advanced-level course moves beyond the basics to explore PySpark, the Python interface for the lightning-fast Spark framework.
This curriculum is designed for those looking to leverage Spark's superior speed—outperforming Hadoop significantly both in-memory and on-disk. You will dive into high-level libraries including:
To ground these technical concepts, the course utilizes diverse datasets, ranging from literary works by William Shakespeare to FIFA 2018 statistics and complex genomic clustering. It is an ideal choice for data professionals seeking an in-depth understanding of large-scale data analysis applications.
There's been a lot of buzz about Big Data over the past few years, and it's finally become mainstream for many companies. But what is this Big Data? This course covers the fundamentals of Big Data via PySpark. Spark is a "lightning fast cluster computing" framework for Big Data. It provides a general data processing platform engine and lets you run programs up to 100x faster in memory, or 10x faster on disk, than Hadoop. You’ll use PySpark, a Python package for Spark programming and its powerful, higher-level libraries such as SparkSQL, MLlib (for machine learning), etc. You will explore the works of William Shakespeare, analyze Fifa 2018 data and perform clustering on genomic datasets. At the end of this course, you will have gained an in-depth understanding of PySpark and its application to general Big Data analysis.
Access
Advertisement
Included with a DataCamp subscription
This course is included with a Subscription subscription.
View on DataCamp →No coupon right now. We'll tell you when there is one.