
UdemyThis course sets up a full analytics pipeline for working with Retrosheet's game-by-game and play-by-play baseball data, starting with installing a virtual Linux machine and the Chadwick software used to extract that data. From there, it moves into filtering with dplyr in R and visualizing results with ggplot.
The first part, covering the Linux VM and Chadwick setup, has no prerequisites, but the dplyr-focused second half assumes that background — obtainable through the instructor's companion course, "Baseball Database Queries with SQL and dplyr." At a relaxed pace, it's designed to take two to three weeks.
This course is for those interested in doing baseball analytics with the Retrosheet game-by-game and play-by-play data. The main tools for working with such data are in the Chadwick software. We install a virtual Linux machine, on which we will install the Chadwick software. We will then learn how to extract baseball data with the Chadwick software, how to further filter the data with dplyr in R, and how to plot our results with ggplot.
For the first part of the course, in which we install the virtual Linux machine and learn how to work with the Chadwick software, there are no prerequisites. To follow the second part of the course, knowledge of dplyr is necessary. This can be obtained through my course "Baseball Database Queries with SQL and dplyr".
At a relaxed pace, the course should take two to three weeks to complete.
Advertisement
Price
This course is free to enrol.
Enrol free on Udemy →More in Development







