Data Cleaning and Wrangling
Coursera
Before you can work with data you have to get some. This course will cover the basic ways that data can be obtained. The course will cover obtaining data from the web, from APIs, from databases and from colleagues in various formats. It will also cover the basics of data cleaning and how to make data “tidy”. Tidy data dramatically speed downstream data analysis tasks. The course will also cover the components of a complete data set including raw data, processing instructions, codebooks, and processed data. The course will cover the basics needed for collecting, cleaning, and sharing data.
More resources on Data Wrangling
Tidyverse
The official home of R's tidyverse packages, with documentation and guides for dplyr, tidyr, and readr. Covers the shared design philosophy so you can reshape, filter, and join data frames consistently.
Pandas Documentation
The official pandas user guide, organized by task: indexing, merging, reshaping, grouping, and handling missing values. Reading it end to end gives you command of DataFrame operations and the idioms experienced Python analysts rely on.
Data Wrangling with Pandas
A single-sitting video walkthrough of pandas by Keith Galli, working through loading CSV files, filtering rows, aggregating, and saving results. By the end you can perform the common cleaning steps on a real dataset yourself.
Cleaning Data with Python
Learn data-wrangling essentials! This DataCamp course teaches you how to clean messy data with Python for effective analysis.
Data Analysis with Python and Pandas Tutorial
First video in Corey Schafer's pandas series, walking through installing pandas and Jupyter, loading a CSV survey dataset into a DataFrame, and inspecting its shape, columns, and rows. Viewers finish able to set up a working environment for data analysis in Python.
Data Visualization Full Course - Learn Data Visualization in 7 Hours
A long-form Python course covering NumPy arrays, Pandas dataframes, and plotting with Matplotlib and Seaborn. Work through it to load, clean and summarize tabular datasets, then chart the results in code.