Introduction to Data Science in Python
University of Michigan
This course will introduce the learner to the basics of the python programming environment, including fundamental python programming techniques such as lambdas, reading and manipulating csv files, and the numpy library. The course will introduce data manipulation and cleaning techniques using the popular python pandas data science library and introduce the abstraction of the Series and DataFrame as the central data structures for data analysis, along with tutorials on how to use functions such as groupby, merge, and pivot tables effectively. By the end of this course, students will be able to take tabular data, clean it, manipulate it, and run basic inferential statistical analyses. This course should be taken before any of the other Applied Data Science with Python courses: Applied Plotting, Charting & Data Representation in Python, Applied Machine Learning in Python, Applied Text Mining in Python, Applied Social Network Analysis in Python.
More resources on Data Science
Data Analysis with Python
Learn data analysis with Python! This free course covers NumPy, Pandas, data cleaning, and visualization. Start your data science journey today!
Practical Deep Learning for Coders
fast.ai's free course teaching deep learning top-down: you train working image, text, and tabular models in the first lessons, then work back to the underlying mechanics. Assumes about a year of coding experience, uses PyTorch and the fastai library.
StatQuest with Josh Starmer - Neural Networks videos
Josh Starmer's StatQuest provides clear, concise, and often humorous explanations of statistical and machine learning concepts. His videos on neural networks break down complex ideas into easily digestible 'quests.'
Kaggle Learn
Free interactive micro-courses from Kaggle, each a few hours of short lessons with in-browser coding exercises. Topics include Python, pandas, data visualization, SQL, feature engineering and introductory machine learning, giving learners working code skills for basic data analysis and modelling.
Harvard University's CS109 Data Science
Harvard's two-semester data science course with public lectures, labs and homework. CS109A covers data collection, cleaning, visualization and regression and classification models; CS109B adds deep learning and probabilistic methods. Students learn to run a full analysis in Python.
Practical Statistics for Data Scientists
A concise statistics book written for programmers that covers sampling, experimental design, A/B testing, regression, classification, and resampling, with R code throughout. Readers learn which statistical ideas matter for data science and how sampling and study design shape the conclusions data can support.