Designing Machine Learning Systems
Covers the full lifecycle of production machine learning: framing business problems, data engineering and sampling, feature engineering, model evaluation, deployment, monitoring for distribution shift, and the team structures that keep systems running reliably.
More resources on MLOps
mlops.community
MLOps.community is a community-driven hub for learning and practicing MLOps, curating articles, tutorials, talks, case studies, tools, and practical resources from practitioners. It also hosts discussions and events that connect you with ongoing conversations and guidance in the MLOps community.
LLMOps
Builds an end-to-end tuning pipeline on Google Cloud: pulling and transforming training data in BigQuery, running a supervised fine-tuning pipeline with Kubeflow, then deploying and safety-checking a chatbot that answers Python questions.
Made With ML
Goku Mohandas's free course builds one production machine learning application end to end: system design, data preparation, training with experiment tracking, testing code and models, then serving, CI/CD, and monitoring. Now maintained under Anyscale.
Full Stack Deep Learning
Learn MLOps best practices to build, deploy, and scale robust machine learning systems with this comprehensive Full Stack Deep Learning course.
Building Machine Learning Pipelines
Hapke and Nelson construct an automated TensorFlow Extended pipeline stage by stage: data ingestion and validation, preprocessing, training, model analysis, and deployment with TensorFlow Serving, then orchestrate the whole thing through Apache Beam, Airflow, and Kubeflow Pipelines.
Machine Learning Engineering for Production
Opening course of DeepLearning.AI's Machine Learning Engineering for Production specialization, taught by Andrew Ng. Covers the production lifecycle, deployment patterns and monitoring, error analysis and baselines, and the data definition work behind consistent labels.