Advanced Monitoring and Logging
Coursera
Part of the AWS Certified DevOps Engineer Professional exam-prep specialization, this course covers monitoring, logging, and governance services in AWS, including CloudWatch and CloudTrail, across three modules of video lessons, hands-on projects, and quizzes.
More resources on Monitoring & Logging
Monitoring and Alerting with Prometheus
Learn Prometheus monitoring and alerting in this comprehensive course! Master metrics collection, visualization, and effective alerting strategies.
Practical Monitoring
Argues for designing monitoring around user-facing symptoms rather than tool checklists, covering alert design, on-call rotation, statistics for metrics, and instrumenting the network and application tiers. Useful for replacing noisy, check-based Nagios-era setups.
prometheus.io
Prometheus.io is the official site for Prometheus, an open-source monitoring and alerting toolkit focused on metrics collection and querying. It provides comprehensive docs, tutorials, installation guides, architecture overview, PromQL references, exporters, and integration guides for Alertmanager and visualization with Grafana.
grafana.com
Grafana is an observability platform for building real-time dashboards of metrics, logs, and traces. The site provides dashboards, plugins, tutorials, and documentation to visualize data from sources like Prometheus, Loki, and Tempo.
Observability Engineering
Defines observability as the ability to ask new questions of production systems, and explains structured events, high-cardinality data, service level objectives, and the shift away from dashboards. Includes guidance on adoption inside existing teams.