Observability Stack

Transform black-box systems into observable infrastructure. You will deploy Prometheus for metrics, Grafana for visualization, and Loki for log aggregation to gain deep insights into system performance.

DevOps EngineerDevOpsLinux

Syllabus

no data

Introduction

An observability pipeline turns host activity into metrics, searchable logs, and actionable alert state. This challenge-based project builds that pipeline locally with Prometheus, Grafana, Loki, Promtail, Node Exporter, and Alertmanager.

You will connect each component through configuration files and live APIs: scrape host metrics, provision a Grafana data source, ship an active system log to Loki, and trigger an InstanceDown alert by stopping Node Exporter. The result is a functioning single-host stack whose data flow you can verify end to end.

What You Will Learn

  • Run Node Exporter on port 9100 and configure Prometheus on 9090 to scrape it as a healthy target
  • Provision a Grafana Prometheus data source in YAML and verify it through Grafana's API on port 8080
  • Start Loki on port 3100 and configure Promtail to follow an active local system log
  • Confirm log ingestion through Loki label APIs and distinguish collector responsibilities from storage responsibilities
  • Run Alertmanager on port 9093 and connect Prometheus to it through alertmanagers configuration
  • Define InstanceDown for up == 0, stop Node Exporter, and verify the alert reaches the firing state

Who This Course Is For

This project is for DevOps and SRE learners ready to assemble a local observability pipeline from configuration and service-level checks.

Prerequisites: Familiarity with Linux services and processes, YAML, HTTP APIs, ports, Prometheus scrape concepts, basic log paths, and curl; this is an assessment-style project.

Learning environment: A browser-accessible single-node Linux host with sudo, Prometheus, Node Exporter, Grafana, downloadable Loki, Promtail, and Alertmanager binaries, an active local log, and local HTTP access; no cloud account is required.

Frequently Asked Questions

Will I build Grafana dashboards and panels?

No. The Grafana phase provisions and verifies a Prometheus data source as code. Dashboard design, panels, variables, and Grafana alert rules are outside this project.

Are logs collected from multiple servers?

No. Promtail follows /var/log/syslog or another active log on the same host and pushes it to the local Loki process. Distributed agents, retention, and object storage are not configured.

Does Alertmanager send email, Slack, or webhook notifications?

No. Prometheus routes the alert to local Alertmanager, and you verify the firing state in the API or UI. External receivers and notification delivery are not part of the challenge.

How is the failure test performed and verified?

You stop Node Exporter so Prometheus reports up == 0. Validation confirms the exporter is stopped and the Prometheus alerts API contains InstanceDown in the firing state.

Teacher

labby
Labby
Labby is the LabEx teacher.