Learning Hub
← DevOps & Infrastructure

Prometheus and Grafana starter lab

6 min readΒ·Updated 2026-09-09

Prometheus scrapes and stores labeled time-series metrics; Grafana queries data sources and presents dashboards and alerts. A useful lab begins with a few service-level signals, not hundreds of panels.

Prometheus scrapes and stores labeled time-series metrics; Grafana queries data sources and presents dashboards and alerts. A useful lab begins with a few service-level signals, not hundreds of panels.

At a glance

Question Practical answer
When is it useful? A sample HTTP service exports request count, error count, and duration; Prometheus scrapes it and Grafana graphs rate, error ratio, and latency percentiles.
What should you do? Run a local sample stack, generate normal and failing traffic, query one metric directly, then build a three-panel dashboard with units and descriptions.
How do you know it worked? Targets are up, queries return expected label sets, the dashboard visibly changes under failure, and an alert condition matches a user-impacting symptom.
Common failure Unbounded labels such as user ID or request ID create dangerous cardinality; keep high-cardinality detail in logs or traces.
flowchart LR
  A[Question] --> B[Prometheus and Grafana starter lab]
  B --> C[Small example]
  C --> D[Evidence]

The important idea is not to stop at a definition: connect the concept to a small example and observable evidence.

Worked example

A sample HTTP service exports request count, error count, and duration; Prometheus scrapes it and Grafana graphs rate, error ratio, and latency percentiles.

Before acting, write the success signal. Change one condition at a time, observe the result, and record assumptions. For Prometheus and Grafana starter lab, this separates what you know from what you are merely guessing.

Practice in 20–30 minutes

Goal: Run a local sample stack, generate normal and failing traffic, query one metric directly, then build a three-panel dashboard with units and descriptions.

  1. Record the starting state and your prediction.
  2. Implement the smallest version without adding unnecessary tools.
  3. Change exactly one input or constraint and repeat.
  4. Save a command, screenshot, output, or checklist as evidence.

Expected result: Targets are up, queries return expected label sets, the dashboard visibly changes under failure, and an alert condition matches a user-impacting symptom.

What can go wrong

Unbounded labels such as user ID or request ID create dangerous cardinality; keep high-cardinality detail in logs or traces.

When the result differs from your prediction, do not change many things at once. Check inputs, versions, environment, permissions, and logs, then repeat from the smallest example.

Definition of done

  • I can explain the concept in my own words.
  • I completed the small example and kept evidence.
  • I know one failure mode and how to check it.
  • Someone else can repeat the work without guessing missing steps.

Go deeper

Use the linked resource or repository at the end of the page when you need a full implementation. Check current versions before applying commands to a real project.