Learning Hub
← DevOps & Infrastructure

Alerts, service-level objectives, and incident response

6 min readΒ·Updated 2026-09-09

An SLO states a reliability target for a user-visible service indicator; alerts should fire on meaningful risk to that objective; incident response restores service, communicates, and learns without blame.

An SLO states a reliability target for a user-visible service indicator; alerts should fire on meaningful risk to that objective; incident response restores service, communicates, and learns without blame.

At a glance

Question Practical answer
When is it useful? A checkout service targets 99.9% successful requests over 30 days and pages when error-budget burn predicts the objective will be exhausted too quickly.
What should you do? Define one indicator and SLO, calculate its error budget, write a symptom-based alert, and run a tabletop incident from detection through review.
How do you know it worked? The alert is actionable, has an owner and runbook, avoids paging on harmless internal noise, and the exercise produces tracked improvements.
Common failure A target of 100% leaves no room for change or recovery and often creates noisy alerts rather than better reliability.
flowchart LR
  A[Question] --> B[Alerts, service-level objectives, and inci]
  B --> C[Small example]
  C --> D[Evidence]

The important idea is not to stop at a definition: connect the concept to a small example and observable evidence.

Worked example

A checkout service targets 99.9% successful requests over 30 days and pages when error-budget burn predicts the objective will be exhausted too quickly.

Before acting, write the success signal. Change one condition at a time, observe the result, and record assumptions. For Alerts, service-level objectives, and incident response, this separates what you know from what you are merely guessing.

Practice in 20–30 minutes

Goal: Define one indicator and SLO, calculate its error budget, write a symptom-based alert, and run a tabletop incident from detection through review.

  1. Record the starting state and your prediction.
  2. Implement the smallest version without adding unnecessary tools.
  3. Change exactly one input or constraint and repeat.
  4. Save a command, screenshot, output, or checklist as evidence.

Expected result: The alert is actionable, has an owner and runbook, avoids paging on harmless internal noise, and the exercise produces tracked improvements.

What can go wrong

A target of 100% leaves no room for change or recovery and often creates noisy alerts rather than better reliability.

When the result differs from your prediction, do not change many things at once. Check inputs, versions, environment, permissions, and logs, then repeat from the smallest example.

Definition of done

  • I can explain the concept in my own words.
  • I completed the small example and kept evidence.
  • I know one failure mode and how to check it.
  • Someone else can repeat the work without guessing missing steps.

Go deeper

Use the linked resource or repository at the end of the page when you need a full implementation. Check current versions before applying commands to a real project.