All topics
DevOps · Learning hub

Alertmanager notes for developers

Master Alertmanager with a curated set of 1 developer notes — core concepts, patterns, and interview prep. Maintained by the DevRecall team.

Save this stack to your DevRecallTest yourself — Alertmanager quizMore DevOps notes
Alertmanager

Alertmanager Essentials

Alertmanager Essentials Alertmanager is Prometheus's companion service for turning firing alerts into actual notifications. Prometheus's job stops at deciding a

Alertmanager Essentials

Alertmanager is Prometheus's companion service for turning firing alerts into actual notifications. Prometheus's job stops at deciding an alert should fire and pushing it over; Alertmanager takes it from there — deduplicating alerts coming from multiple Prometheus instances, grouping related ones into a single notification, silencing known issues, suppressing noisy alerts that are implied by a more important one, and finally routing each alert to the right receiver (Slack, PagerDuty, email, a webhook).

Config Basics

Everything lives in alertmanager.yml: a single top-level route tree, a list of named receivers, and optional inhibit_rules.

global:
  resolve_timeout: 5m
  slack_api_url: 'https://hooks.slack.com/services/XXX/YYY/ZZZ'

route:
  receiver: 'default-slack'
  group_by: ['alertname', 'cluster']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

receivers:
  - name: 'default-slack'
    slack_configs:
      - channel: '#alerts'
        send_resolved: true
        title: '{{ .CommonAnnotations.summary }}'
        text: '{{ .CommonAnnotations.description }}'

inhibit_rules:
  - source_matchers: ['severity = critical']
    target_matchers: ['severity = warning']
    equal: ['alertname', 'cluster', 'service']

The Routing Tree

Routes form a tree, not a flat list. Every alert enters at the root route and walks down through routes children whose matchers match its labels. By default a matching child route is used exclusively — the alert stops descending further siblings — unless continue: true tells it to keep evaluating later routes too (useful when you want both a team channel and a global catch-all to fire).

route:
  receiver: 'default-slack'
  group_by: ['alertname', 'cluster', 'service']
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers:
        - severity = "critical"
      receiver: 'pagerduty-oncall'
      group_wait: 10s
      repeat_interval: 1h
      continue: true      # also fall through to the team route below

    - matchers:
        - team = "payments"
      receiver: 'payments-slack'

    - matchers:
        - alertname = "Watchdog"
      receiver: 'null'    # discard heartbeat alerts entirely

receivers:
  - name: 'null'
  - name: 'pagerduty-oncall'
    pagerduty_configs:
      - routing_key: '<integration-key>'
  - name: 'payments-slack'
    slack_configs:
      - channel: '#payments-alerts'

Routing to a receiver literally named null (with no configs) is the standard way to intentionally swallow an alert — like a synthetic `Watchdog` heartbeat that exists only to confirm the pipeline itself is alive.

Grouping, Inhibition & Silencing

These three mechanisms exist to stop a single incident from becoming dozens of separate pages.

  • Grouping (group_by) bundles alerts sharing the listed labels into one notification — e.g. 20 pods failing health checks at once becomes one Slack message listing 20 instances, not 20 messages.

  • Inhibition suppresses a lower-priority alert when a related higher-priority one is already firing — e.g. don't page about elevated latency on a service that's already `InstanceDown`, since the latency alert is a symptom, not new information.

  • Silencing is a manually created, time-bound mute matched by label — used during planned maintenance so real work doesn't trigger pages for a known, expected condition.

# Create a silence via amtool for a 2-hour maintenance window
amtool silence add \
  alertname="HighLatency" cluster="eu-west-1" \
  --duration=2h \
  --comment="Planned DB migration - jira PROJ-1234" \
  --author="lubomir"

# List active silences, expire one early
amtool silence query
amtool silence expire <silence-id>

# Check where an alert with given labels would route, without firing it
amtool config routes test --config.file=alertmanager.yml \
  severity=critical team=payments

Timing Parameters

These three settings are the ones people get wrong most often, because their names sound interchangeable but they control very different things:

  • `group_wait` — how long to wait after the *first* alert in a new group before sending the initial notification, to let a few more related alerts arrive and land in the same message.

  • `group_interval` — how long to wait before sending a notification about *new* alerts added to an *already-notified* group.

  • `repeat_interval` — how long to wait before re-sending a notification for an alert that is still firing and unchanged — this is your re-page cadence for an unresolved incident.

Practical Tips & Gotchas

  • Run Alertmanager as a cluster (3 replicas, gossiping over `--cluster.peer`) — a single instance is a silent single point of failure for every alert in your stack.

  • `matchers` (the current syntax, e.g. `severity = "critical"`) replaced the older `match`/`match_re` maps — new configs should use matchers; both still work but shouldn't be mixed carelessly.

  • A route with no matching child falls back to using its own (or the root's) receiver — always make sure the root route's receiver is something a human actually monitors, not silence-by-accident.

  • `send_resolved: true` on a receiver is what sends the follow-up "this is now resolved" message — without it, alerts appear to fire and then just vanish with no closure.

  • Inhibition only works within a single Alertmanager instance's view of currently firing alerts — it does not reach across Alertmanager clusters or delay alerts, it purely suppresses notification for alerts already flagged as firing.

Keep your Alertmanager knowledge sharp.

Save this stack to your personal DevRecall — add your own notes, track what you're learning, and share what you know with the community.

Get started — free forever