prometheus Runbook¶
Metadata¶
| Field | Value |
|---|---|
| Service | prometheus |
| Criticality | Tier 1 |
| Owner | Platform / Observability owner |
| Namespace | prometheus, mimir |
| Clusters | jls |
| Last validated | 2026-08-24 |
| Related service page | ../services/prometheus.md |
Trigger Conditions¶
- Alerting stops.
- Scrape targets are missing.
- TSDB storage is full or unavailable.
- External Prometheus or Alertmanager endpoints fail.
1. Health Checks¶
kubectl -n prometheus get pods,svc,pvc,ingressroute
kubectl -n prometheus logs deploy/prometheus-server --tail=200
kubectl -n prometheus logs statefulset/prometheus-alertmanager --tail=100
Verify the metrics pipeline end to end (all three must return data):
# 1. Prometheus scrapes Mimir (feeds cortex_* metrics and MimirUnhealthy)
kubectl -n prometheus exec deploy/prometheus-server -c prometheus-server -- \
wget -qO- 'http://localhost:9090/api/v1/query?query=up{job="mimir-monolith"}'
# 2. Remote write to Mimir is not failing
kubectl -n prometheus exec deploy/prometheus-server -c prometheus-server -- \
wget -qO- 'http://localhost:9090/api/v1/query?query=sum(rate(prometheus_remote_storage_failed_samples_total[5m]))'
# 3. The single Alertmanager receives alerts (Watchdog always firing)
kubectl -n prometheus exec deploy/prometheus-server -c prometheus-server -- \
wget -qO- 'http://localhost:9090/api/v1/query?query=ALERTS{alertname="Watchdog"}'
Mimir (namespace mimir) must show 1 replica; its ruler logs should not show
alertmanager dispatch errors.
2. Troubleshooting Workflows¶
Check scraping, alerting, and storage:
kubectl -n prometheus describe statefulset prometheus-server
kubectl -n prometheus get configmap
kubectl -n prometheus get secret
Look for rule errors, remote-write failures, and disk pressure.
3. Disaster Recovery¶
- Restore TSDB storage and alerting secrets.
- Reconcile the overlay (Fleet re-applies from Git, or
kubectl kustomize prometheus/overlays/jls --enable-helm | kubectl apply -f -). - Confirm scrape targets recover.
- Trigger and validate a test alert path.
4. Scaling and Resource Management¶
Increase storage, memory, or retention controls in Git when scrape volume outgrows the current profile.
5. Maintenance Procedures¶
- Rotate alert receiver and remote-write secrets.
- Review rule changes before merges.
- Plan retention changes carefully because TSDB rewrites can be expensive.
6. Rollback Strategy¶
- Revert the overlay to the previous working revision.
- Restore the prior TSDB snapshot if a config or upgrade damaged the data path.
7. Post-Incident Actions¶
- Add a changelog fragment covering manual recovery.
- Update the service page if endpoints or integrations changed.
- Add the observed failure mode to this runbook.