Skip to content

grafana: page when OpenBao cannot unseal or a site fails outright, and label alerts by stage - #201

Merged
Peeja merged 5 commits into
mainfrom
claude/happy-bell-cfh5qb-outage-rules
Oct 8, 2026
Merged

Peeja merged 5 commits into
mainfrom
claude/happy-bell-cfh5qb-outage-rules

Conversation

@Peeja

@Peeja Peeja commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Written by Claude.

Part of FIL-1164.

Adds the two outage-class appliance rules that FIL-1164 still lacks. The third, Piri's proving deadline, is #163.

Rules

Both rules are severity = "critical". They sit in a new group, Forge Regions outages, evaluated every 60s. At the 300s the other appliance rules use, evaluation alone could use up FIL-1164's five-minute budget.

OpenBao cannot unseal. The node auto-unseals through the transit key at Central. When that fails, OpenBao logs failed to unseal core at WARN and retries every five seconds for as long as the failure lasts (command/server.go, runUnseal, at v2.6.2, the version the nodes pin). The only failure that loop treats as fatal is a failed declarative self-init, and the nodes initialise with bao operator init, so the rule matches only the WARN line. It counts it over a one-minute window, with for = 2m:

  • A single failed attempt during a restart drops out of the window before two minutes of pending build up.
  • A real failure keeps every window full, so it fires about three minutes after the first failure.

An OpenBao that was started but never initialised logs the same line, so a region whose provisioning stops between starting OpenBao and bao operator init fires it too; that node is sealed all the same.

It doesn't catch a node sealed by hand with bao operator seal. That logs vault is sealed once and doesn't retry, and every graceful restart logs the same line.

Appliance site is failing most requests. This is the outage tier of "Appliance 5xx rate too high", which stays as the warning. It fires when more than half of a site's requests get a 5xx for two minutes. It is split by host, so Ingot's site is judged on its own: Ingot exports no HTTP metrics, so Caddy's view of its site is its error rate. It uses a three-minute rate window, because Caddy is scraped every minute and rate() needs two samples, and links to the same Error rate panel.

It only fires once a site has served at least five 5xx responses in the window. The warning's floor of 0.1 requests per second would hide a quiet site that fails everything, and a busy one whose clients back off once it fails. Five errors in three minutes still keeps a single failed request on an idle site from paging. A site whose clients stop sending entirely has no errors to count, and this rule can't see it.

docs/observability.md's alert table has a row for each.

A stage label on every alert

The aim is for production critical alerts to page where staging's notify. An IRM route can only tell the two apart by a label the alert carries:

  • Central's service rules already get stage from their queries.
  • Every appliance rule now derives it from the appliance label, which is <stage>-<region> by construction. It is cut at the first hyphen, the way grafana: make an alert's dashboard link open on the stage that fired #199's dashboard links do.
  • "Provision Lambda errors" derives it from its function name, fc-<stage>-provision.

An alert raised because a query failed (DatasourceError) or came back empty (DatasourceNoData) carries none of the query's labels, so it has no stage, and a route matching stage = "prod" won't see it. docs/observability.md now describes the routing and says IRM should page on a critical alert with no stage as well as on stage = "prod".

Adding a label changes every appliance alert instance's identity. Their state resets once, when this is applied.

Whichever of this and #163 merges second needs stage on #163's four rules too.

Not checked

  • The OpenBao log line comes from reading the source, not from seeing it in staging's Loki. Before relying on the rule, run this query in Explore over a recent restart of the staging node's OpenBao: {service_name=~"appliance-.*-openbao"} |~ "unseal". It should show the unseal lines, and none of failed to unseal core while healthy.
  • Whether Grafana drops an empty stage or keeps it as "" on a DatasourceError instance. Either way it doesn't match prod; the IRM route for "no stage" should be tested with a forced datasource error on a staging rule.
  • The service_name the rule matches is inferred. It is appliance-<stage>-<region>-openbao, from the Alloy rule that builds names out of the Compose service, as dev's config does. Staging's Alloy is the host's own and lives outside these repositories.
  • The stage label is unverified against live data. Check it after applying: the rules' instances in Grafana Alerting should show stage=staging, and not an empty value or the whole appliance name.
  • The 50% threshold and the five-error gate are chosen, not derived, like the warning's 5%. FIL-1242 would set them.

Testing

  • make check passes, including tofu fmt and the stage-picker check.
  • tofu validate passes on the cached grafana provider 4.46.0.

✴️

🤖 Generated with Claude Code

https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ


Generated by Claude Code

claude and others added 2 commits October 7, 2026 22:23
…ls outright

Two outage-class rules FIL-1164 asks for, in a new 60s group so detection
fits inside five minutes. Both are severity = "critical", which is what
production routing keys on.

OpenBao cannot unseal reads OpenBao's own retry loop: when auto-unseal
through Central's transit key fails, it logs "failed to unseal core" and
retries every five seconds (openbao command/server.go, runUnseal, v2.6.2).

Appliance site is failing most requests is the outage tier of the 5xx
warning: more than half of a site's requests answered with a 5xx for two
minutes, per host, so Ingot's site is judged on its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Production critical alerts are to page where staging's notify, and an
IRM route can only tell them apart by a label the alert carries.
Central's service rules already take stage from their queries. The
appliance rules now derive it from the appliance label, <stage>-<region>,
as #199's dashboard links do, and the provision Lambda rule from its
function name, fc-<stage>-provision.

This changes every appliance alert instance's identity once, so their
state resets when it is applied.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
@Peeja Peeja changed the title grafana: page when an appliance's OpenBao cannot unseal or a site fails outright grafana: page when OpenBao cannot unseal or a site fails outright, and label alerts by stage Oct 7, 2026
@Peeja
Peeja marked this pull request as ready for review October 8, 2026 14:50
Copilot AI balanced review requested due to automatic review settings October 8, 2026 14:50

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Peeja
Peeja force-pushed the claude/happy-bell-cfh5qb-outage-rules branch from 8230f1d to 741e6dd Compare October 8, 2026 15:01
claude and others added 2 commits October 8, 2026 15:33
… unstaged alerts route

A review of the outage rules found:

- The site-failing rule's 0.1 req/s floor hid a quiet site failing every
  request, and a busy one whose clients back off once it fails. It now
  needs at least five 5xx responses in its three-minute window instead.
- An alert raised because a query failed or returned nothing carries none
  of the query's labels, so it has no stage and a stage="prod" route
  misses it. The stage label's comment and docs/observability.md now say
  so, and the docs ask for critical alerts with no stage to page.
- OpenBao's "error unsealing core" only follows a failed declarative
  self-init, which these nodes don't use, so the rule matches only the
  "failed to unseal core" retry line. Its comment also notes that an
  uninitialised OpenBao fires it.
- The stage label's local now sits above the dashboard-link comment
  instead of splitting it from the locals it describes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
#196 appended the Forge Usage group to alerts.tf where this branch appended
Forge Regions outages; both are kept, outages first. The routing paragraph in
docs/observability.md now allows for #196's info severity and for its
stack-wide rule, which has no stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
@Peeja
Peeja merged commit 427dff6 into main Oct 8, 2026
22 checks passed
@Peeja
Peeja deleted the claude/happy-bell-cfh5qb-outage-rules branch October 8, 2026 20:54
Peeja added a commit that referenced this pull request Oct 8, 2026
#201 merged first, so its stage label is on every appliance rule but these
four. They get local.appliance_stage_label like the rest; each query keeps
the appliance label it reads from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants