Repository navigation
grafana: page when OpenBao cannot unseal or a site fails outright, and label alerts by stage - #201
Merged
Merged
Conversation
…ls outright Two outage-class rules FIL-1164 asks for, in a new 60s group so detection fits inside five minutes. Both are severity = "critical", which is what production routing keys on. OpenBao cannot unseal reads OpenBao's own retry loop: when auto-unseal through Central's transit key fails, it logs "failed to unseal core" and retries every five seconds (openbao command/server.go, runUnseal, v2.6.2). Appliance site is failing most requests is the outage tier of the 5xx warning: more than half of a site's requests answered with a 5xx for two minutes, per host, so Ingot's site is judged on its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Production critical alerts are to page where staging's notify, and an IRM route can only tell them apart by a label the alert carries. Central's service rules already take stage from their queries. The appliance rules now derive it from the appliance label, <stage>-<region>, as #199's dashboard links do, and the provision Lambda rule from its function name, fc-<stage>-provision. This changes every appliance alert instance's identity once, so their state resets when it is applied. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Peeja
marked this pull request as ready for review
October 8, 2026 14:50
Peeja
force-pushed
the
claude/happy-bell-cfh5qb-outage-rules
branch
from
October 8, 2026 15:01
8230f1d to
741e6dd
Compare
… unstaged alerts route A review of the outage rules found: - The site-failing rule's 0.1 req/s floor hid a quiet site failing every request, and a busy one whose clients back off once it fails. It now needs at least five 5xx responses in its three-minute window instead. - An alert raised because a query failed or returned nothing carries none of the query's labels, so it has no stage and a stage="prod" route misses it. The stage label's comment and docs/observability.md now say so, and the docs ask for critical alerts with no stage to page. - OpenBao's "error unsealing core" only follows a failed declarative self-init, which these nodes don't use, so the rule matches only the "failed to unseal core" retry line. Its comment also notes that an uninitialised OpenBao fires it. - The stage label's local now sits above the dashboard-link comment instead of splitting it from the locals it describes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
#196 appended the Forge Usage group to alerts.tf where this branch appended Forge Regions outages; both are kept, outages first. The routing paragraph in docs/observability.md now allows for #196's info severity and for its stack-wide rule, which has no stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Peeja
added a commit
that referenced
this pull request
Oct 8, 2026
#201 merged first, so its stage label is on every appliance rule but these four. They get local.appliance_stage_label like the rest; each query keeps the appliance label it reads from. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Written by Claude.
Part of FIL-1164.
Adds the two outage-class appliance rules that FIL-1164 still lacks. The third, Piri's proving deadline, is #163.
Rules
Both rules are
severity = "critical". They sit in a new group, Forge Regions outages, evaluated every 60s. At the 300s the other appliance rules use, evaluation alone could use up FIL-1164's five-minute budget.OpenBao cannot unseal. The node auto-unseals through the transit key at Central. When that fails, OpenBao logs
failed to unseal coreat WARN and retries every five seconds for as long as the failure lasts (command/server.go,runUnseal, at v2.6.2, the version the nodes pin). The only failure that loop treats as fatal is a failed declarative self-init, and the nodes initialise withbao operator init, so the rule matches only the WARN line. It counts it over a one-minute window, withfor= 2m:An OpenBao that was started but never initialised logs the same line, so a region whose provisioning stops between starting OpenBao and
bao operator initfires it too; that node is sealed all the same.It doesn't catch a node sealed by hand with
bao operator seal. That logsvault is sealedonce and doesn't retry, and every graceful restart logs the same line.Appliance site is failing most requests. This is the outage tier of "Appliance 5xx rate too high", which stays as the warning. It fires when more than half of a site's requests get a 5xx for two minutes. It is split by host, so Ingot's site is judged on its own: Ingot exports no HTTP metrics, so Caddy's view of its site is its error rate. It uses a three-minute rate window, because Caddy is scraped every minute and
rate()needs two samples, and links to the same Error rate panel.It only fires once a site has served at least five 5xx responses in the window. The warning's floor of 0.1 requests per second would hide a quiet site that fails everything, and a busy one whose clients back off once it fails. Five errors in three minutes still keeps a single failed request on an idle site from paging. A site whose clients stop sending entirely has no errors to count, and this rule can't see it.
docs/observability.md's alert table has a row for each.A
stagelabel on every alertThe aim is for production critical alerts to page where staging's notify. An IRM route can only tell the two apart by a label the alert carries:
stagefrom their queries.appliancelabel, which is<stage>-<region>by construction. It is cut at the first hyphen, the way grafana: make an alert's dashboard link open on the stage that fired #199's dashboard links do.fc-<stage>-provision.An alert raised because a query failed (
DatasourceError) or came back empty (DatasourceNoData) carries none of the query's labels, so it has nostage, and a route matchingstage = "prod"won't see it.docs/observability.mdnow describes the routing and says IRM should page on a critical alert with nostageas well as onstage = "prod".Adding a label changes every appliance alert instance's identity. Their state resets once, when this is applied.
Whichever of this and #163 merges second needs
stageon #163's four rules too.Not checked
{service_name=~"appliance-.*-openbao"} |~ "unseal". It should show the unseal lines, and none offailed to unseal corewhile healthy.stageor keeps it as""on aDatasourceErrorinstance. Either way it doesn't matchprod; the IRM route for "no stage" should be tested with a forced datasource error on a staging rule.service_namethe rule matches is inferred. It isappliance-<stage>-<region>-openbao, from the Alloy rule that builds names out of the Compose service, as dev's config does. Staging's Alloy is the host's own and lives outside these repositories.stagelabel is unverified against live data. Check it after applying: the rules' instances in Grafana Alerting should showstage=staging, and not an empty value or the whole appliance name.Testing
make checkpasses, includingtofu fmtand the stage-picker check.tofu validatepasses on the cached grafana provider 4.46.0.✴️
🤖 Generated with Claude Code
https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Generated by Claude Code