Skip to content

Alert when Piri stops receiving chain notifications - #162

Merged
Peeja merged 3 commits into
mainfrom
claude/happy-bell-cfh5qb-piri-chain-alert
Sep 30, 2026
Merged

Peeja merged 3 commits into
mainfrom
claude/happy-bell-cfh5qb-piri-chain-alert

Conversation

@Peeja

@Peeja Peeja commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Written by Claude.

Part of FIL-1383

Problem

On 2026-09-28, staging's Lotus stopped following the chain. Piri stayed up, answered its health check and kept sending metrics, but it stopped proving. Nothing alerted, and the problem surfaced more than a day later, when a deploy hung.

Piri exports no chain or proving metrics yet; FIL-1383 adds them. Until then, the one signal is in its logs. Curio's chain scheduler, which Piri runs, logs this while the head isn't moving:

no notifications received in 5m0s, resubscribing to ChainNotify

The text is at curio lib/chainsched/chain_sched.go:209, at the version Piri pins. It recurs about every ten minutes, and staging's Piri logged it throughout the incident.

Change

terraform/envs/grafana/alerts.tf: a fifth rule in the appliance group, "Piri has stopped receiving chain notifications". It's the first rule here that reads logs rather than metrics.

  • Query A, a Loki instant query. With alert_stages = ["staging"], as terraform.tfvars sets it, it renders as:
    sum by (appliance, region, node) (
      count_over_time(
        {appliance=~"(staging)-.*", service_name=~"appliance-.*-piri"}
          |~ "no notifications received in .* resubscribing to ChainNotify"
        [15m]
      )
    )
    
    The (staging) group comes from local.stages, which joins alert_stages with |. Adding production would make it (staging|prod).
  • Stage B: threshold > 0, the same shape as the other rules.
  • for = 15m, so one line doesn't fire the alert by itself. It fires about twenty minutes after the head stops moving.
  • no_data_state = OK: an empty result is the healthy state, as in the 5xx rule.
  • severity = critical: a node that isn't proving is what matters most.
  • Routing: team = forge, like every other rule here.
  • runbook_url: the runbook's "When something is wrong" section, as the other appliance rules use.
  • The stage filter uses alert_stages, as the other appliance rules do, so today it watches staging and not dev.

loki_datasource_uid: a new variable, set in terraform.tfvars like the Prometheus one. Its value, grafanacloud-logs, is confirmed from the data source's settings page URL.

Testing

  • Against the incident: query A, run in Explore as a range query over the stuck window (2026-09-28 23:00 to 2026-09-29 21:40 UTC), returns a series with node="staging/eu-central-3". The stream carries appliance="staging-eu-central-3" and service_name="appliance-staging-eu-central-3-piri". That confirms the selector, the log text and the labels the summary uses.
  • tofu fmt -check -recursive passes.
  • tofu validate passes, with init -backend=false and the pinned grafana provider 4.46.0. The lock file is unchanged.

Notes

  • Applies on merge if GRAFANA_APPLY_ENABLED is set.
  • Staging's Piri is stopped at the moment, so this won't fire there until it runs again. Once it does, the alert should fire until the host's Lotus is fixed, which is correct.
  • Not caught: a Lotus that's down outright logs a different error, so this rule doesn't catch it. Neither does a Lotus that keeps sending stale heads. Alert on Piri's chain head and PDP proving metrics #163's metric-based alerts cover both, once Export chain head and PDP proving metrics piri#152 is deployed, and this rule can go then.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

✴️

A stopgap for a Lotus whose head has stopped moving. Piri's chain
scheduler (curio lib/chainsched) logs "no notifications received in
5m0s, resubscribing to ChainNotify" about every ten minutes while the
head is stuck, and nothing Piri exports shows it otherwise.

The rule counts that line per node over fifteen minutes in Piri's log
stream, narrowed to the alerting stages by the appliance label the way
the metric rules are, and fires after fifteen minutes pending so a
single line does not page. It joins the appliance group with the same
team = "forge" routing label as every other rule.

It is the first rule here to read logs, so this adds a
loki_datasource_uid variable alongside the Prometheus one, set to
grafanacloud-logs by the same Grafana Cloud convention.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
@Peeja
Peeja marked this pull request as ready for review September 30, 2026 18:08
Copilot AI balanced review requested due to automatic review settings September 30, 2026 18:08

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Peeja
Peeja added this pull request to stack #165 September 30, 2026 18:11
@Peeja
Peeja merged commit 2439d2e into main Sep 30, 2026
16 checks passed
@Peeja
Peeja deleted the claude/happy-bell-cfh5qb-piri-chain-alert branch September 30, 2026 19:18
Peeja added a commit that referenced this pull request Sep 30, 2026
…ame the Regions groups (#166)

* grafana: say every data source a rule reads needs a query grant, and rename the Regions group

Applying #162 failed with a 403 putAlertRuleGroupForbidden: its rule is
the first to read grafanacloud-logs, and forge-terraform could query
only grafanacloud-prom. The handler refuses a group that reads any data
source the caller cannot query. The header now says the grant is needed
on every data source a rule reads, today prom and logs, and before the
rule merges.

The appliance group is renamed Forge Regions, matching its dashboard.
The provider replaces a rule group whose name changes, so its rules are
recreated with new uids and fresh state. observability.md's rule table
follows, and gains the chain-notification rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

* grafana: rename the appliance containers group Forge Regions containers

It stays its own group: a group's rules share one evaluation interval,
and this one runs every 60s where Forge Regions runs every 300s. Like
the Regions rename, the provider replaces the group, so its rule is
recreated with a new uid and fresh state.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

* grafana: don't list the data sources the query grant covers

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

* grafana: rewrap the query-grant comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

* grafana: reflow the query-grant comment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants