Skip to content

Alert on Piri's chain head and PDP proving metrics - #163

Draft
Peeja wants to merge 16 commits into
mainfrom
claude/happy-bell-cfh5qb-piri-proving-alerts
Draft

Peeja wants to merge 16 commits into
mainfrom
claude/happy-bell-cfh5qb-piri-proving-alerts

Conversation

@Peeja

@Peeja Peeja commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Written by Claude.

Part of FIL-1383

Uses the metrics from fil-forge/piri#152, and follows the alerts its docs settled on after five review rounds. It was stacked on #162, which has merged.

Problem

#162 catches a stuck chain head from a log line. It misses a Lotus that keeps sending stale heads, and it says nothing about proving itself. fil-forge/piri#152 adds gauges for Piri's chain head, each proof set's schedule and recorded proving failures, the proof sets Curio has given up on, and the last run of each proving task. This PR adds the alerts that use them.

The gauges are what Curio recorded in its database, not chain state, so none of these rules confirms that a proof landed. Each rule pages on one failure, and is named for what its metric actually shows.

Change

terraform/envs/grafana/alerts.tf: six rules in the "Forge Regions" group. They use a new piri_metric_matcher local (appliance=~"(<stages>)-.*", job="forge/piri", where <stages> is alert_stages joined with |, today staging|prod) and follow the Prometheus rules' existing shape: an instant query, then a threshold stage. Each carries team_name, component, severity and stage = local.appliance_stage_label, and groups by appliance, region, node, plus proof_set where the failure is a proof set's.

Rule Severity for Fires when
Piri's chain head is stale critical 5m the last applied tipset is more than 300s old
Piri is reporting but has no chain head critical 10m target_info and piri_pdp_proofsets_unrecoverable are present, but no head gauge is
Piri's proving period has not advanced critical 5m a proof set's recorded challenge window has closed, against the head extrapolated from wall-clock time, and the set is not in a failure backoff
Piri proving is failing critical 0m increase(piri_pdp_proofset_consecutive_prove_failures[1h]) > 0 for a proof set
Piri proof set is unrecoverable critical 0m increase(piri_pdp_proofsets_unrecoverable[1h]) > 0
Piri's Prove task has not run in 1.5 proving periods critical 10m time since the last PDPv0_Prove run without a retryable error, over the longest proving period seen in a day, is > 1.5

Why each is shaped as it is:

  • Head stale, 5m: a run of null rounds or a slow chain scheduler handler can leave five minutes between heads, so the head has to stay stale across two of the group's five-minute evaluations.
  • No chain head, the extra gate: Piri reports the unrecoverable count on every collection, proof sets or none, and it does not depend on the chain. Requiring it keeps the rule quiet on a Piri from before chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152, and covers a node with no proof sets.
  • Not advanced: it catches scheduling that has stopped, such as a stuck head or the task engine not running. It does not catch a missed proof: Curio schedules the next period as soon as the window closes, whether or not the proof landed. A set in a failure backoff is left out with unless … next_prove_attempt_epoch > <extrapolated head>, since Curio holds it back on purpose and "proving is failing" has already paged. 5m is enough, since a healthy set is past its window for only a minute or two.
  • Proving failing, increase over 1h: only a successful prove send resets the count, so the level stays up for up to a proving period after any revert, and indefinitely if proving is then disabled. That makes it a graph, not a page. increase rather than delta, so a reset followed by a new failure within the hour still fires. The window is the debounce, so for is 0m. During a long backoff it pages after each failure rather than throughout.
  • Unrecoverable, increase over 1h: Piri doesn't run Curio's data set deletion, so the count never goes down on its own and the level would page forever. A set already unrecoverable when the metric first arrives is never paged for, and removing a row by hand while others remain reads as a reset and fires once.
  • Prove task, critical: the metric is the last run without a retryable error, which includes runs that woke too late or had their proof rejected. So it shows the task is running, not that proofs land. It adds a Prove task that keeps failing with a retryable error, or never runs, while scheduling carries on. Critical, as chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152's table recommends. The period is read over a day, so a node whose last set goes unrecoverable loses it before the threshold, and this rule doesn't repeat that page.
  • no_data_state = OK on all six: an empty result is healthy, or means a Piri older than the metrics. The two increase rules evaluate to 0 on a healthy set, so they are empty only when there is nothing to watch.
  • exec_err_state = Error, like the other appliance metric rules.

Every group_left/group_right carries () on the same line. On the next line, ( is parsed as a label list, which breaks the whole rule (piri#152's fourth round).

docs/observability.md: the six rules are in the alert table.

This PR also deletes the Loki rule from #162, "Piri has stopped receiving chain notifications", with its matcher and docs row. These rules replace it, and running both would page twice for a stuck Lotus. So the PR must merge only once every stage the rules watch runs a Piri that exports these gauges: until then, the Loki rule is the only thing watching the chain there.

Draft until

fil-forge/piri#152 is deployed to the alerting stages, today staging and prod. Prod has no nodes yet, so staging is the one to check. Before these rules run, check in Explore that:

  • the metrics arrive under these names, for example piri_chain_head_timestamp_seconds{job="forge/piri"} and piri_pdp_proofsets_unrecoverable{job="forge/piri"};
  • staging's Piri series carry the labels the rules select and group by: count by (node, appliance, region) (target_info{job="forge/piri"}) should show staging/eu-central-3, staging-eu-central-3 and eu-central-3;
  • piri_pdp_proofsets_unrecoverable reads 0, since a set that is already unrecoverable won't be paged for.

Testing

  • tofu fmt -check -recursive and tofu validate pass, with the pinned grafana provider, after merging main.
  • scripts/check-stage-picker.sh, scripts/normalise-dashboard.sh --check and git diff --check pass.
  • All six expressions, with the matcher filled in, parse in Prometheus v0.314's PromQL parser. The file's other fourteen also parse, with placeholder matchers.
  • In v0.314's test engine, with synthetic series:
    • "not advanced" fires for a set past its window;
    • it leaves out a set whose backoff deadline is ahead of the head, and takes it back once the head passes;
    • increase fires on a rise, a reset and a new rise, where delta gives 0;
    • a flat count gives 0.
  • Not run: any rule against real series, and Grafana's own validation, which only happens on apply.

Notes

  • Not caught: a proof Curio sent that then failed on chain, and a Prove task that woke after its window closed. Curio records both as successful runs and leaves the failure count alone. Catching them needs "last proven" from chain state, which chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152 leaves as follow-up.
  • A Piri that stops reporting entirely isn't caught by any rule in this file, because an empty result is healthy for all of them. A separate absent(target_info{…}) rule would cover it. I haven't written one, because it would need to know which nodes are meant to exist.
  • A node added later must run a Piri that includes Export chain head and PDP proving metrics piri#152. One that predates it exports none of these gauges, so no rule here fires for it, and the Loki rule that would have caught a stuck head is gone.
  • Staging's Alloy now runs Select staging's Forge telemetry by service.namespace infra-nodes#113's namespace-based configuration, which keeps forge/.+ jobs.
  • The 30-second epoch is hard-coded, which is right for Calibration and mainnet.

✴️

🤖 Generated with Claude Code

https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ


Generated by Claude Code

@Peeja
Peeja added this pull request to stack #165 September 30, 2026 18:11
Base automatically changed from claude/happy-bell-cfh5qb-piri-chain-alert to main September 30, 2026 19:18
Four rules in the appliance group over the gauges Piri's PDP pipeline
exports (piri_chain_head_*, piri_pdp_proofset_*,
piri_pdp_task_last_success_timestamp_seconds), starting from the PromQL
in Piri's monitoring docs:

- the chain head is more than five minutes old (critical);
- Piri is reporting but its chain scheduler has never applied a head
  (warning), gated on the database-backed gauges so a Piri too old to
  export the head does not match;
- a proof set is past its challenge window, with the current epoch
  extrapolated from the wall clock so a stale head still fires, and a
  `for` long enough to ride out curio scheduling the next period just
  after the window closes (critical);
- no successful PDPv0_Prove in 1.5 proving periods (warning).

They select Piri's series by job="forge/piri" and the appliance label,
since the host exporter's service_name in appliance_matcher does not
apply to them. The chain-notification log rule stays until the alerting
stages run a Piri with these gauges, and its comment now says so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
claude and others added 10 commits September 30, 2026 19:58
main's #139 rewrote the header, dropping the rule-count paragraph this
branch had brought to nine, and added container_matcher beside
piri_metric_matcher. The header is main's; both locals stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
main's #166 renames the group these rules join to Forge Regions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Main's #167 added the Postgres and 5xx rules at the end of the Forge
Regions group, where this branch's Piri rules also go. The Piri rules
now follow main's, and piri_metric_matcher follows piri_log_matcher
ahead of main's new matchers. The Loki rule's superseded note is
reflowed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
The table names every rule in alerts.tf, and #183 brought it up to date
for #167's. Add the four this branch introduces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
#189 renamed the routing label from team to team_name on every rule in
main. These four were written before it and still carried team, so a route
matching team_name would have missed them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Through #192. #190 added Ingot's matcher and rules where this branch adds
Piri's; both are kept, the Piri ones after the Ingot ones.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
Through #200. #197's scrape rules now sit where this branch appended the
Piri rules; they are kept, with the Piri rules still last in the
appliance group.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Co-Authored-By: Petra Jaros <peeja@peeja.com>
claude and others added 3 commits October 8, 2026 21:02
#201 merged first, so its stage label is on every appliance rule but these
four. They get local.appliance_stage_label like the rest; each query keeps
the appliance label it reads from.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
fil-forge/piri#152 went through five review rounds that changed what its
gauges mean and which alerts they support, and these rules were written
against its first version. Two of them claimed more than the metrics say:

- "Piri proof set is past its challenge window" could not see a missed
  proof, because Curio schedules the next period as soon as the window
  closes whether or not the proof landed. It is now "Piri's proving period
  has not advanced", catching scheduling that has stopped (a stuck head,
  the task engine not running), and leaves out proof sets in a failure
  backoff, as #152 documents, so one rejected transaction does not page
  twice. `for` drops to 5m, the few minutes #152 recommends.
- "Piri has not proved in 1.5 proving periods" read a gauge that counts
  any Prove run without a retryable error, including rejected and late
  ones. It is now "Piri's Prove task has not run in 1.5 proving periods",
  reads the period over a day as #152 does, and is critical as #152's
  table recommends.

Two rules are added for the failures #152 added gauges for:

- "Piri proving is failing": increase of the per-proof-set failure count
  over an hour. The level stays raised for up to a proving period after any
  revert, so it is not a paging condition; increase, not delta, so a reset
  followed by a new failure still counts.
- "Piri proof set is unrecoverable": increase of the unrecoverable count
  over an hour. Piri does not run Curio's data set deletion, so the level
  never falls on its own and would page forever.

The chain head stale comment now gives #152's null round and slow handler
caveat, and "Piri is reporting but has no chain head" is gated on the
unrecoverable count, which Piri reports on every collection, so a node with
no proof sets is covered too. Every group_left/group_right carries () on
the same line, the parse failure #152's fourth round found.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

Peeja commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Claude reviewer says… (1st round, ce80842)

Nothing blocking. The six expressions parse, match fil-forge/piri#152's final alerts, and bring back none of the bugs fixed there. Two should-fix findings: one failure gets only a warning until well after a proof is missed, and the Loki rule this PR keeps would double-page with the new ones.

Verified:

  • tofu fmt -check -recursive, tofu validate, scripts/check-stage-picker.sh and git diff --check pass.
  • All six expressions parse in Prometheus v0.314.0's parser, with the matcher filled in.
  • Each matches chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152's monitoring.md:
    • increase rather than delta;
    • the backoff exclusion against the extrapolated head;
    • group_left()/group_right() with () on the same line.
  • The metric names survive OTLP translation unchanged.
  • appliance, region and node survive every aggregation, so the stage template resolves.
  • Labels, no_data_state/exec_err_state, the missing dashboard link and for = 0m all follow the neighbouring rules.
  • At the group's 300s interval, for = 5m fires on the second evaluation and 0m on the first.
  • In promqltest, a Piri with target_info but no head fires "no chain head", while "not advanced" returns empty for the same node.
  • Not run: anything against live data, and Grafana's apply-time validation.
  1. should-fix: "Piri is reporting but has no chain head" is a warning, and its comment says the proving rules page if the failure lasts long enough (alerts.tf ~1525-1540). They don't:

    • "Not advanced" needs the head series, so it returns nothing.
    • "Head stale" has no series either.
    • "Proving failing" can't fire, because nothing is sent.
    • The Loki rule doesn't fire when Lotus is down outright.

    So a Piri against an unreachable Lotus gets only a warning until the Prove-task rule fires, about 36h later, while a proof is missed within about a day.

  2. should-fix: the Loki rule "Piri has stopped receiving chain notifications" stays (~765-793). Merging at the draft gate runs it alongside the new rules, so a stuck Lotus raises two critical alerts. Its comment "Nothing else Piri exports says so" also stops being true.

  3. nit: "for is 5m, the shortest pending period a group evaluated every five minutes can give". 0m is shorter, so it should say the shortest non-zero one (~1600-1610).

  4. nit: "loses the period series within the day, before the threshold" holds only while 1.5 proving periods is longer than a day (~1915-1920).

  5. note: one cause can raise several pages, for example a stuck head pages "head stale", then "not advanced", then "Prove task". That's chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152's design, and worth a sentence in the header comment.

  6. nit: the new docs/observability.md table rows don't line up with the table's columns.

✴️ Claude Code session


Generated by Claude Code

Review round 1 on this PR:

- "Piri is reporting but has no chain head" was a warning on the grounds
  that the proving rules would page if it lasted. None of them can fire
  without a head series or a proving send, so the only later page came
  after a proof was missed. It is critical now, like "chain head is
  stale" for the same cause.
- The Loki rule on Piri's resubscribe line was to be deleted once these
  gauges replaced it, but this PR kept it, so merging at the gate would page
  twice for a stuck Lotus. It goes, with its matcher and docs row; the PR
  merges only once every alerting stage runs a Piri with the gauges.
- Comments: the shortest non-zero pending period, the assumption behind the
  Prove-task rule lapsing before its threshold, and a note that one cause
  can raise several of these over time. The docs table is realigned.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

Peeja commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Claude author says…

The fixes are in a89fbb4.

  1. Fixed. "Piri is reporting but has no chain head" is now critical, like "chain head is stale", which has the same cause. Its comment now says why: none of the proving rules can fire for a node with no head, so the only later page would come after a proof was missed.
  2. Fixed by deleting the Loki rule "Piri has stopped receiving chain notifications", along with its now-unused piri_log_matcher and its docs row.
    • Why: these rules replace it, and the header comment now says that running both would page twice.
    • When to merge: the PR body now says it must merge only once every stage the rules watch runs a Piri that exports the gauges, because until then the Loki rule is all that watches the chain there.
  3. Fixed: "the shortest non-zero pending period".
  4. Fixed. The comment now gives the assumption: the Prove-task rule lapses before its threshold only while 1.5 proving periods is longer than a day, as with mainnet's 2880 epochs.
  5. Added to the header comment: one cause can raise more than one of these over time, and each names the failure as it has become by then.
  6. Fixed. The docs/observability.md table is realigned, including the rows that predate this PR.

Checks:

  • tofu fmt -check -recursive, tofu validate on terraform/envs/grafana, scripts/check-stage-picker.sh and git diff --check pass.
  • The PromQL is unchanged since round 1's parse check.

Round 1 had should-fix findings, so a round 2 follows.

✴️ Claude Code session


Generated by Claude Code

Give both five-minute pending periods the same wording, rewrap an
overlong line, move the container log stream naming onto the Postgres
matcher now that the Piri log matcher is gone, and say that every Piri
chain and proving rule pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Petra Jaros <petra@fil.org>
Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ

Peeja commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Claude reviewer says… (2nd round, a89fbb4)

There's nothing blocking and nothing that needs a should-fix: only comment nits and one note on the merge gate.

Verified:

  • tofu fmt, tofu validate, scripts/check-stage-picker.sh and git diff --check pass.
  • All six expressions parse with the matcher filled in.
  • Apart from comments and the deletion, a89fbb4's only change is warning → critical on "reporting but has no chain head".
  • Nothing in infra-central or infra-nodes' docs/RUNBOOK.md still refers to the deleted Loki rule or piri_log_matcher.
  • The docs table realignment changed only whitespace in existing rows.
  • Coverage holds for nodes that export the gauges. piri_chain_head_timestamp_seconds is the tipset's own timestamp, so a Lotus that resubscribes with a stuck tipset still trips "chain head is stale".
  • The no-head rule's claim holds: of the other rules, only the Prove-task rule can still fire for a Piri with no head.
  • The metric names, backoff exclusion, increase window and thresholds match Export chain head and PDP proving metrics piri#152's final monitoring.md.

Findings:

  1. nit (alerts.tf, "proving period has not advanced"): it still says for is "the shortest a five-minute group gives"; 0m is shorter. Use the "shortest non-zero pending period" wording that the chain-head rule now has.
  2. nit (alerts.tf, chain-head rule): one comment line runs to 85 columns.
  3. nit (alerts.tf, postgres_log_matcher/openbao_log_matcher): "same scheme" referred to the deleted piri_log_matcher comment, the one that explained how container log streams are named. It now follows piri_metric_matcher, which uses a different scheme.
  4. nit (alerts.tf, Piri group header): "The proving failures page (critical)" undersells it, since the chain-head rules page too.
  5. note (PR body): the merge gate is right for today. But once this merges, a node added later that runs a Piri older than chore(deps): Bump github.com/ethereum/go-ethereum from 1.17.5 to 1.17.6 #152 has no chain or proving coverage at all, where the Loki rule used to cover it. Worth a line in Notes.

✴️


Generated by Claude Code

Peeja commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Claude author says…

The fixes are in 2e4e8e8.

  1. Both five-minute pending periods now say "the shortest non-zero pending period a five-minute group gives".
  2. Rewrapped.
  3. The explanation of container log stream naming (Alloy names each stream appliance-<stage>-<region>-<service> from the Compose service and stamps appliance, region and node) now sits on postgres_log_matcher. OpenBao's "same scheme" refers back to it again.
  4. The header now says all of these rules page, and why.
  5. Added under Notes: a node added later must run a Piri that includes Export chain head and PDP proving metrics piri#152.

Checks: tofu fmt -check -recursive, scripts/check-stage-picker.sh and git diff --check pass. The commit changes only comments, so I didn't run the PromQL checks again.

This round found only nits, so there won't be another. The PR stays a draft until its gate is met.

✴️


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants