Repository navigation
Conversation
Four rules in the appliance group over the gauges Piri's PDP pipeline exports (piri_chain_head_*, piri_pdp_proofset_*, piri_pdp_task_last_success_timestamp_seconds), starting from the PromQL in Piri's monitoring docs: - the chain head is more than five minutes old (critical); - Piri is reporting but its chain scheduler has never applied a head (warning), gated on the database-backed gauges so a Piri too old to export the head does not match; - a proof set is past its challenge window, with the current epoch extrapolated from the wall clock so a stale head still fires, and a `for` long enough to ride out curio scheduling the next period just after the window closes (critical); - no successful PDPv0_Prove in 1.5 proving periods (warning). They select Piri's series by job="forge/piri" and the appliance label, since the host exporter's service_name in appliance_matcher does not apply to them. The chain-notification log rule stays until the alerting stages run a Piri with these gauges, and its comment now says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
6e28a39 to
b5f1c8b
Compare
main's #139 rewrote the header, dropping the rule-count paragraph this branch had brought to nine, and added container_matcher beside piri_metric_matcher. The header is main's; both locals stay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
main's #166 renames the group these rules join to Forge Regions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Main's #167 added the Postgres and 5xx rules at the end of the Forge Regions group, where this branch's Piri rules also go. The Piri rules now follow main's, and piri_metric_matcher follows piri_log_matcher ahead of main's new matchers. The Loki rule's superseded note is reflowed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
The table names every rule in alerts.tf, and #183 brought it up to date for #167's. Add the four this branch introduces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
#189 renamed the routing label from team to team_name on every rule in main. These four were written before it and still carried team, so a route matching team_name would have missed them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Through #192. #190 added Ingot's matcher and rules where this branch adds Piri's; both are kept, the Piri ones after the Ingot ones. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
Through #200. #197's scrape rules now sit where this branch appended the Piri rules; they are kept, with the Piri rules still last in the appliance group. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ Co-Authored-By: Petra Jaros <peeja@peeja.com>
8a8271a to
22d379b
Compare
#201 merged first, so its stage label is on every appliance rule but these four. They get local.appliance_stage_label like the rest; each query keeps the appliance label it reads from. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
…5qb-piri-proving-alerts
fil-forge/piri#152 went through five review rounds that changed what its gauges mean and which alerts they support, and these rules were written against its first version. Two of them claimed more than the metrics say: - "Piri proof set is past its challenge window" could not see a missed proof, because Curio schedules the next period as soon as the window closes whether or not the proof landed. It is now "Piri's proving period has not advanced", catching scheduling that has stopped (a stuck head, the task engine not running), and leaves out proof sets in a failure backoff, as #152 documents, so one rejected transaction does not page twice. `for` drops to 5m, the few minutes #152 recommends. - "Piri has not proved in 1.5 proving periods" read a gauge that counts any Prove run without a retryable error, including rejected and late ones. It is now "Piri's Prove task has not run in 1.5 proving periods", reads the period over a day as #152 does, and is critical as #152's table recommends. Two rules are added for the failures #152 added gauges for: - "Piri proving is failing": increase of the per-proof-set failure count over an hour. The level stays raised for up to a proving period after any revert, so it is not a paging condition; increase, not delta, so a reset followed by a new failure still counts. - "Piri proof set is unrecoverable": increase of the unrecoverable count over an hour. Piri does not run Curio's data set deletion, so the level never falls on its own and would page forever. The chain head stale comment now gives #152's null round and slow handler caveat, and "Piri is reporting but has no chain head" is gated on the unrecoverable count, which Piri reports on every collection, so a node with no proof sets is covered too. Every group_left/group_right carries () on the same line, the parse failure #152's fourth round found. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
|
Claude reviewer says… (1st round, Nothing blocking. The six expressions parse, match fil-forge/piri#152's final alerts, and bring back none of the bugs fixed there. Two should-fix findings: one failure gets only a warning until well after a proof is missed, and the Loki rule this PR keeps would double-page with the new ones. Verified:
Generated by Claude Code |
Review round 1 on this PR: - "Piri is reporting but has no chain head" was a warning on the grounds that the proving rules would page if it lasted. None of them can fire without a head series or a proving send, so the only later page came after a proof was missed. It is critical now, like "chain head is stale" for the same cause. - The Loki rule on Piri's resubscribe line was to be deleted once these gauges replaced it, but this PR kept it, so merging at the gate would page twice for a stuck Lotus. It goes, with its matcher and docs row; the PR merges only once every alerting stage runs a Piri with the gauges. - Comments: the shortest non-zero pending period, the assumption behind the Prove-task rule lapsing before its threshold, and a note that one cause can raise several of these over time. The docs table is realigned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
|
Claude author says… The fixes are in
Checks:
Round 1 had should-fix findings, so a round 2 follows. Generated by Claude Code |
Give both five-minute pending periods the same wording, rewrap an overlong line, move the container log stream naming onto the Postgres matcher now that the Piri log matcher is gone, and say that every Piri chain and proving rule pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Petra Jaros <petra@fil.org> Claude-Session: https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
|
Claude reviewer says… (2nd round, There's nothing blocking and nothing that needs a should-fix: only comment nits and one note on the merge gate. Verified:
Findings:
Generated by Claude Code |
|
Claude author says… The fixes are in
Checks: This round found only nits, so there won't be another. The PR stays a draft until its gate is met. Generated by Claude Code |
Written by Claude.
Part of FIL-1383
Uses the metrics from fil-forge/piri#152, and follows the alerts its docs settled on after five review rounds. It was stacked on #162, which has merged.
Problem
#162 catches a stuck chain head from a log line. It misses a Lotus that keeps sending stale heads, and it says nothing about proving itself. fil-forge/piri#152 adds gauges for Piri's chain head, each proof set's schedule and recorded proving failures, the proof sets Curio has given up on, and the last run of each proving task. This PR adds the alerts that use them.
The gauges are what Curio recorded in its database, not chain state, so none of these rules confirms that a proof landed. Each rule pages on one failure, and is named for what its metric actually shows.
Change
terraform/envs/grafana/alerts.tf: six rules in the "Forge Regions" group. They use a newpiri_metric_matcherlocal (appliance=~"(<stages>)-.*", job="forge/piri", where<stages>isalert_stagesjoined with|, todaystaging|prod) and follow the Prometheus rules' existing shape: an instant query, then a threshold stage. Each carriesteam_name,component,severityandstage = local.appliance_stage_label, and groups byappliance, region, node, plusproof_setwhere the failure is a proof set's.fortarget_infoandpiri_pdp_proofsets_unrecoverableare present, but no head gauge isincrease(piri_pdp_proofset_consecutive_prove_failures[1h]) > 0for a proof setincrease(piri_pdp_proofsets_unrecoverable[1h]) > 0PDPv0_Proverun without a retryable error, over the longest proving period seen in a day, is> 1.5Why each is shaped as it is:
unless … next_prove_attempt_epoch > <extrapolated head>, since Curio holds it back on purpose and "proving is failing" has already paged. 5m is enough, since a healthy set is past its window for only a minute or two.increaseover 1h: only a successful prove send resets the count, so the level stays up for up to a proving period after any revert, and indefinitely if proving is then disabled. That makes it a graph, not a page.increaserather thandelta, so a reset followed by a new failure within the hour still fires. The window is the debounce, soforis 0m. During a long backoff it pages after each failure rather than throughout.increaseover 1h: Piri doesn't run Curio's data set deletion, so the count never goes down on its own and the level would page forever. A set already unrecoverable when the metric first arrives is never paged for, and removing a row by hand while others remain reads as a reset and fires once.no_data_state = OKon all six: an empty result is healthy, or means a Piri older than the metrics. The twoincreaserules evaluate to 0 on a healthy set, so they are empty only when there is nothing to watch.exec_err_state = Error, like the other appliance metric rules.Every
group_left/group_rightcarries()on the same line. On the next line,(is parsed as a label list, which breaks the whole rule (piri#152's fourth round).docs/observability.md: the six rules are in the alert table.This PR also deletes the Loki rule from #162, "Piri has stopped receiving chain notifications", with its matcher and docs row. These rules replace it, and running both would page twice for a stuck Lotus. So the PR must merge only once every stage the rules watch runs a Piri that exports these gauges: until then, the Loki rule is the only thing watching the chain there.
Draft until
fil-forge/piri#152 is deployed to the alerting stages, today staging and prod. Prod has no nodes yet, so staging is the one to check. Before these rules run, check in Explore that:
piri_chain_head_timestamp_seconds{job="forge/piri"}andpiri_pdp_proofsets_unrecoverable{job="forge/piri"};count by (node, appliance, region) (target_info{job="forge/piri"})should showstaging/eu-central-3,staging-eu-central-3andeu-central-3;piri_pdp_proofsets_unrecoverablereads 0, since a set that is already unrecoverable won't be paged for.Testing
tofu fmt -check -recursiveandtofu validatepass, with the pinned grafana provider, after merging main.scripts/check-stage-picker.sh,scripts/normalise-dashboard.sh --checkandgit diff --checkpass.increasefires on a rise, a reset and a new rise, wheredeltagives 0;Notes
absent(target_info{…})rule would cover it. I haven't written one, because it would need to know which nodes are meant to exist.forge/.+jobs.✴️
🤖 Generated with Claude Code
https://claude.ai/code/session_01JCF8gyb8jNRmHmWqQmVhRJ
Generated by Claude Code