Skip to content

Detect frozen PVs with an independent CA context - #50

Merged
slominskir merged 1 commit into
mainfrom
frozen-pv-detection
Oct 4, 2026
Merged

slominskir merged 1 commit into
mainfrom
frozen-pv-detection

Conversation

@slominskir-coding-agent

Copy link
Copy Markdown
Contributor

Step 2 of the plan discussed for #28: detect "frozen" PVs, so the periodic Tomcat restart can eventually be replaced by evidence. Report-only for now.

A frozen PV is one whose monitor has stopped working while the IOC still serves it, so restarting epics2web would be expected to fix it. A PV whose IOC is down isn't frozen.

How it works

FrozenPvDetector runs every FROZEN_CHECK_SECONDS (default 10). It gives suspicious PVs a short-lived subscription in a second CA context, which has its own virtual circuits. The subscription is made like ChannelMonitor's: same type, count and Monitor.VALUE mask. A PV is frozen when either:

  • lost connection: its monitor has been DISCONNECTED, or still CONNECTING, longer than the healthcheck grace period, but the independent subscription connects and receives the PV;
  • lost updates: its monitor used to get updates and has gone quiet, and the independent subscription receives a change while the monitor gets nothing. The change must be half an interval old before it counts, so a monitor that's merely a little behind isn't flagged.

Why subscriptions and not value comparison: comparing a fresh read with the monitor's last value would flag PVs with a monitor deadband (MDEL), whose value can drift within the deadband without an update being posted. Subscriptions made the same way get the same deadband from the IOC.

Load: only suspicious PVs are probed, at most 20 at a time, each for at most 6 intervals, and the same PV at most once every 30 intervals. Suspicious means not connected past the grace period, or quiet for 6 intervals after having changed at least once. A PV that has never changed can't show missing updates, so it's never probed. With the defaults, a probe lasts at most 60 s and a PV is re-probed at most every 5 minutes.

Through a gateway: both contexts reach PVs through the gateway, so this finds problems between epics2web and the gateway, not inside the gateway.

A frozen PV stops being frozen when its monitor gets an update.

Reporting

  • Each frozen PV is logged as a WARNING (PV … is frozen: <reason>), and its recovery at INFO.
  • /healthcheck lists frozen PVs with frozen, frozen_minutes and frozen_reason. They're added to a PV's existing entry, or get an entry of their own if the PV is still marked connected.
  • Status codes: the default and ?strict=true responses are unchanged. The new ?frozen=true answers 503 when any PV is frozen, for restart automation once it's trusted.
  • FROZEN_PV_CHECK=false turns detection off. The README documents both settings.

Suggested rollout

Deploy with the cron restart still in place, and watch the logs for is frozen warnings:

Checks

  • FrozenPvDetectorTest (unit, 11 tests, controlled clock and fake probes) covers:
    • a quiet PV that changed for the probe;
    • recovery;
    • a quiet PV that doesn't change;
    • a monitor update during a probe;
    • never-changing PVs;
    • disconnected and never-connected PVs the IOC serves;
    • an IOC that doesn't serve the PV;
    • the probe cap, removed PVs, and close.
  • FrozenPvDetectionTest (integration, about 19 s): runs the real detector in the test JVM against the test IOC, with two real CA contexts. For 10 s with HELLO, channel1 (never changes) and a missing PV, nothing is flagged. Then the test clears HELLO's CAJ subscription behind its monitor's back, as a lost subscription would look. HELLO, and only HELLO, is then flagged for a missed change. Passed 3 runs in a row.
    • This test caught a bug in my first version. The detector waited for the probe's latest change to be half an interval old, which never happens for a PV changing every 0.2 s. It now counts from the first change.
  • HealthcheckTest:
    • new workingMonitorsAreNotFrozen: 10 s of HELLO and channel1 gives ?frozen=true 200;
    • added to the disconnect test: with the IOC stopped, channel1 isn't flagged frozen, and ?frozen=true returns 200.
  • No false positives: the full integration suite (22 tests) passed twice with FROZEN_CHECK_SECONDS: 1, with no frozen warnings in the server log. A hung IOC (docker pause for 45 s) and an IOC restart weren't flagged either.
  • ./gradlew spotlessCheck test passes.

🤖 Generated with Claude Code

IOC and network restarts have historically left PVs "frozen": a monitor
that stops delivering while the IOC still serves the PV, which a
restart of epics2web fixes. Nothing reported that, so Tomcat is
restarted periodically by cron.

FrozenPvDetector gives suspicious PVs a short-lived subscription in a
second CA context, with its own virtual circuits, made like the
monitor's so the IOC applies the same deadband to both. Comparing
values instead would mistake a change within a deadband for a missed
update. A PV is frozen when its monitor has been disconnected, or
still connecting, past the grace period while the independent
subscription connects, or when a monitor that used to update has gone
quiet and the independent subscription receives a change it doesn't.
Only PVs not connected past the grace period, and quiet PVs that have
changed before, are probed, at most 20 at a time and each at most once
every 30 check intervals. A PV stops being frozen when its monitor gets
an update.

Detection is report-only: frozen PVs are logged as warnings and listed
by /healthcheck, and the default and strict responses are unchanged.
?frozen=true answers 503 when a PV is frozen, for restart automation
once it's trusted. FROZEN_CHECK_SECONDS (default 10) sets the check
interval, from which the other timings derive, and FROZEN_PV_CHECK=false
turns detection off.

FrozenPvDetectorTest covers the decisions with a controlled clock.
FrozenPvDetectionTest runs the detector against the test IOC with two
real CA contexts, and simulates a lost subscription by clearing a
monitor's CAJ subscription behind its back. HealthcheckTest checks that
working monitors and a stopped IOC aren't reported frozen.

Part of #28

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@slominskir
slominskir merged commit ec9df53 into main Oct 4, 2026
6 checks passed
@slominskir
slominskir deleted the frozen-pv-detection branch October 4, 2026 03:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

source::ai Work done by an AI agent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant