Skip to content

Report PVs by time since disconnect, and fail the healthcheck only in strict mode - #49

Merged
slominskir merged 1 commit into
mainfrom
healthcheck-connection-time
Oct 4, 2026
Merged

slominskir merged 1 commit into
mainfrom
healthcheck-connection-time

Conversation

@slominskir-coding-agent

Copy link
Copy Markdown
Contributor

Fixes #28. Step 1 of the plan discussed for #28: make /healthcheck an accurate report. Detecting frozen PVs with an independent check, and recovery tests, follow separately.

What was wrong

  • The check measured time since a PV's last value change, not since it disconnected. A PV that rarely changes was flagged the moment it disconnected, and disconnected_minutes showed how long since its value last changed. In a test with the IOC paused, /healthcheck returned 503 at the moment the monitor went disconnected.
  • PVs that never connected were never listed. That's the stuck "connecting" state /caget creates and destroys channels with no coordination with monitors #29's race left behind.
  • Any listed PV made the response 503. Both epics2web instances behind the httpd balancer see the same IOCs, so one IOC going down marks both backends down at once.

Changes

  • ChannelMonitor records when its connection state last changed (getStateChanged()), through a single setState that every state change now uses.
  • /healthcheck lists PVs not connected for longer than HEALTHCHECK_GRACE_SECONDS (default 30, as before), counted from that change, or from when monitoring began for a PV that never connected. The JSON array keeps its format, and each entry gains state: DISCONNECTED, or CONNECTING if the PV never connected.
  • Status codes:
    • default: 200 whenever the server is up, for the balancer;
    • ?strict=true: 503 when a PV that was connected has disconnected, for Nagios. PVs that never connected are listed but don't fail strict mode, because they may not exist, such as a mistyped PV name in a client.
  • build.yaml sets HEALTHCHECK_GRACE_SECONDS: 2, and the README documents /healthcheck.

Deployment

  • Nagios: change its check to /epics2web/healthcheck?strict=true to keep alerting on disconnected PVs. Without that, it would always see 200.
  • httpd balancer: keep /epics2web/healthcheck. It now reports a backend as down only when that server is down.

Checks

  • New HealthcheckTest (integration, about 14 s):
    • neverConnectedPvIsListedWithoutFailingStrictMode: a monitored missing PV is listed as CONNECTING, with 200 in both modes.
    • disconnectedPvFailsStrictModeOnly: monitors channel1 (unchanged since the IOC started) for 3 s, then stops the softioc container with docker stop. One second later it isn't listed yet. Once the grace period passes, it's listed as DISCONNECTED, with well under a minute disconnected, 200 by default and 503 in strict mode. After docker start softioc it reconnects and strict mode returns 200. The test skips if the docker command isn't available.
    • Both fail on the server without this change.
  • Hung IOC (manual, docker pause softioc for 45 s): when CA marked HELLO disconnected (about 30 s in), /healthcheck still returned 200 with an empty list. Polled from 41 s on, strict mode returned 503 while the default stayed 200. After the unpause everything reconnected.
  • HealthcheckTest passed 3 runs in a row, and the full integration suite (20 tests) passed twice, on this change stacked on Stop monitors racing their own subscription when the channel closes #47's commit. That commit has the same content as the merged ce1000b, and this branch is rebased onto it. ./gradlew spotlessCheck test passes after the rebase.

🤖 Generated with Claude Code

… strict mode

The healthcheck flagged a PV that wasn't connected once its last value
change was more than 30 seconds old. So a PV that rarely changes was
flagged as soon as it disconnected, its disconnected_minutes showed
time since its value last changed, and a PV that never connected was
never listed. Any listed PV made the response 503. Two epics2web
instances behind a load balancer see the same IOCs, so one IOC going
down took both out of the balancer.

ChannelMonitor now records when its connection state last changed. The
healthcheck lists PVs not connected for longer than
HEALTHCHECK_GRACE_SECONDS (default 30) since that change, or since
monitoring began for a PV that never connected, with their state. It
answers 200 by default. With ?strict=true it answers 503 when a PV that
was connected has disconnected, as monitoring such as Nagios needs. PVs
that never connected are listed but don't fail strict mode, since they
may not exist.

HealthcheckTest covers both kinds of PV. The disconnect case stops and
starts the softioc container. build.yaml sets a 2 second grace period.

Fixes #28

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@slominskir
slominskir merged commit 6952730 into main Oct 4, 2026
6 checks passed
@slominskir
slominskir deleted the healthcheck-connection-time branch October 4, 2026 03:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

source::ai Work done by an AI agent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Healthcheck flags a PV by time since its last value change, not since it disconnected

1 participant