Repository navigation
Report PVs by time since disconnect, and fail the healthcheck only in strict mode - #49
Merged
Merged
Conversation
… strict mode The healthcheck flagged a PV that wasn't connected once its last value change was more than 30 seconds old. So a PV that rarely changes was flagged as soon as it disconnected, its disconnected_minutes showed time since its value last changed, and a PV that never connected was never listed. Any listed PV made the response 503. Two epics2web instances behind a load balancer see the same IOCs, so one IOC going down took both out of the balancer. ChannelMonitor now records when its connection state last changed. The healthcheck lists PVs not connected for longer than HEALTHCHECK_GRACE_SECONDS (default 30) since that change, or since monitoring began for a PV that never connected, with their state. It answers 200 by default. With ?strict=true it answers 503 when a PV that was connected has disconnected, as monitoring such as Nagios needs. PVs that never connected are listed but don't fail strict mode, since they may not exist. HealthcheckTest covers both kinds of PV. The disconnect case stops and starts the softioc container. build.yaml sets a 2 second grace period. Fixes #28 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
slominskir
approved these changes
Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #28. Step 1 of the plan discussed for #28: make
/healthcheckan accurate report. Detecting frozen PVs with an independent check, and recovery tests, follow separately.What was wrong
disconnected_minutesshowed how long since its value last changed. In a test with the IOC paused,/healthcheckreturned 503 at the moment the monitor went disconnected./cagetcreates and destroys channels with no coordination with monitors #29's race left behind.Changes
ChannelMonitorrecords when its connection state last changed (getStateChanged()), through a singlesetStatethat every state change now uses./healthchecklists PVs not connected for longer thanHEALTHCHECK_GRACE_SECONDS(default 30, as before), counted from that change, or from when monitoring began for a PV that never connected. The JSON array keeps its format, and each entry gainsstate:DISCONNECTED, orCONNECTINGif the PV never connected.?strict=true: 503 when a PV that was connected has disconnected, for Nagios. PVs that never connected are listed but don't fail strict mode, because they may not exist, such as a mistyped PV name in a client.build.yamlsetsHEALTHCHECK_GRACE_SECONDS: 2, and the README documents/healthcheck.Deployment
/epics2web/healthcheck?strict=trueto keep alerting on disconnected PVs. Without that, it would always see 200./epics2web/healthcheck. It now reports a backend as down only when that server is down.Checks
HealthcheckTest(integration, about 14 s):neverConnectedPvIsListedWithoutFailingStrictMode: a monitored missing PV is listed asCONNECTING, with 200 in both modes.disconnectedPvFailsStrictModeOnly: monitorschannel1(unchanged since the IOC started) for 3 s, then stops thesoftioccontainer withdocker stop. One second later it isn't listed yet. Once the grace period passes, it's listed asDISCONNECTED, with well under a minute disconnected, 200 by default and 503 in strict mode. Afterdocker start softiocit reconnects and strict mode returns 200. The test skips if thedockercommand isn't available.docker pause softiocfor 45 s): when CA markedHELLOdisconnected (about 30 s in),/healthcheckstill returned 200 with an empty list. Polled from 41 s on, strict mode returned 503 while the default stayed 200. After the unpause everything reconnected.HealthcheckTestpassed 3 runs in a row, and the full integration suite (20 tests) passed twice, on this change stacked on Stop monitors racing their own subscription when the channel closes #47's commit. That commit has the same content as the merged ce1000b, and this branch is rebased onto it../gradlew spotlessCheck testpasses after the rebase.🤖 Generated with Claude Code