Skip to content

v2.3.0 - #54

Merged
slominskir merged 1 commit into
mainfrom
release-2.3.0
Oct 4, 2026
Merged

slominskir merged 1 commit into
mainfrom
release-2.3.0

Conversation

@slominskir-coding-agent

Copy link
Copy Markdown
Contributor

Bumps VERSION to 2.3.0. Merging this releases v2.3.0: the CD workflow, limited to main by #53, tags the release, attaches the war and publishes the Docker image. Pushing this branch didn't start a CD run.

The upgrade notes below are for pasting into the GitHub release.


Upgrade notes

This release hardens the WebSocket monitor, /caget and /healthcheck. It fixes several ways PVs could get stuck after IOC, network or gateway disruptions, and adds detection of PVs that freeze anyway.

Action needed when upgrading

  • Nagios, or any monitoring that alerts on /healthcheck: use /epics2web/healthcheck?strict=true. The default /healthcheck now answers 200 whenever the server is up, even if PVs are disconnected. That's so a load balancer doesn't take every instance out when an IOC goes down. Strict mode answers 503 when a PV that was connected has stayed disconnected past the grace period.
  • Load balancers: keep using /epics2web/healthcheck.
  • Tomcat: tested with Tomcat 11.0.26, which the Docker image now pins. That version includes WebSocket fixes relevant to this release, such as write timeouts covering whole messages. Upgrading from earlier 11.0.x releases is recommended.

Behaviour changes

  • Unresponsive clients are closed. The server pings WebSocket clients every 30 s, and closes a client that sends no message and no pong for 60 s, or whose writes stay blocked for 20 s. Browsers answer pings automatically, and epics2web.js reconnects.
  • Clients that fall behind get the latest values. A client that falls behind gets each PV's latest update instead of a backlog. Connect and disconnect (info) messages are never dropped.
  • /caget validates requests:
    • it answers HTTP 400 for more than 500 PVs in one request, or for a jsonp callback that isn't a JavaScript name such as a.b_c;
    • a PV that doesn't connect now fails only the request that asked for it, with an error naming the PV.
  • /healthcheck reports more accurately:
    • it counts the grace period from when a PV disconnected, not from its last value change;
    • entries gain a state field;
    • PVs that never connected are listed, as CONNECTING, but don't fail strict mode;
    • frozen PVs are listed with frozen, frozen_minutes and frozen_reason;
    • ?frozen=true answers 503 when any PV is frozen.
  • No session cookies: responses no longer set a JSESSIONID cookie, except for WebSocket handshakes.
  • Console changes: the console lists connected sessions that don't monitor anything yet, and shows the number of open Channel Access channels.
  • -Infinity is now reported with its sign.

Frozen PV detection (new, report-only)

epics2web now looks for PVs whose monitor has stopped working while the IOC still serves them. It gives suspicious PVs a short-lived subscription in a second, independent CA context, which opens its own connection to each IOC or gateway and probes at most 20 PVs at a time.

  • A frozen PV is logged as WARNING ... PV <name> is frozen: <reason>.
  • It's listed by /healthcheck.
  • Through a gateway, it finds problems between epics2web and the gateway, not inside the gateway.

It doesn't change the default or strict healthcheck responses.

New settings (environment variables, all optional)

Variable Default Meaning
WEBSOCKET_PING_INTERVAL_SECONDS 30 How often clients are pinged
WEBSOCKET_TIMEOUT_SECONDS 60 Close a client with no message or pong for this long
WEBSOCKET_SEND_TIMEOUT_SECONDS 20 Close a client whose write blocks this long (Tomcat only)
HEALTHCHECK_GRACE_SECONDS 30 How long a PV may be disconnected before /healthcheck lists it
FROZEN_CHECK_SECONDS 10 How often to look for frozen PVs
FROZEN_PV_CHECK true false turns frozen PV detection off

Fixes

Known issue

Suggested rollout

  1. Test first: deploy to a test instance, and check WEDM screens, the console and /caget users.

  2. One production instance at a time: with two instances behind a load balancer, upgrade one, compare the two for a few days, then upgrade the other.

  3. Keep any periodic Tomcat restart for now, and watch for:

    • is frozen warnings, especially around IOC and network maintenance;
    • forcing disconnect messages naming epics2web hosts in gateway logs, which should stop with this release;
    • a "Channel Access Channels" count on the console that keeps climbing above "Unique PVs (Monitors)".

    If none of these show up over a few maintenance windows, the periodic restart can go.


🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@slominskir
slominskir merged commit b26aa41 into main Oct 4, 2026
6 checks passed
@slominskir
slominskir deleted the release-2.3.0 branch October 4, 2026 14:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

source::ai Work done by an AI agent

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant