Skip to content

Confirm in production that v2.3.0 stops frozen PVs, then drop the periodic Tomcat restart #57

Description

@slominskir-coding-agent

Related: #28, #47, #48, #50, #56

Production restarts Tomcat periodically because IOC, network and gateway disruptions have left PVs "frozen" until a restart. v2.3.0 fixes the likely causes and adds detection, but the freezes never reproduced in testing. This issue collects the production evidence, and the decision to drop the restart.

Evidence so far, from production logs

  • Gateway logs (April): bad resource id ... casStrmClient.cc ... CAS Request: acetomcat on epicswebops9: cmd=2 ... Bad resource identifier - unexpected problem with client's input - forcing disconnect, then a pcas assertion: the expression "this->eventLogQue.count() == 0" didnt evaluate to boolean true.
    • cmd=2 is EVENT_CANCEL: epics2web's Tomcat cancelled a subscription the gateway didn't know. The gateway then dropped epics2web's whole connection, so every PV through that gateway disconnected at once.
    • This is the race fixed in Stop monitors racing their own subscription when the channel closes #47: CAJ sends cancels at once but queues subscription requests, so a cancel can arrive first.
    • In the code production ran, monitors closing at that moment could stay stuck as disconnected until a restart.
    • The most likely cause found so far.
  • Gateway logs: Identical process variable names on multiple servers. For example, VIP1L07BC.SEVR is served by both opsbat6.acc.jlab.org:37825 and iocnl1b.acc.jlab.org:5064. The gateway uses whichever server answers first. If one copy doesn't update, the PV looks frozen, and which copy wins can change after restarts. epics2web can't detect or fix this, because its independent check also goes through the gateway. It's for the IOCs' owners.
  • epics2web logs:
    • no resubscribeSubscriptions or IndexOutOfBoundsException, so the JCA bug fixed in 2.4.12 (resubscribe failure epics-base/jca#86) isn't confirmed here;
    • User destroyed channel for comm8 is /caget requests timing out, which is noise.
  • Production gateways run ca-gateway 2.1.3 and 2.1.3.J1.

Deploying v2.3.0

Per the release notes:

  • Test instance first; check WEDM screens, the console and /caget users.
  • Upgrade Tomcat to 11.0.26.
  • Nagios: switch to /epics2web/healthcheck?strict=true. The load balancer stays on /epics2web/healthcheck.
  • One production instance at a time; compare the two for a few days.
  • Keep the periodic restart for now.

What to watch

  • epics2web logs: WARNING ... PV <name> is frozen: <reason>, especially around IOC and network maintenance. The reason says which kind of failure it was.
  • Gateway logs: forcing disconnect naming epicsweb hosts. Expect none from upgraded instances. While one instance is still old, compare the two.
  • Console: "Channel Access Channels" should stay close to "Unique PVs (Monitors)". A steady climb would mean leaking channels.
  • Users: reports of frozen or stale values, with the PV name and time, to compare with the above.

Questions for the gateways' owner

Deciding

  • Nothing frozen and no forcing disconnect from epicsweb hosts over a few maintenance windows: stop the periodic restart. Optionally trial /epics2web/healthcheck?frozen=true, which answers 503 when a PV is frozen, as the trigger for an automatic restart.
  • Frozen warnings appear: their reasons, times and PVs point at the next fix. One option is for epics2web to recreate just the frozen PV's monitor instead of restarting Tomcat. That would be off by default, and turned on once the detector has a track record.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

source::aiWork done by an AI agent

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions