You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Confirm in production that v2.3.0 stops frozen PVs, then drop the periodic Tomcat restart #57
Production restarts Tomcat periodically because IOC, network and gateway disruptions have left PVs "frozen" until a restart. v2.3.0 fixes the likely causes and adds detection, but the freezes never reproduced in testing. This issue collects the production evidence, and the decision to drop the restart.
Evidence so far, from production logs
Gateway logs (April):bad resource id ... casStrmClient.cc ... CAS Request: acetomcat on epicswebops9: cmd=2 ... Bad resource identifier - unexpected problem with client's input - forcing disconnect, then a pcas assertion: the expression "this->eventLogQue.count() == 0" didnt evaluate to boolean true.
cmd=2 is EVENT_CANCEL: epics2web's Tomcat cancelled a subscription the gateway didn't know. The gateway then dropped epics2web's whole connection, so every PV through that gateway disconnected at once.
In the code production ran, monitors closing at that moment could stay stuck as disconnected until a restart.
The most likely cause found so far.
Gateway logs: Identical process variable names on multiple servers. For example, VIP1L07BC.SEVR is served by both opsbat6.acc.jlab.org:37825 and iocnl1b.acc.jlab.org:5064. The gateway uses whichever server answers first. If one copy doesn't update, the PV looks frozen, and which copy wins can change after restarts. epics2web can't detect or fix this, because its independent check also goes through the gateway. It's for the IOCs' owners.
epics2web logs:
no resubscribeSubscriptions or IndexOutOfBoundsException, so the JCA bug fixed in 2.4.12 (resubscribe failure epics-base/jca#86) isn't confirmed here;
User destroyed channel for comm8 is /caget requests timing out, which is noise.
Production gateways run ca-gateway 2.1.3 and 2.1.3.J1.
Nothing frozen and no forcing disconnect from epicsweb hosts over a few maintenance windows: stop the periodic restart. Optionally trial /epics2web/healthcheck?frozen=true, which answers 503 when a PV is frozen, as the trigger for an automatic restart.
Frozen warnings appear: their reasons, times and PVs point at the next fix. One option is for epics2web to recreate just the frozen PV's monitor instead of restarting Tomcat. That would be off by default, and turned on once the detector has a track record.
Related: #28, #47, #48, #50, #56
Production restarts Tomcat periodically because IOC, network and gateway disruptions have left PVs "frozen" until a restart. v2.3.0 fixes the likely causes and adds detection, but the freezes never reproduced in testing. This issue collects the production evidence, and the decision to drop the restart.
Evidence so far, from production logs
bad resource id ... casStrmClient.cc ... CAS Request: acetomcat on epicswebops9: cmd=2 ... Bad resource identifier - unexpected problem with client's input - forcing disconnect, then a pcas assertion:the expression "this->eventLogQue.count() == 0" didnt evaluate to boolean true.cmd=2isEVENT_CANCEL: epics2web's Tomcat cancelled a subscription the gateway didn't know. The gateway then dropped epics2web's whole connection, so every PV through that gateway disconnected at once.Identical process variable names on multiple servers. For example,VIP1L07BC.SEVRis served by bothopsbat6.acc.jlab.org:37825andiocnl1b.acc.jlab.org:5064. The gateway uses whichever server answers first. If one copy doesn't update, the PV looks frozen, and which copy wins can change after restarts. epics2web can't detect or fix this, because its independent check also goes through the gateway. It's for the IOCs' owners.resubscribeSubscriptionsorIndexOutOfBoundsException, so the JCA bug fixed in 2.4.12 (resubscribe failure epics-base/jca#86) isn't confirmed here;User destroyed channelforcomm8is/cagetrequests timing out, which is noise.ca-gateway2.1.3 and 2.1.3.J1.Deploying v2.3.0
Per the release notes:
/cagetusers./epics2web/healthcheck?strict=true. The load balancer stays on/epics2web/healthcheck.What to watch
WARNING ... PV <name> is frozen: <reason>, especially around IOC and network maintenance. The reason says which kind of failure it was.forcing disconnectnaming epicsweb hosts. Expect none from upgraded instances. While one instance is still old, compare the two.Questions for the gateways' owner
forcing disconnectfrom epicsweb hosts appear before v2.3.0? Did it line up with frozen-PV reports?Identical process variable namesduplicates, for the IOCs' owners.Deciding
forcing disconnectfrom epicsweb hosts over a few maintenance windows: stop the periodic restart. Optionally trial/epics2web/healthcheck?frozen=true, which answers 503 when a PV is frozen, as the trigger for an automatic restart.