Skip to content

Add recovery tests to CI: IOC restart and hang, and through a CA gateway #56

Description

@slominskir-coding-agent

Related: #28, #48, #50

Why

Frozen PVs after IOC, network and gateway disruptions are why production restarts Tomcat periodically. v2.3.0 fixes several likely causes. But CI checks none of these scenarios, so a regression in how epics2web recovers would go unnoticed. During the v2.3.0 work they were only checked by hand, with temporary scripts.

What was checked by hand

Against docker compose -f build.yaml up, at ec9df53 with JCA 2.4.10, and again with 2.4.12. A raw WebSocket client monitored HELLO (5 updates/s), and each event was followed by an 8 s count of updates and of repeated values:

Event Result
IOC restart (docker stop -t 1 softioc, 5 s, docker start softioc), ×5 Updates resume within about 2 s; no duplicates
IOC hang (docker pause softioc for 40 s, then docker unpause), ×3 About 30 s connected but silent, then disconnected; reconnects within 1 s of the unpause; then every update arrives twice (#48)
The same, through a CA gateway: IOC restart ×3, IOC hang ×2 Recovers; no duplicates (the gateway's C client handles the hang)
Gateway restart ×3 Recovers; no duplicates
Gateway hang ×2 Recovers; duplicates (#48, on epics2web's connection to the gateway)

The frozen-PV detector (#50) flagged nothing in any of these, which is correct.

Proposed tests

Integration tests can already control containers: HealthcheckTest runs docker stop and docker start through ProcessBuilder, and skips without the docker command.

  1. RecoveryTest, against the test IOC directly:

    • IOC restart: updates resume within a few seconds;
    • IOC outage longer than the grace period: the monitor reports disconnected, then recovers;
    • IOC hang (docker pause longer than CA's 30 s echo timeout): the monitor reports disconnected, then updates resume after the unpause;
    • no duplicates afterwards. This fails today because of Monitors deliver every update twice after an IOC circuit becomes unresponsive and recovers (CAJ) #48: mark it @Ignore("#48") until that's fixed, or assert it with the issue named.

    Each hang costs about 30 s, so the class takes about a minute.

  2. The same tests through a CA gateway, as a second CI job or a compose profile. Production gateways run ca-gateway 2.1.3 and 2.1.3.J1. The public images are stale (pklaus/ca-gateway 2020, dmscid/epics-gateway 2016), so build one from epics-extensions/ca-gateway v2.1.3, which needs EPICS base 7.0, pcas and caPutLog. Ideally publish it as a JLab image, like jeffersonlab/softioc, since building EPICS in every CI run would add several minutes. Ask the gateway's owner for its pvlist, access file and command line, to match production.

  3. Network faults, optional: with fixed container IPs, docker network disconnect and docker network connect --ip break and restore the network without the IOC's address changing.

The setup used by hand (pklaus/ca-gateway:1.4, whose gateway can't resolve hostnames, hence the fixed IPs), as an override to build.yaml:

networks:
  default:
    ipam:
      config:
        - subnet: 172.30.0.0/24
services:
  softioc:
    networks:
      default:
        ipv4_address: 172.30.0.10
  gateway:
    image: pklaus/ca-gateway:1.4
    container_name: gateway
    command: ["-cip", "172.30.0.10", "-sport", "5064"]
    environment:
      EPICS_CA_AUTO_ADDR_LIST: "NO"
    networks:
      default:
        ipv4_address: 172.30.0.20
    depends_on:
      - softioc
  epics2web:
    environment:
      EPICS_CA_ADDR_LIST: 172.30.0.20

Done when

Not in scope

  • Changing production gateways or IOCs.
  • VERSION.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions