Summary
Worker node k8s-worker-01 (172.16.101.2) became unreachable on 2025-12-02 ~03:49 MSK, causing multiple pods to be stuck in Terminating state.
Symptoms
- Node status:
NotReady
- Kubelet stopped posting node status at 03:49
- Node unreachable via ping (100% packet loss)
- 19 pods stuck on the node, most in
Terminating state
transmission-0 stuck in Terminating for 43+ hours
- StatefulSet shows 0/1 ready replicas
Affected Services
- transmission (StatefulSet, stuck in Terminating)
- argocd-server
- argocd-notifications-controller
- grafana-deployment
- papermc (StatefulSet)
- authelia
- cloudflare-tunnel
- hello (test deployment)
- longhorn instance-manager
Investigation Done
- Node conditions show
NodeStatusUnknown since 03:49
- Network unreachable (ping timeout to 172.16.101.2)
- All pods on node are in Terminating or problematic state
- Memory was overcommitted at 116% on this node
Potential Root Causes
Based on cluster history (CLAUDE.md errors):
-
NFS mount freeze - transmission-config PVC uses StorageClass truenas which no longer exists in cluster. Mount options unknown, possibly hard mount causing D-state freeze.
-
Watchdog disabled - Per Error 13 in CLAUDE.md, watchdog was disabled on worker node due to previous reboot loop issues. No automatic recovery on hang.
-
OOM / Kernel panic - Node memory was 116% overcommitted (9212Mi limits vs 8124Mi available).
-
Hardware failure - Raspberry Pi specific (SD card, power, thermal).
Next Steps
Related
- Error 14 in CLAUDE.md: NFS hard mount causing node freeze
- Error 13 in CLAUDE.md: Watchdog configuration issues
Summary
Worker node
k8s-worker-01(172.16.101.2) became unreachable on 2025-12-02 ~03:49 MSK, causing multiple pods to be stuck in Terminating state.Symptoms
NotReadyTerminatingstatetransmission-0stuck in Terminating for 43+ hoursAffected Services
Investigation Done
NodeStatusUnknownsince 03:49Potential Root Causes
Based on cluster history (CLAUDE.md errors):
NFS mount freeze -
transmission-configPVC uses StorageClasstruenaswhich no longer exists in cluster. Mount options unknown, possibly hard mount causing D-state freeze.Watchdog disabled - Per Error 13 in CLAUDE.md, watchdog was disabled on worker node due to previous reboot loop issues. No automatic recovery on hang.
OOM / Kernel panic - Node memory was 116% overcommitted (9212Mi limits vs 8124Mi available).
Hardware failure - Raspberry Pi specific (SD card, power, thermal).
Next Steps
journalctl --boot=-1for crash logstransmission-configPV mount options (hard vs soft)Related