Symptom
Checking out another branch deleted a service's scripts from the working tree
while its timers were live. The unit points straight at the checkout, so it began
failing every minute:
dragonwilds-status.service: Failed at step EXEC spawning
.../linux-server/dragonwilds/dragonwilds-status.sh: No such file or directory
status=203/EXEC
Seven minutes of failures and a stale dashboard before it was noticed. Nothing
alerted, because a timer unit failing is not something anything is watching.
Cause
setup.sh renders ExecStart= to $SCRIPT_DIR/<script>.sh — a path inside the
repo. Any ordinary git operation can therefore take down a running service:
- checking out a branch that does not track those files (what happened here)
- stashing, discarding untracked files, or an interrupted rebase
- deleting or moving the checkout
Scope
Not specific to one service. linux-server/backup, linux-server/forgejo, and
linux-server/dragonwilds all render ExecStart into the checkout the same way,
so a single branch switch can take out backups, the runner monitor, and the game
server's status and auto-update timers at once.
The game server process itself was unaffected, since it runs from outside the
repo — which is exactly the property the scripts lack.
Options
- Install scripts to a stable path —
setup.sh copies to
/usr/local/lib/<service>/ and renders ExecStart there. Units stop depending
on the checkout. The tradeoff is that editing a script in the repo no longer
takes effect until setup.sh is re-run, which is arguably correct for
something systemd runs unattended.
- Symlink
/usr/local/lib/<service> at the checkout. Keeps live editing, but
keeps the failure mode too.
- Leave it, add monitoring — an
OnFailure= handler or an Uptime Kuma check
so a broken timer is at least noticed. Complementary to 1, not an alternative.
Option 1 is the real fix. Worth pairing with 3, since a silently failing timer is
what turned this into seven minutes rather than seven seconds.
Note
Pre-existing pattern rather than a regression from any one PR — it simply had not
been triggered before.
Symptom
Checking out another branch deleted a service's scripts from the working tree
while its timers were live. The unit points straight at the checkout, so it began
failing every minute:
Seven minutes of failures and a stale dashboard before it was noticed. Nothing
alerted, because a timer unit failing is not something anything is watching.
Cause
setup.shrendersExecStart=to$SCRIPT_DIR/<script>.sh— a path inside therepo. Any ordinary git operation can therefore take down a running service:
Scope
Not specific to one service.
linux-server/backup,linux-server/forgejo, andlinux-server/dragonwildsall renderExecStartinto the checkout the same way,so a single branch switch can take out backups, the runner monitor, and the game
server's status and auto-update timers at once.
The game server process itself was unaffected, since it runs from outside the
repo — which is exactly the property the scripts lack.
Options
setup.shcopies to/usr/local/lib/<service>/and rendersExecStartthere. Units stop dependingon the checkout. The tradeoff is that editing a script in the repo no longer
takes effect until
setup.shis re-run, which is arguably correct forsomething systemd runs unattended.
/usr/local/lib/<service>at the checkout. Keeps live editing, butkeeps the failure mode too.
OnFailure=handler or an Uptime Kuma checkso a broken timer is at least noticed. Complementary to 1, not an alternative.
Option 1 is the real fix. Worth pairing with 3, since a silently failing timer is
what turned this into seven minutes rather than seven seconds.
Note
Pre-existing pattern rather than a regression from any one PR — it simply had not
been triggered before.