Skip to content

Time out deploy commands after 10 minutes - #19

Merged
slominskir merged 1 commit into
mainfrom
deploy-timeout
Oct 1, 2026
Merged

slominskir merged 1 commit into
mainfrom
deploy-timeout

Conversation

@slominskir-coding-agent

Copy link
Copy Markdown
Contributor

Fixes #1.

What changed

  • Timeout: SSHFacade waits at most 10 minutes (commandTimeout) for the deploy command. After that it closes the channel and the job records a SocketTimeoutException as its stack trace, with no exit code, and the stdout and stderr received so far.
  • No transaction while the command runs: asyncExecuteRemoteCommand is now @TransactionAttribute(NOT_SUPPORTED). Before, it ran in a container transaction for the whole command. The jeffersonlab/wildfly:3.0.1 image keeps Wildfly's default transaction timeout of 300 seconds, so a command running longer than that would leave the transaction rolled back, and saving the job at the end would fail. deployJobFacade.edit(job) now saves the job in its own transaction.

The 10 minutes is a hard-coded constant, like the other SSH timeouts. If some deploys need longer, it could become an environment variable or a column on APP_ENV.

No deployment steps.

What happens on the remote host

The sshd library has no client API to signal the remote process, so the timeout only closes the channel:

  • A command waiting for input, like the certificate prompt in the issue, gets end-of-input and exits.
  • A command that is busy, like sleep, keeps running on the remote host until it ends on its own.

Checks

  • gradle spotlessApply build passes (JDK 21, in the gradle:9-jdk21 container).
  • I ran the real executeRemoteCommand (through reflection, with commandTimeout set to 3 seconds) against the project's Dockerfile-sshd container:
deploy_command (version 1.0.0 appended) Time Exit code Stdout / stderr Stack trace
touch /tmp/hello && echo 1.2 s 0 1.0.0 none
sh -c 'echo out; echo err >&2; exit 3' # 0.3 s 3 out / err none
sh -c 'echo partial; sleep 30; echo done' # 3.3 s none partial SocketTimeoutException
sh -c 'echo asking; read -r x; echo got:$x' # 3.2 s none asking SocketTimeoutException

After the run, the sleep command was still running in the sshd container; the read command had exited.

Not checked: the transaction change and the /log page in the full Compose stack. Port 8081 on the VM is taken by another project's Keycloak.

🤖 Generated with Claude Code

A deploy command could hang forever, such as on a prompt to accept a
certificate, holding its thread with the job never finished. Wait at
most 10 minutes for the command, then close its channel and record a
SocketTimeoutException with the output so far.

The async deploy also no longer runs in a transaction: Wildfly's
default transaction timeout is 300 seconds, so a command running
longer than that left a rolled-back transaction and the job was never
saved. The job is saved by deployJobFacade.edit in its own transaction.

Fixes #1

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@slominskir-coding-agent slominskir-coding-agent Bot added the enhancement New feature or request label Oct 1, 2026
@slominskir-coding-agent slominskir-coding-agent Bot added the enhancement New feature or request label Oct 1, 2026

@slominskir slominskir left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some C++ app deploys build from source, and that can take a very long time, especially if Clang Tidy static analysis is repeated (generally done in CI). We'll likely need to revisit this hard-coded 10 minute timeout, but good enough for now.

@slominskir
slominskir merged commit a1a8c85 into main Oct 1, 2026
5 checks passed
This was referenced Oct 1, 2026
@slominskir-coding-agent
slominskir-coding-agent Bot deleted the deploy-timeout branch October 1, 2026 19:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add timeout on deploy

1 participant