You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #581, which covers transient D1 errors and the undocumented retry contract. This is a different failure: verification on the server stalls or dies, and the client has no way to recover the asset.
What happens
On 2026-09-30, three publishes of ferry-cli (a cli-binary app, asset ingest protocol v1, from GitHub Actions) failed at POST /api/apps/{appId}/builds/{buildId}/assets/{assetId}/upload/complete. Each time it was a different asset. Twenty minutes before the first failure, 0.1.6 published the same set of 12 assets without trouble: each ~135 MB binary was uploaded and completed in 15–20 s.
complete sent no response headers for 5 minutes (09:52:2x → 09:57:25 UTC; the client's fetch failed is undici's 300 s header timeout). The retry got 409 ASSET_UPLOAD_BUSY.
complete answered 422 {"error":"failed to verify and seal asset","code":"ASSET_UPLOAD_SEAL_FAILED"} at 10:10:17 UTC. We uploaded exactly the bytes whose SHA-256 and size we declared, and nothing says what failed.
complete timed out after 180 s (10:12:2x → 10:15:26 UTC). We then polled complete every 15 s for 10 minutes and got 409 ASSET_UPLOAD_BUSY every time, until we gave up at 10:22:20.
The re-run started a new build, so its idempotency keys were new, and the three assets before ferry-darwin-x64 completed in normal time. So a stuck asset doesn't seem to be a poisoned key or a bad object. Some verifications just never finish, and some fail without saying why.
Why it matters
We can't tell a slow verification from a dead one, and there is no way to get that asset back:
ASSET_UPLOAD_BUSY has no upper bound. Nothing shows whether the lease will expire, whether its holder is still alive, or when it would be reasonable to stop waiting.
Once the lease is held, re-declaring and re-uploading doesn't help, because the asset is stuck until the lease clears.
ASSET_UPLOAD_SEAL_FAILED has no detail (hash mismatch? size? storage read error?), so we can't tell whether a retry could succeed.
The only recovery left to a client is to abandon the whole build (mark it failed) and start again with a new build and re-upload every asset. That is what Ferry now does (see below).
What would help
Verification leases that expire. Document how long ASSET_UPLOAD_BUSY can last. Once the holder is gone, a later complete should take over, or answer something like ASSET_UPLOAD_STALE so the client knows to upload the asset again.
Error detail for ASSET_UPLOAD_SEAL_FAILED: expected vs. actual SHA-256 and size, or "storage read failed". Also say whether a retry of the same asset can succeed.
Expected verification time per size, so clients can choose their timeouts. 0.1.6 had ~15 s for 135 MB; the failures above waited 5–10 minutes.
What Ferry does meanwhile
tools/release/src/publish-hands.ts in botiverse/ferry already polls a busy asset for up to 10 minutes (#581). From the next release, when an asset is still busy after that, or fails with ASSET_UPLOAD_SEAL_FAILED, it marks the build failed and starts a new build, up to three times. The abandoned builds for 0.1.8 are marked failed. The 0.1.7 build predates that code and may still be pending.
Follow-up to #581, which covers transient D1 errors and the undocumented retry contract. This is a different failure: verification on the server stalls or dies, and the client has no way to recover the asset.
What happens
On 2026-09-30, three publishes of
ferry-cli(acli-binaryapp, asset ingest protocol v1, from GitHub Actions) failed atPOST /api/apps/{appId}/builds/{buildId}/assets/{assetId}/upload/complete. Each time it was a different asset. Twenty minutes before the first failure, 0.1.6 published the same set of 12 assets without trouble: each ~135 MB binary was uploaded and completed in 15–20 s.7031f106-2a31-4741-996d-73a946e711ccferry-darwin-arm64(133.9 MB), asset14118511-64d2-4a2a-a95c-68b98b91517ecompletesent no response headers for 5 minutes (09:52:2x → 09:57:25 UTC; the client'sfetch failedis undici's 300 s header timeout). The retry got409 ASSET_UPLOAD_BUSY.d528ac3c-6a6c-43da-8543-438e46898490ferry-darwin-arm64.gz(42.9 MB), assetd1f29cbf-395f-4a65-8d90-9ce6c7171f1bcompleteanswered422 {"error":"failed to verify and seal asset","code":"ASSET_UPLOAD_SEAL_FAILED"}at 10:10:17 UTC. We uploaded exactly the bytes whose SHA-256 and size we declared, and nothing says what failed.ae304ea1-ea1d-45d8-9a11-5a7ab7247167ferry-darwin-x64(136.3 MB), assetf9db7547-dcf1-4616-8259-b8a3a40e300bcompletetimed out after 180 s (10:12:2x → 10:15:26 UTC). We then polledcompleteevery 15 s for 10 minutes and got409 ASSET_UPLOAD_BUSYevery time, until we gave up at 10:22:20.The re-run started a new build, so its idempotency keys were new, and the three assets before
ferry-darwin-x64completed in normal time. So a stuck asset doesn't seem to be a poisoned key or a bad object. Some verifications just never finish, and some fail without saying why.Why it matters
We can't tell a slow verification from a dead one, and there is no way to get that asset back:
ASSET_UPLOAD_BUSYhas no upper bound. Nothing shows whether the lease will expire, whether its holder is still alive, or when it would be reasonable to stop waiting.ASSET_UPLOAD_SEAL_FAILEDhas no detail (hash mismatch? size? storage read error?), so we can't tell whether a retry could succeed.The only recovery left to a client is to abandon the whole build (mark it
failed) and start again with a new build and re-upload every asset. That is what Ferry now does (see below).What would help
ASSET_UPLOAD_BUSYcan last. Once the holder is gone, a latercompleteshould take over, or answer something likeASSET_UPLOAD_STALEso the client knows to upload the asset again.ASSET_UPLOAD_SEAL_FAILED: expected vs. actual SHA-256 and size, or "storage read failed". Also say whether a retry of the same asset can succeed.GET …/assets/{assetId}returninguploading | verifying (since …) | ready | failed (why). Or the asynchronouscompletesuggested in Asset upload complete: transient D1 errors answer 409, and retrying a long verification is undocumented #581. Either way a client can wait on something that has a known end.What Ferry does meanwhile
tools/release/src/publish-hands.tsin botiverse/ferry already polls a busy asset for up to 10 minutes (#581). From the next release, when an asset is still busy after that, or fails withASSET_UPLOAD_SEAL_FAILED, it marks the buildfailedand starts a new build, up to three times. The abandoned builds for 0.1.8 are markedfailed. The 0.1.7 build predates that code and may still bepending.