Skip to content

ci(image): make the workbench asset install resilient to pseudo-dojo.org outages #219

Description

@sigilmakes

Problem

The production Workbench image downloads every pseudo-dojo archive from www.pseudo-dojo.org at build time (Dockerfile core-build stage: goldilocks assets install workbench && goldilocks assets verify workbench). No cache, no mirror.

When the runner cannot reach pseudo-dojo.org for ~15 minutes, the build fails. src/goldilocks_core/assets/download.py retries 3× on connect errors, but each attempt blocks the full 300s connect timeout, so one outage window exhausts the retry budget and wastes ~18 minutes before the build dies.

Evidence (PR #218, 2026-09-18 18:14 UTC): requests.exceptions.ConnectTimeout ... 'Connection to www.pseudo-dojo.org timed out. (connect timeout=300)', job burned 20m6s; a rerun of the same commit passed in 4m51s. Scheduled release builds on 09-19 and 09-20 succeeded — the site was reachable, so this is availability, not correctness.

Approach

Pick one (or combine); decide in this issue before implementing.

  1. Cache across runs. Give the image build a buildx cache (cache-from/cache-to: type=gha). The assets-install layer survives across builds when its inputs are unchanged, so the download is skipped on cache hits. Cheapest durable option; keeps pseudo-dojo.org as the source of truth.
  2. Step-level retry. Wrap assets install workbench in a bounded retry (the whole command, not per-request) so one full outage window no longer fails the build. Cheapest immediate fix; still re-downloads on every cold build.
  3. Mirror. Publish the archives to a more reliable host and point src/goldilocks_core/pseudo/registry.toml at it (checksums stay). Most durable; changes the fetch surface for everyone, and needs a licence-aware host (precedent: feat(pseudo): publish the PseudoDojo PBEsol table to PSDI and register it #88 published the PBEsol table to PSDI).

Independent of the choice above: cut the connect timeout. 300s × 4 attempts is the 18-minute burn; a fast-fail timeout (~30–60s) with the existing retries bounds a total outage to a few minutes and makes option 2 viable.

Acceptance criteria

  • A CI image build that hits a pseudo-dojo.org outage either succeeds (cache/mirror) or fails within a bounded, documented time
  • The chosen approach is recorded in this issue with a short rationale
  • just image-e2e passes locally; CI image builds pass

Written by an agent on behalf of Willow Sparks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions