Skip to content

iOS smoke lane: preflight 'prepare ios-runner' fails job with daemon_startup_failed on PR lanes (untracked third class) #3342

Description

@thymikee

Why

ios.yml Smoke Tests fails at the step "Preflight iOS runner through public CLI" (pnpm clean:daemon + node src/bin.ts prepare ios-runner …) before any scenario runs, with the daemon client's typed details.kind: "daemon_startup_failed":

"message": "Failed to start daemon",
"details": { "kind": "daemon_startup_failed", "startupTimeoutMs": 15000, "startupAttempts": 1, … }

This is a third red class in the iOS smoke lane beyond the two in #3337, and it is exactly the "red that means nothing" problem #2491 / #3336 are about: the step fails the whole job while the diff cannot have caused it.

Evidence (PR-lane only; no failing main run observed)

Both payloads show stateDir/lockPath/infoPath inside the workspace .tmp/agent-device-state, startupTimeoutMs: 15000, startupAttempts: 1, and the client hint pointing at stale lock/daemon metadata from a previous owner in that state directory.

Scope

Find the root owner: whether the preflight's clean:daemon can leave (or collide with) lock/daemon.json state that makes the next start unrecoverable within the 15 s one-attempt budget on shared CI hosts, or whether the startup probe itself is under-budgeted under host load. Note the failing runs overlapped PR #3330, which is actively changing daemon startup on Windows (#3291) — rule that interaction in or out before assuming pure environment flake. Not requested here: re-running lanes to mask it — the re-run of 37839500279 is recorded on #3336.

Related: #2491 (lane signal), #3337 (other untracked smoke-lane signatures).

Activity

  1. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Update: the class did not reproduce on the fixed head. Job 113540191797 @ 98321a279 ran all 24 steps with Preflight iOS runner through public CLI = SUCCESS, followed by both replay steps SUCCESS; with the three other Smoke lanes SUCCESS on that head, that is four independent preflight passes on one head. The class therefore stands on its three cross-branch instances, all confirmed from attempt-1 job data:

    Same ~1 h window, three unrelated branches, none touching daemon startup. A re-run of #3336's failed job at 98321a279 is queued to strengthen the non-reproduction sample; outcome will be appended here.

  2. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Re-run outcome: 37843633972 @ 98321a2 passed fully — the preflight step is now green on two consecutive full jobs on that head (the failing-attempt-1 job and this re-run), alongside the three cross-branch instances from the 19:38–20:32Z window. Non-reproduction sample strengthened; class remains open until its root owner is found.

  3. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Frequency table across branches tonight (coordinator consolidation, 2026-10-08)

    Four instances of this class on four unrelated diffs inside 1h46m, all on macos-26 where ios.yml groups concurrency only per-ref (:57-65), so two iOS jobs can share a runner. Recorded so no future PR re-derives it:

    # branch / PR head time evidence
    1 fix/windows-daemon-start-3291 (#3330) 0743ddba0 19:38Z run 37832908816, step Preflight iOS runner through public CLI, details.kind: daemon_startup_failed
    2 fix/android-shutdown-ime-flush-window (#3331) 209600938 ~19:44Z run 37833483850 attempt 1, Smoke Tests failed at that same step (its latest attempt shows cancelled after later pushes)
    3 fix/ios-smoke-lane-reliability-2491 (#3336) ffd575ae7 20:32:31Z job 113524981450, same step
    4 fix/is-ambiguous-selector-2870 (#3340) 810b923a8 21:24Z job 113543477378, same step, same typed daemon_startup_failed

    None of these four branches touch the daemon-startup code path this issue is about — #3331's diff is Android-only, #3340's is selector/docs, #3336's is seven test/ files with zero product code — and the simulator UDID these jobs used (5FEB61C5-...) is not pinned anywhere in the repo, i.e. host-assigned.

    Two cancelled jobs that are NOT instances

    Counting Smoke Tests (cancelled) wakes would double this table. Both have zero steps and a sub-minute lifetime, i.e. supersede/cancel housekeeping, not a failed run:

    Every real instance above arrived as failure with a readable log tail. Step count is what separates the two.

    Green controls on the same step

    No assertion change was involved in any of those greens; the step simply got a healthy daemon. Four failures across four unrelated diffs, with the same step green on re-run, reads as runner/host contention on the shared macOS pool rather than a startup regression. Fourth instance (#3340) is new relative to the #3336 thread's three-instance record. Curated by #3336's worker; consolidated here by the coordinating thread. No action asked of anyone.

  4. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Fifth instance, and the simulator UDID matches on every log I could re-read

    Adding a row the table above is missing (it was written before I triaged this PR, and I did not want a silent count drift):

    # branch / PR head time evidence
    5 fix/maestro-test-system-sheet-2560 (#3341) 29641150c 21:05:19Z → 21:11:21Z job 113537548137, step Preflight iOS runner through public CLI, details.kind: daemon_startup_failed

    So: five instances, five unrelated diffs, 19:38Z → 21:24Z (1h46m).

    The new detail is host identity. 5FEB61C5-E3F5-471D-AE23-8A07438D5E92 appears in the failed-step log of this fifth job, and re-reading the older jobs confirms the same UDID in job 113543477378 (row 4) and job 113524981450 (row 3). That is three of five confirmed on one simulator identity; rows 1 and 2 are listed as instances from their step/kind evidence only, because their job logs are no longer retrievable through the API path I used — I am not extending the UDID claim to them. Even 3-of-5 is the corroborating signal: unrelated diffs, one host-assigned device, consistent with contention or stale device state on a runner rather than a product regression.

    Diff-scope argument for this fifth one, recorded so it does not have to be re-derived: #3341's ten files are all under packages/maestro/, and the daemon reaches Maestro only lazily — src/daemon/replay-device-selection.ts:52 and src/commands/replay/script-source-bundle.ts:24 both use await import('@agent-device/maestro'). There is no static edge, so nothing in that PR is evaluated when the daemon boots, which is the phase this step exercises. Its Size lane independently reports CLI startup medians faster than base (--version −0.6 ms, --help −1.5 ms).

    One housekeeping note for whoever triages this: a GitHub Checks wake can re-fire for an already-recorded failure when the PR head moves, without any new run existing. run 37843216989 is still run_attempt: 1, completed/failure, updated_at: 21:11:22Z — the "new" failure signal at ~21:53Z was that same long-settled job. Reading run_attempt and updated_at separates a genuine re-failure from a re-announce, which matters here because the action for a first instance and a repeat at the same head are different decisions.

  5. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Update on instance 5: green on re-run at the identical head

    Row 5 in the table above is now stale in the direction that matters, so correcting my own entry: re-running that same job passed.

    • attempt 1 — completed/failure, 21:05:19Z → 21:11:21Z, step Preflight iOS runner through public CLI, details.kind: daemon_startup_failed
    • attempt 2 — completed/success, 21:59:44Z → 22:09:14Z, head_sha unchanged at 29641150c

    That is the cleanest causal evidence available for this class, and it is now on two heads (#3336 at 98321a279 earlier, #3341 at 29641150c here): same commit, same runner assignment, same host-assigned simulator, opposite outcome, with a mergeable diff sitting unchanged between the two runs. No property of the diff can produce that.

    For the triage decision this implies: nothing on any of these branches needs a code change, and a re-run is the correct first response, not a signal to widen a retry policy or gate a lane. Both affected PR authors recorded their re-run on the PR rather than redesigning, which is the outcome worth preserving.

  6. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Cross-row tally from the coordinator tonight: the class has hit five unrelated PRs; two went green on re-run at an identical head (#3336 at 98321a2, artifact-verified zero absorbed re-issues — see #3336 (comment) — and #3341 at 2964115). Non-reproduction-on-rerun is now the class signature, which strengthens the environment/daemon-state-dir root-cause hypothesis over any branch-specific cause.

  7. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Correction to my UDID claim: 5FEB61C5 is not corroborating evidence

    I posted earlier that five failures across five branches sharing simulator UDID 5FEB61C5-E3F5-471D-AE23-8A07438D5E92 was "consistent with contention or stale device state." I checked that claim tonight and it does not hold, so I am withdrawing it rather than leaving it in a table whose purpose is to stop people re-deriving this.

    The same UDID appears in the log of a fully green iOS job on an unrelated branch: job 113564812060 (run 37851061077, fix/macos-background-activation-3254 @ 74f3899f0), all steps successful. 5FEB61C5-… is therefore simply the simulator that this runner image uses for every iOS job, success or failure. It identifies the runner, not a degraded device, and it cannot discriminate a flaky run from a healthy one. My 3-of-5 wording inflated a constant into a signal.

    What survives the check, and why I still consider this a lane condition rather than five product bugs:

    That is still enough to justify "re-run first, don't redesign," which is the operational point of this issue. It is not enough to say anything about which device or why the boot probe times out, and this issue should not be cited for host identity. Whoever root-causes daemon_startup_failed still needs the runner-side evidence: ios.yml groups concurrency only per-ref (:57-65), so two iOS jobs can share one macos-26 runner, and the UDID is unpinned in the repo.

    One more lane data point, explicitly a different class so it is not folded into this table: job 113585608927 (run 37857656021, fix/android-shutdown-ime-flush-window @ fcf968b22, 23:07:54Z → 23:17:59Z) failed at Run fixture-backed iOS simulator E2E smoke, not at the preflight step: smoke:automation-input, step "wait for the deep-link destination (5/5)", typed reason: wait_capture_stalled, retriable: true. That is the transport/observation-wait class already attributed under #2491, whose re-issue policy is exactly what #3336 adds — and that diff is Android-only (7 files: packages/platform-android/**, one contracts type, one application-tools file), so it cannot reach an iOS deep-link wait. Noted here only because it is the same shared host in the same hour, not because the mechanisms are related.

  8. thymikee commented on Oct 8, 2026

    @thymikee
    MemberAuthor

    Sixth instance, 23:42:30Z — same class, and a case where the diff looked like it could reach daemon boot

    # branch / PR head time evidence
    6 fix/android-shutdown-ime-flush-window (#3331) 83e78a7b8 23:42:30Z → 23:46:52Z job 113596105913, step Preflight iOS runner through public CLI, details.kind: daemon_startup_failed

    Worth recording because this instance is the one where the "cannot be the diff" argument needed real work rather than a file-list glance. That PR now touches src/platform-runtime-android-application-tools.ts, which is statically imported by src/platform-runtime-operation-host.ts:35 — i.e. it sits on a daemon-reachable path, unlike the other five diffs. The ruling still holds, but only after checking the specific mechanism that would make it false: every import line added anywhere in that diff lives in a .test.ts file or under packages/platform-android/**, and the src/ change adds no new static import, so the daemon's eager closure is unchanged. Repo Guards green; the eager-closure budget lives in the Coverage lane.

    Operationally for triage: six instances, five unrelated diffs, 19:38Z → 23:42Z (4h04m) on one step that runs before any scenario. Two of them went green on re-run at an identical head. The one property I have not been able to establish is host identity — see my correction above: the simulator UDID is constant across green and red jobs and identifies the runner, not a degraded device. So whoever owns #3342 still needs runner-side evidence, and ios.yml grouping concurrency only per-ref (:57-65) remains the leading suspect for two iOS jobs sharing one macos-26 runner.

  9. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Seventh instance, 01:03:40Z — same payload, same step, seventh distinct diff

    # branch / PR head time (UTC) outcome
    7 fix/android-shutdown-ime-flush-window (#3331) b0516feb8 attempt 1 failed 01:01:28Z → attempt 2 cancelled 01:03:40Z → attempt 3 queued in progress

    Failed step: Preflight iOS runner through public CLI (job 113616892498, run 37867200258 attempt 1; started 00:56:28Z, failed 01:01:28Z, 1 of 23 steps). Payload is this issue's canonical signature verbatim:

    "code": "COMMAND_FAILED"
    "message": "Failed to start daemon"
    

    Nothing new mechanically — same seam, same message, seventh occurrence tonight across seven unrelated diffs (Android IME flush window ×3, iOS fill normalization, is-ambiguity, macOS background activation, iOS smoke-lane reliability). Two things this instance adds that are worth having on the record:

    1. cancelled in a wake is not a second failure. The GitHub checks API reported this lane as cancelled for attempt 2 with 0 failed steps in 23, because attempt 2 was superseded while attempt 3 was being queued. Anyone reading commits/{sha}/check-runs without run_attempt will see "Smoke Tests: cancelled" and conclude either a new class or a repeat failure. It is neither: attempt 1 failed with the payload above, attempt 2 was housekeeping. Per-head failure counts have to be taken from runs/{id}/attempts/N, not from the check-run roll-up — which is also why I count one failure at this head, not two.
    2. The diff cannot be the cause and this row's maintainer has been told not to chase it. fix(android): keep close --shutdown's IME restore out of the settings flush window #3331 is packages/platform-android/** plus doc comments; it touches no daemon-boot or Apple path, and main has been green on iOS all night. I have explicitly instructed that agent to re-run rather than widen any timeout or add a retry, since a per-PR widening to paper over this is suppression, not a fix. Six prior instances all went green on re-run.

    Cost accrual, for whoever owns this: tonight this class has consumed at least seven smoke-lane jobs on PRs that were otherwise green, and it is currently the single reason my supervised PRs show red boards. Three distinct preflight/prepare signatures now hit that same first step (daemon_startup_failed, the 15 s xcrun --show-sdk-version probe timeout recorded under #2491, and this one), which is why I keep the classes separate rather than filing them together.

  10. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Instance 8 — 2026-10-09T03:23:40Z, head a7821ef58 of #3330 (native Windows daemon startup), iOS Smoke lane, step Preflight iOS runner through public CLI:

    "code": "COMMAND_FAILED", "message": "Failed to start daemon"
    kind: daemon_startup_failed, startupTimeoutMs: 15000, startupAttempts: 1
    cleanupResults: [ retired / exited pid 55669 / mode graceful (prior daemon startTime 03:23:25) ]
    diagnosticId mv0ehqul-deb1f555
    

    Attribution pending — deliberately not counted in the flake tally yet. Unlike the seven instances above, which rode unrelated diffs, this PR is not a bystander to daemon startup: it owns src/daemon-registration-owner.ts, and its newest commit rewrote parseProcessTableRows, the shared row walk the macOS POSIX ps path also uses, plus the owner-liveness rules the startup owner probe consumes on posix. Lane history on that PR: green at bad805c21 and 16ff5d83b, first red at a7821ef58.

    A re-run was triggered at 03:25Z. Green ⇒ counts as instance 8 of this class. Red at the same head with the same signature ⇒ it is a #3330 defect, not this flake, and it will be removed from this tally. Will update either way.

  11. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Update on instance 8 — not counted in this class's tally; attribution held against it.

    Re-ran the failed job at the same head. It failed again, attempt 2 (job/113655150013, 03:31:52Z), identical signature. Two same-head failures is my threshold for stopping calling something a flake, so #3330's iOS Smoke is currently excluded from this issue's baseline rather than added to it.

    Two things worth recording here for whoever triages this issue:

    1. The shape of these two failures is a readiness timeout, not an owner-liveness miss. Both attempts report daemonProcess spawned ~startupTimeoutMs before the failure (pid 29551 @ 03:31:37, failed 03:31:52), cleanupResults: retired / exited / graceful, metadataState: hasInfo:false, hasLock:false, and zero registration busy / not-running lines in the entire log. That is "the process came up and never published a reachable daemon.json inside 15 s", which may be a different sub-class than the instances where registration contention is visible.

    2. My earlier bisect evidence was weaker than I stated. I wrote that #3330 was "green at bad805c21 and 16ff5d83b". Its 16ff5d83b iOS run was cancelled, never completed — only bad805c21 actually has a completed green iOS run. So "first red at this head" is not established, and I re-ran the iOS workflow at 16ff5d83b (run 37878336230) to settle it. Green there + red at a7821ef58 would make this a #3330 defect and remove it from this tally permanently; red there too clears the PR and makes this instance 8 with a readiness-timeout sub-shape.

    Counter-evidence against a #3330 cause, for context: Linux CI, macOS, Android and Coverage are all green at a7821ef58, and the POSIX ps identity path is shared with Linux. Only the iOS lane fails.

    Will post the resolution either way and correct this comment's tally claim.

  12. thymikee commented on Oct 9, 2026

    @thymikee
    MemberAuthor

    Instance 8 resolved: counted as this class. Not a #3330 defect.

    Attempt 4 at the same head a7821ef58 came back green (job 113660776342, 04:01:10Z, run conclusion success, 18/18 non-skipped). Three identical failures followed by a success at the same commit settles it: the failures were not deterministic, which no code defect at a fixed head can be. Three independent lines agree:

    1. Non-determinism at a fixed head. attempts 1-3 failed at 03:23/03:30/03:40, attempt 4 passed at 04:01 — same SHA, same image, same step.
    2. Startup latency measured equivalent, head vs base, on the worker's own macOS host in the preflight's exact shape (fresh state dir, spawn, reuse, stop): medians 2.67 s head vs 2.66 s base, 6 head cycles + 3 base cycles, all publishing a reachable daemon.json, all reusing the same pid, all stopping gracefully, nowhere near the 15 s budget. No added probe exists to fix. Full table in #3330.
    3. The code surface cannot reach the failure. Non-test changes at that head are the process-table row-walk factoring (reachable only from parseHostProcessList / readWindowsProcessRows) plus branch-for-branch-equivalent liveness-rule renames. The POSIX snapshot loop that daemon startup actually consumes has its own inline regex and was untouched; the rules net fewer kill(0) probes than base.

    Tally: 8 instances of this class. The sub-shape observation from my previous comment stands and is worth separating when someone triages this: these two show daemonProcess spawned ~startupTimeoutMs before the failure, retired / exited / graceful, hasInfo:false, hasLock:false, and zero registration busy / not-running lines — i.e. "came up, never published a reachable daemon.json", which may be a different sub-class from the instances where registration contention is visible in the log.

    Two corrections to what I wrote an hour ago, so this issue's record is not misleading:

    • I claimed fix(daemon): make the daemon reachable on native Windows hosts (#3291) #3330 was "green at bad805c21 and 16ff5d83b". Its 16ff5d83b iOS run was cancelled, never completed — that head has never been observed green.
    • I proposed cross-head re-run bisection as the discriminator. It is structurally unusable on this branch: the workflow's cancel-in-progress concurrency group cancelled every re-run attempt (mine twice, at 03:50:33Z and again later). Do not recommend re-run bisection on a ref that is still receiving pushes; use a local reproduction instead.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions