Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 20 additions & 13 deletions docs/CLI.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,7 +153,8 @@ granted.
- `--allow-download` — permit downloading a missing runtime / system image
(multi-GB; never implicit). Without it, a missing runtime is exit 12.
iOS runtimes remain Xcode-managed in v1: `--allow-download` cannot install
them; install the runtime through Xcode first.
them; install the runtime through Xcode first. Through a gateway the flag
has no effect.
- `--ttl <duration>` — the lease's initial TTL, replacing
`lease.defaultTtlMs` (15m) for this lease. Asking for more than
`lease.maxTtlMs` (4h) is a `BAD_REQUEST` (exit 2), not a silent clamp. See
Expand Down Expand Up @@ -567,12 +568,19 @@ differs.

**Leasing is identical.** `simlock lease` takes the same flags and prints the
same grant line, including `--ttl`, `--no-wait`, `--timeout`, and
`--allow-download` (forwarded to the chosen worker, which clamps it through
its own `downloads.policy`). The request waits in the gateway's own
fleet-wide FIFO queue, reporting `queued` with a `queuePosition` exactly as a
worker's queue does, and is dispatched to the worker best placed to serve it
— a machine with a matching warm device first, otherwise the one with the
most free capacity. You do not name a machine and there is no flag to; where
`--allow-download`, which is accepted and has no effect through a gateway:
only runtimes already installed on a worker count, and no download is
started. The request waits in the gateway's own fleet-wide FIFO queue,
reporting `queued` with a `queuePosition` exactly as a worker's queue does.

The gateway sends a request only to a worker that can serve it. `--device`
matches a worker's model in any letter case and by any other name that
worker's catalog lists for it, such as an Android AVD id (`pixel_7`). When
`--os` is given, the worker must pair that runtime with the model; without
it, the model must pair with at least one installed runtime. A worker that
has the model and the runtime but cannot pair them is passed over. Among the
workers that can serve it, the request goes to a machine with a matching warm
device first, otherwise the one with the most free capacity. You do not name a machine and there is no flag to; where
a device lives is the gateway's decision.

The grant carries one additional block so you can see where it landed:
Expand Down Expand Up @@ -628,9 +636,8 @@ runtime annotated with the workers that have it — so a `--device` the
catalog lists is leasable *somewhere*, not necessarily everywhere. A model is
paired with a runtime when at least one connected worker pairs them itself;
one worker having the model and another having the runtime does not make a
pair. `simlock worker list --json` shows each worker's own pairings. The
gateway does not yet pick a worker by its pairings, so a listed pair can
still be sent to a worker that cannot pair them.
pair. `simlock worker list --json` shows each worker's own pairings. A
request goes only to a worker that pairs the model with the runtime.

**`simlock events`** shows the fleet: every worker's business events are
republished on the gateway's bus with `workerId` added to the payload,
Expand Down Expand Up @@ -803,9 +810,9 @@ why its Android catalog looks thin, trimmed to one worker:
```

`downloads.policy` and `lease.maxTtlMs` are that worker's own effective
config, read when its uplink connects and again on every periodic refresh —
routing needs the policy to know whether a machine may install a missing
runtime before sending it a request that needs one. `catalog` is what that
config, read when its uplink connects and again on every periodic refresh.
The policy is shown for reference; routing does not read it, since no
download is started through a gateway. `catalog` is what that
worker can lease, each model with the runtimes it pairs with. `host` is the
machine: operating system, its version, CPU architecture, and the version of
each platform tool its drivers use (`xcode` with its build; the Android
Expand Down
11 changes: 6 additions & 5 deletions docs/CLIENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -217,11 +217,12 @@ const { platforms } = await client.getCatalog({ platform: "ios" });
Against a gateway the catalog is the union of the connected workers'. A
model is paired with a runtime when at least one worker pairs them, and
`modelWorkers` and `runtimeWorkers` say which workers have each model and
runtime. The gateway does not yet use the pairings to pick a worker, so a
pair it lists can still go to a worker that has the model and the runtime
but cannot pair them, and that request fails there. `modelAliases` and
`images` are the unions of each worker's own. The gateway does not yet route
by another name, so ask it for a model by its name in `models`.
runtime. The gateway sends a request only to a worker that pairs the model
with the runtime. `modelAliases` and `images` are the unions of each
worker's own. A model may be asked for by any name a worker lists for it, in
any letter case, and the gateway sends that worker its own name for it.
`allowDownload` has no effect through a gateway: only installed runtimes
count.

## What machine answered: `getStatus().host`

Expand Down
10 changes: 9 additions & 1 deletion docs/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ a warning. Inspect the effective, merged configuration at any time with
| `gateway.token` | **Worker side.** The join token (`simlock token create --role worker`, minted on the gateway) this worker presents when it opens its uplink. Required whenever `gateway.url` is set. | unset |
| `gateway.label` | **Worker side.** Display name for this worker in `simlock worker list`, `status`, the console, and on the lease's `worker` block. Display-only: nothing routes on it and it need not be unique. | the worker's own id |
| `exec.timeoutMs` | **Worker side.** How long one `device.exec` command (`simlock simctl` / `simlock adb` against a gateway or over HTTP) may run before the worker kills it and the operation fails with `EXEC_TIMEOUT`. Authoritative: it bounds the process that actually runs. | `10 minutes` |
| `gateway.routing` | **Gateway side.** Which routing policy places a queued request on a worker. `warm-then-free` is the only policy in v1: warm hit first, then the most free running capacity for the platform. | `warm-then-free` |
| `gateway.routing` | **Gateway side.** Which routing policy places a queued request on a worker. `warm-then-free` is the only policy in v1: among workers that can serve the request, warm hit first, then the most free running capacity for the platform. | `warm-then-free` |
| `gateway.disconnectedRetentionMs` | **Gateway side.** How long a disconnected worker is kept (greyed, never dispatched to) before the gateway forgets it. The clock is held while the gateway still knows of gateway-issued leases on that worker, and that hold ends when the last of those leases passes its deadline. | `24 hours` |
| `gateway.execTimeoutMs` | **Gateway side.** How long the gateway waits on a proxied `device.exec` before giving up. A backstop for a worker that never answers at all — deliberately longer than the worker's own `exec.timeoutMs`, which is authoritative because that side owns the process and can kill it, so an ordinary timeout surfaces as the worker's `EXEC_TIMEOUT` rather than racing this one. | `11 minutes` |
| `gateway.leaseRequestTimeoutMs` | **Gateway side.** How long the gateway waits on a forwarded `lease.request` before giving up on that worker for this request, answering `WORKER_UNREACHABLE`. Bounds the one uplink call that otherwise had no timeout of its own, so a wedged worker cannot park a request where neither a deadline nor `lease.cancel` could ever reach it again. Generous against a cold device provision-plus-boot; well below `gateway.execTimeoutMs`, since granting a lease should never take as long as a command run against the device afterward. | `5 minutes` |
Expand Down Expand Up @@ -207,6 +207,14 @@ does not apply to it. It reads:
| `lease.*` | `defaultTtlMs`/`maxTtlMs` bound what its own clients may ask for, before a request is dispatched — see below |
| `log.*`, `eventBuffer.*`, `eventLog.*` | logging and the event history, as anywhere |

**A worker's `downloads.policy` does not apply to requests through a
gateway.** The gateway sends a request only to a worker whose catalog already
has what it asks for: the model under any name the worker lists for it, in
any letter case, paired with the requested runtime (or, with none requested,
with at least one installed runtime). It never asks a worker to download, so
`--allow-download` has no effect through a gateway, whatever each worker's
policy says.

**Both ends have a `lease.*` block, and on a fleet lease the gateway's is the
one that decides the width.** A request arriving at a gateway with no `ttlMs`
is filled in with the *gateway's* `lease.defaultTtlMs` before it is dispatched
Expand Down
25 changes: 16 additions & 9 deletions docs/HTTP-API.md
Original file line number Diff line number Diff line change
Expand Up @@ -187,9 +187,8 @@ Each platform entry carries `modelRuntimes`: for every name in `models`, the
installed runtimes that model pairs with. A pair listed there can be leased;
a model and a runtime that are each listed but not paired cannot. An empty
list means nothing installed pairs with that model. On a gateway a model is
paired with a runtime when at least one connected worker pairs them; the
gateway does not yet pick a worker by its pairings, so such a request can
still go to a worker that cannot pair them and fail there.
paired with a runtime when at least one connected worker pairs them, and the
gateway sends a request only to a worker that pairs them.

Each entry also carries `modelAliases`: for a name in `models`, the other
names a lease request may use for it, in any letter case. Only models that
Expand All @@ -199,8 +198,9 @@ system image with its API level (`runtime`, a value from `runtimes`), `tag`,
and `abi`, including an image whose ABI the host cannot run natively; an iOS
entry has no `images`. On a gateway `modelAliases` is the union per model and
`images` the union of each worker's images, absent when no worker's entry
for that platform has an `images` field. The gateway does not yet route by another name: send it the name from
`models`.
for that platform has an `images` field. A gateway accepts any name a worker
lists for a model, in any letter case, and sends that worker its own name for
it.

```json
{ "platforms": [ {
Expand Down Expand Up @@ -277,6 +277,14 @@ try again, use a new key. Repeating works across a daemon restart, for
belong to your token: another token sending the same key starts a request of
its own.

Through a **gateway**, `allowDownload` does not change which worker is picked
and never starts a download: only runtimes already installed on a worker
count. It still makes the `POST` answer early, as described below. The
gateway sends the request only to a worker whose catalog can serve it: one
that lists `device` as a model or another name for one, in any letter case,
and pairs that model with `os` (or, without `os`, with at least one installed
runtime). The worker is sent its own name for the model.

With `allowDownload: true` the `201` is returned as soon as the request is
stored — resolving a downloadable runtime can take minutes, so progress and
any later failure surface on the request resource instead of on the `POST`
Expand Down Expand Up @@ -704,10 +712,9 @@ simulated (hence the thin Android catalog), trimmed to one worker:
```

`downloads.policy` is that worker's own effective policy, read when its
uplink connects and again on every periodic refresh. Routing needs it to know
whether a worker may install a missing runtime at all before sending it a
request that depends on one; it is never an override, since the worker clamps
`allowDownload` through the same policy regardless.
uplink connects and again on every periodic refresh. It is shown for
reference; routing does not read it, since no download is started through a
gateway.

`catalog` is what that worker can lease, each model with the runtimes it
pairs with. `host` is the worker's machine, the same block its own
Expand Down
38 changes: 24 additions & 14 deletions docs/internal/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -479,11 +479,8 @@ carries none. A gateway's own `status.get` reports the gateway's machine with
no tools.

From `config.get` the gateway keeps two fields. The effective
`downloads.policy`, because routing has to know whether a worker is even
*allowed* to install a missing runtime before it sends that worker a request
which depends on one. It is a routing input and never an override: the
worker still clamps `allowDownload` through its own policy, whatever the view
said. And `lease.maxTtlMs`, compared against the gateway's own to warn when a
`downloads.policy`, for display only: routing counts installed runtimes and
never reads it (ADR 0009 §3). And `lease.maxTtlMs`, compared against the gateway's own to warn when a
worker's cap is lower. Config is daemon input, read at start, so these change
only across a worker restart; re-reading them on the tick costs one call.
`config.get` is an admin operation, which the uplink session is.
Expand Down Expand Up @@ -578,17 +575,30 @@ that removed a worker, or the last that decided when none removed any — is
reported on `request.dispatched`. `gateway.routing` names a whole list; the
lists are code, and no config key lists or orders stages.

The v1 policy (`warm-then-free`) is three stages:

1. `eligible` (filter): drop workers that are disconnected, drained,
incompatible, or lacking the requested platform, model, or runtime — a
download counts as available only on a worker whose own `downloads.policy`
would allow it;
2. `warm-hit` (rank, settles): prefer a worker with an unleased `ready` device
matching the request — a **warm hit**, and a sub-second grant;
3. `free-capacity` (rank): otherwise the worker with the **most free running
The v1 policy (`warm-then-free`) is four stages:

1. `takes-requests` (filter): drop workers that are disconnected,
incompatible, drained, or whose capacity has not been read;
2. `can-serve` (filter): drop workers whose catalog cannot serve the request
(ADR 0009 §3, `routing/request-match.ts`). The model is the first entry of
the worker's `models` whose name or `modelAliases` entry equals the
requested name, ignoring letter case. A named runtime must be in that
model's `modelRuntimes`; with none named the list must be non-empty. Only
installed runtimes count, so a download never makes a worker able to
serve;
3. `warm-hit` (rank, settles): prefer a worker with an unleased `ready` device
matching the request, compared against the worker's own name for the
model — a **warm hit**, and a sub-second grant;
4. `free-capacity` (rank): otherwise the worker with the **most free running
capacity** for that platform.

The same matcher gives the name the gateway forwards: the worker is sent its
own name for the model, so it resolves exactly what routing matched, and
`allowDownload` is always forwarded as `false`. `lease.requested` keeps the
name the client sent. The `eligible` stage that predates this split stays in
the code only so the conformance tests can run the three stages that
reproduce the policy before ADR 0009.

There is no other placement rule in v1: no requester affinity, no label
selectors, no per-worker platform exclusions. Each of those is a future
routing policy behind `gateway.routing`, not a change to the request shape —
Expand Down
48 changes: 48 additions & 0 deletions e2e/gateway-fleet.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -451,6 +451,54 @@ describe("gateway fleet", () => {
});
});

it("grants --device 'iphone 16 pro' and --device pixel_7 through a gateway by the worker's own names", async () => {
const port = await freeLoopbackPort();
const gateway = await withDaemon({
configOverrides: { http: { host: "127.0.0.1", port }, mode: "gateway" },
driver: "none",
});
const minted = await gateway.cli(["token", "create", "--role", "worker"]);
const { secret } = minted.json as { secret: string };
// The fake driver resolves model names exactly, so a grant proves the gateway sent the
// worker its own name for the model.
await withDaemon({
configOverrides: {
gateway: { label: "worker", token: secret, url: `ws://127.0.0.1:${port}` },
},
driverScript: {
android: {
availableOsVersions: ["35"],
knownModels: ["Pixel 7"],
modelAliases: { "Pixel 7": ["pixel_7"] },
},
ios: { availableOsVersions: ["26.0"], knownModels: ["iPhone 16 Pro"] },
},
});
await waitForWorkers(
gateway,
(views) =>
views.length === 1 && views[0]?.connection === "connected" && views[0].catalog.length === 2,
"the worker connected with both catalogs",
);

const lease = async (platform: string, device: string) => {
const leased = await gateway.cli(
["lease", "--platform", platform, "--device", device, "--detach", "--no-wait"],
{ timeout: 30_000 },
);
expect(leased.code, leased.stderr).toBe(0);
const grant = leased.json as {
readonly device: { readonly spec: { readonly model: string } };
readonly lease: { readonly id: string };
};
expect((await gateway.cli(["release", grant.lease.id], { timeout: 30_000 })).code).toBe(0);
return grant.device.spec.model;
};

await expect(lease("ios", "iphone 16 pro")).resolves.toBe("iPhone 16 Pro");
await expect(lease("android", "pixel_7")).resolves.toBe("Pixel 7");
});

it("refuses an uplink whose token is not a worker join token", async () => {
const port = await freeLoopbackPort();
const gateway = await withDaemon({
Expand Down
9 changes: 4 additions & 5 deletions src/contract/schemas.ts
Original file line number Diff line number Diff line change
Expand Up @@ -725,11 +725,10 @@ export const workerViewSchema = z.object({
capacity: statusCapacitySchema.optional(),
/**
* The worker's effective `downloads.policy`, read once with `config.get` when the uplink
* connects. It is on the view because it is a *routing input*, not decoration: ADR 0005 §13
* says a request that would need a download is only eligible on a worker whose policy allows
* one, and #118's policy reads it from here rather than asking at dispatch time. Absent for
* a worker whose `config.get` the gateway could not read (an incompatible one, or a call
* that failed).
* connects. Display only: routing counts installed runtimes and never reads it, and the
* gateway forwards every request with `allowDownload: false` (ADR 0009 §3). Absent for a
* worker whose `config.get` the gateway could not read (an incompatible one, or a call that
* failed).
*/
downloads: z.object({ policy: z.enum(["never", "on-request", "always"]) }).optional(),
/**
Expand Down
3 changes: 2 additions & 1 deletion src/gateway/dispatcher.ts
Original file line number Diff line number Diff line change
Expand Up @@ -370,7 +370,8 @@ export class GatewayDispatcher {
...(input.mode === undefined ? {} : { mode: input.mode }),
},
{
allowDownload: input.allowDownload ?? false,
// `input.allowDownload` stays accepted and is not passed on: only installed runtimes
// count through a gateway (ADR 0009 §3).
noWait: input.noWait ?? false,
ownerId: input.owner ?? session.principal,
requesterId: input.requesterId ?? session.principal,
Expand Down
Loading
Loading