Deployment configuration for Forge central services on AWS ECS/Fargate.
Six services plus their dependencies, across as many stages as you need: sprue, hilt, swarf, piri-signing-service, delegator and plc, each on its own public hostname. All of them are backed by a shared RDS Postgres instance and an OpenBao that also serves as the root of trust for regional appliances.
This replaces the single-VM Docker Compose deployment in smelt, and carries over its secret and key generation code with one substantial change: keys are minted inside AWS by a Lambda rather than on an operator's laptop.
- How it fits together
- Architecture decisions
- What each service needs — including Sharp edges
- Repository layout
- Stages
- DNS
- What survives a destroy
- Runbook
- Development
- Planned work
- Related
ALB (*.latest.dev.fil-forge.com in dev)
│
┌───────┬────────────┬─────────┼─────────┬───────────────┬───────────┐
upload auth revoke plc delegator signer ssm
│ │ │ │ │ (OpenBao)
│ └── AppRole ─┼─────────┼─────────┼───────────────┼───────────┘
│ │ │ │ │
├──────── RDS Postgres (one database per service) ───────┤
│ │
S3 DynamoDB
Regional appliances reach OpenBao at ssm.<hostname_suffix> to unseal at boot,
and the Ingot on an appliance reaches plc at plc.<hostname_suffix>. In dev
these end in latest.dev.fil-forge.com. sprue, hilt and swarf reach plc over
private DNS instead, which keeps the call inside the VPC.
piri-signing-service is spelled signing-service in AWS resource names and
SSM parameter paths, but uses the stable public label signer. Similarly,
sprue, hilt and swarf keep their implementation names internally while serving
at upload, auth and revoke, as specified by the
Forge service identity RFC.
| Decision | Why |
|---|---|
| Secrets minted by a Lambda in the VPC | Private keys are generated where they will be used. Nothing is written to a laptop, and no key enters Terraform state: the function returns only DIDs, addresses and names. |
| SSM Parameter Store, per-service prefixes | Each task execution role reads only /forge-central/<stage>/<service>/*. A compromised sprue task cannot read hilt's AppRole secret or the delegator's transactor key. smelt's 1Password item is all-or-nothing by comparison. |
| OpenBao stores in Postgres, not on a volume | Fargate has no durable local disk. Reusing the RDS instance avoids an EFS filesystem and keeps a replaced task's data intact. |
| OpenBao seals with KMS | There is no unseal key to store, share or leak, and no sidecar polling to apply one. A restarted task comes back ready with no operator step. smelt runs 1-of-1 Shamir with the key in 1Password. |
| hilt authenticates with AppRole, not root | Its policy reaches only forge-central/hilt/data/tenant/*. smelt hands hilt the Vault root token and tracks that as debt. |
| Directory per stage over shared modules | Each stage's root says what differs and nothing else; the shared modules stop the stages drifting apart. |
| Two roots per stage | platform holds the VPC, RDS, OpenBao and ingress; apps holds the services. A routine image bump plans in seconds and never touches the database. |
| State in S3, one bucket per account | The state lives in the account it describes, locked by S3 conditional writes rather than a DynamoDB table. Nothing outside AWS has to be reachable for a deploy to work. |
| OpenTofu, and Terraform refused outright | A versions.tofu / versions.tf pair per root. Terraform stamps a version marker OpenTofu then reads as being from the future, so one stray terraform apply would lock OpenTofu out of that state. |
| Images pinned by digest | A git SHA names the last commit, not the code you just built, so it collides with itself against a dirty tree. A manifest digest is content-derived and cannot move underneath a deploy. |
| S3 and DynamoDB through task roles | No static access keys anywhere. This is what replaces MinIO's root user and password. |
How a regional appliance is admitted to a stage is decided in docs/decisions/2026-08-region-onboarding.md: the transit key it seals against, the wrapped token that delivers its unseal credential, and the registration writes central performs on its behalf.
| Service | Port | Health | Postgres | Other |
|---|---|---|---|---|
| sprue | 8080 | /health |
yes | 3 S3 buckets, plc |
| hilt | 8080 | /health |
yes | OpenBao, plc, calls sprue |
| swarf | 8080 | /health |
yes | plc, serves an SSE stream |
| delegator | 8080 | /healthcheck |
no | 2 DynamoDB tables, chain RPC |
| piri-signing-service | 7446 | /healthcheck |
no | chain RPC |
| plc | 3000 | /_health |
yes | nothing else |
Two more names appear in the parameter store: indexer and etracker. They are not deployed, and they get identities anyway, because the delegator validates two UCAN proofs at startup that must be signed by their keys, exactly as in smelt. Both are expected to become real services, so their keys are kept rather than discarded, which is also what lets the proofs be re-signed after a rotation.
These each cost an afternoon to rediscover.
- hilt and swarf bind
127.0.0.1by default. Without an explicitHILT_SERVER_HOST/SWARF_SERVER_HOSTthe health check can never pass. - sprue, hilt and swarf generate an ephemeral identity key when none is supplied, silently changing their DID on every restart. Supplying the key is mandatory, not an optimisation.
- hilt and swarf accept the identity key only as a file path, and the
delegator's UCAN proofs are file-only too (the inline variant panics). ECS
injects secrets as environment variables, so the
ecs-servicemodule wraps the entrypoint to write them out before exec'ing the process. - Health paths disagree:
/health,/healthcheckand/_healthall appear. - Migrations run in-process via goose for sprue, hilt, swarf and plc.
Concurrent starts race on the goose lock, so services run at
desired_count = 1until someone sets the relevant*_SKIP_MIGRATIONS. - No service exposes Prometheus metrics. Observability is JSON logs on
stdout, collected by CloudWatch into
/forge-central/<stage>/<service>and the provision Lambda into/aws/lambda/fc-<stage>-provision, both kept for 30 days, and forwarded from there to Grafana Cloud together with the AWS metrics for ECS, the ALB, RDS and the NAT gateway. See docs/observability.md for the labels and queries. - swarf's
/revocations/:sinceis a long-lived SSE stream, so the ALB idle timeout is raised well above its 60-second default. - did:web resolution goes over the public internet. hilt resolves sprue at
https://upload.<hostname_suffix>/.well-known/did.json, so a task in a private subnet reaches the public ALB back out through the NAT gateway. - Every plan warns that
failure_thresholdis deprecated. Expected, and the alternatives are worse: AWS fixed the Cloud Map custom health check wait at one 30-second interval and deprecated the parameter, but leaving it out makes the provider create the service with no custom health config at all, after which every plan schedules a replacement that lands in the same state. The comment interraform/modules/shared/ecs-service/routing.tfhas the full story and tracks the upstream issues that would end the warning.
# Go binary executed in AWS to provision DB & secrets
cmd/provision/ the Lambda: phase dispatch, seeding, OpenBao, funding, appliances
internal/keygen/ Ed25519 identities, secp256k1 wallets, UCAN proofs
internal/dbinit/ idempotent role and database creation
internal/vaultinit/ OpenBao init, mounts, hilt's AppRole
internal/ssmstore/ the never-overwrite parameter store
internal/fund/ the three FilecoinPay transactions
internal/onboard/ the appliance registration writes and their verification
build/ Lambda container image
scripts/fund-payer.sh invokes the fund phase, with a confirmation prompt
scripts/mint-appliance-token.sh issues an appliance's unseal credential, wrapped
scripts/onboard-appliance.sh registers an appliance and returns its S3 proof
scripts/retire-region.sh removes a region's Ingot identity from hilt and SSM
scripts/refresh-bump-prs.sh rebuilds every open bump branch on top of main
scripts/set-dev-pin.sh pins one dev service at one image digest
scripts/smoke-test.sh checks a deployed stage over public HTTPS
scripts/tail-logs.sh prints the tail of every log group a stage owns
scripts/wait-services-stable.sh waits for a cluster's ECS services to reach steady state
# Documentation beyond this file
docs/appliance-onboarding.md the runbook for admitting a regional appliance
docs/observability.md what a stage ships to Grafana and the queries that find it
docs/decisions/ why a thing is the way it is, one file per subject
# Infra configuration
terraform/
modules/ the wiring; see Stages below
envs/ one directory per root module
bootstrap/<account>/account/ state bucket, CI roles
bootstrap/<account>/<region>/ the image registry, telemetry egress to Grafana
dev/platform/ dev/apps/ applied on every push to main
staging/platform/ staging/apps/ applied on every push to main
prod/platform/ prod/apps/ platform applied on every push to main; apps not yet
# Deployment
.github/workflows/check-and-deploy.yml check, then plan on a PR or apply and smoke-test on main
A stage is a directory pair under terraform/envs/, backed by two state files in
the account's bucket. Everything is namespaced by the stage name, so stages coexist in one
AWS account without colliding:
fc-<stage>-*resources/forge-central/<stage>/*parameters- RFC service identities under the stage's
hostname_suffix(for example,upload.latest.dev.fil-forge.comin dev)
fc is short for forge-central, this repository's own deployment (as opposed to
deployments of regional nodes). It is kept short because a target group name is
capped at 32 characters, and <prefix>-<stage>-signing-service has to fit
inside it. Only AWS resource names abbreviate; paths and namespaces spell
forge-central out, since nothing there is close to a length limit.
Files:
terraform/envs/<stage>/platform/ VPC, RDS, S3, DynamoDB, ALB, OpenBao, provision Lambda
main.tf module "platform" plus what this stage overrides
terraform.tfvars committed, non-secret: DNS, chain, contracts
outputs.tf re-exported for the apps root
image.auto.tfvars committed Lambda digest; published in dev, copied on promotion
versions.tofu OpenTofu version, S3 backend, providers
versions.tf refuses Terraform; OpenTofu never reads it
terraform/envs/<stage>/apps/ the six ECS services
main.tf reads platform outputs via terraform_remote_state
terraform.tfvars committed: image digests
versions.tofu as above, with this root's own state key
versions.tf as above
Both roots stay short because modules/platform and modules/apps hold the
wiring. That is the point of the split: a stage's root says what differs, and
nothing else can drift between stages.
The module tree mirrors that split, so each root can name the directories it depends on:
terraform/modules/
platform/ everything the platform root builds
main.tf the wiring, calling the eight below
network/ kms/ database/ storage/ ingress/ provision/ openbao/
aurora/ prod's database, in place of database/
log-forwarding/ the role CloudWatch Logs ships a stage's groups to Grafana with
apps/ the six ECS services
shared/ used by more than one root
ecs-service/ apps, and openbao inside platform
constants/ every root, bootstrap included
ecr/ regional bootstrap only: the image registry
telemetry/ regional bootstrap only: the Firehoses and metric stream to Grafana
tfstate/ account bootstrap only: the state bucket
github-actions-iam/ account bootstrap only: the two CI roles
A module used by exactly one root lives under that root's composite module.
shared/ holds the two that genuinely cross the boundary. Adding a module
therefore never means editing a trigger pattern anywhere, because the workflow
applies both roots on every push rather than choosing between them by path.
Chain configuration lives in the platform root and the apps root
reads it from there, so a stage has one set of contract addresses rather than
two copies to keep in step. That mirrors smelt's shared smart-contracts.env.
The public Forge domains delegate the zones used by this deployment to Route53.
Dev stages share dev.fil-forge.com, staging uses staging.fil-forge.com, and
production uses fil-forge.com directly.
fil-forge.com DNS
├── NS dev ──► Route53 zone dev.fil-forge.com (non-prod account)
│ ├── upload.latest.dev.fil-forge.com
│ ├── ssm.latest.dev.fil-forge.com
│ └── upload.<STAGE>.dev.fil-forge.com
├── NS staging ─► Route53 zone staging.fil-forge.com (non-prod account)
│ ├── upload.staging.fil-forge.com
│ └── ssm.staging.fil-forge.com
├── NS upload ──► Route53 zone upload.fil-forge.com (production account)
├── NS ssm ─────► Route53 zone ssm.fil-forge.com (production account)
└── … one zone per public service name
Adding a personal stage beneath dev.fil-forge.com requires no change to the
DNS project. Shared domain roots such as staging are delegated once before a
stage uses them.
Production carries no stage label: upload.fil-forge.com. Ephemeral and
personal stages use <STAGE>.dev.fil-forge.com; this repository's dev stage is
the long-lived latest stage, so Sprue is upload.latest.dev.fil-forge.com.
The shared staging deployment is the RFC's separate
<service>.staging.fil-forge.com namespace; it is not a stage label beneath
dev.fil-forge.com.
Public labels are stable identities rather than implementation names: Sprue is
upload, Hilt is auth, Swarf is revoke, piri-signing-service is signer,
and Delegator and Indexer use delegator and indexer.
Two per-stage settings follow, and this is where they diverge:
zone_nameis the delegated Route53 zone records are written into. Dev stages sharedev.fil-forge.com; staging usesstaging.fil-forge.com.hostname_suffixis what that stage's hostnames end with, which for non-prod includes the stage label.ingot_hostname_suffixis the corresponding suffix in thefilonecontent.comnamespace. Ingot identities aredid:web:s3.<REGION>.<ingot_hostname_suffix>.
The delegation itself lives in
fil-one/infrastructure and is added
once per dev/staging domain root: an aws_route53_zone for the delegated name, plus a
Cloudflare NS record carrying that zone's four name servers.
Production has no domain root to delegate, because its service names sit
directly beneath fil-forge.com. Each public service name is a Route53 zone of
its own instead, created by terraform/envs/bootstrap/prod/account from the
constants module's public_hostname_labels, so the zones survive a rebuild of
the prod stage. fil-one/infrastructure carries one Cloudflare NS record per
zone, copied from that root's service_zone_name_servers output. The prod
platform root sets no zone_name, which tells the ingress module to write each
record into its hostname's own zone.
Those records are created with proxied = false, which matters: these hostnames
serve did:web documents and terminate their own TLS at the ALB, so Cloudflare
must not sit in front of them.
Certificates belong here, not in the fil-one/infrastructure project.
The ingress module issues *.<hostname_suffix>, writes the DNS validation
records into the delegated zone, and waits for validation. In production a
wildcard would validate through a record in the Cloudflare apex, so the
certificate names every public hostname and validates each one in its own zone. Two reasons it
cannot be one central certificate:
- An ALB needs its certificate in the ALB's own region. A
us-east-1certificate, which is what CloudFront requires, cannot be attached. - A wildcard covers exactly one label, so
*.dev.fil-forge.comdoes not matchupload.latest.dev.fil-forge.com. Each stage needs its own.
terraform destroy deletes no parameter this project generates. The
provision Lambda creates them, so Terraform has no record of them and never
removes them. An accidental destroy therefore cannot burn a funded wallet or
invalidate a DID that storage providers have already registered against.
They also stay readable, which takes deliberate arrangement. SecureStrings are encrypted under the account's AWS-managed SSM key rather than the stage's own customer-managed key. The stage's key is destroyed with the stage, and a key in PendingDeletion stops serving decryption at once, so tying the parameters to it would leave every secret unreadable the moment the stage came down and would fail the next apply that tried to rebuild it. The stage's key seals OpenBao and nothing else, and what it protects is meant to die with the stage: OpenBao's storage sits in the stage's database and goes at the same time.
So a destroyed and recreated stage silently comes back with its previous identities and wallets. An appliance's stored delegation is one of them: a rebuilt stage still holds the proof hilt signed for whatever Ingot DID it was addressed to, and onboarding returns that copy rather than issuing a new one. That is usually what you want, and it is occasionally a surprise, so check before assuming a rebuilt stage is fresh:
aws ssm get-parameters-by-path --path /forge-central/dev --recursive \
--query 'Parameters[].Name' --output textTo retire a stage, delete the parameters after the destroy, having first confirmed the wallets hold no funds:
# Check the balances first. This is not reversible.
aws ssm get-parameter --name /forge-central/dev/signing-service/payer-key.address
aws ssm get-parameter --name /forge-central/dev/delegator/transactor-key.address
aws ssm get-parameters-by-path --path /forge-central/dev --recursive \
--query 'Parameters[].Name' --output text \
| xargs -n 10 aws ssm delete-parameters --names- AWS CLI, with credentials for the target account.
- OpenTofu 1.12 or newer, for the bootstrap roots and
for reading a stage's outputs. Stage applies themselves run in GitHub Actions.
Terraform is not an alternative here and every root refuses it outright: it
stamps a version marker into state that OpenTofu reads as being from the future,
so one
terraform applywould lock OpenTofu out of that state. - Docker with buildx, for
make publish. - Go and make, for
make checkandmake test. - ShellCheck, for the shell half of
make check. CI pins 0.11.0, so that build is the one that decides a merge; an older local one can pass a script CI rejects. - Foundry's
cast, only to read chain balances by hand. Nothing in the deploy path needs it.
| Part | How it is deployed |
|---|---|
bootstrap roots |
tofu apply run locally, always |
| provision image | make publish run locally, pushed to ECR by hand |
dev platform, apps |
GitHub Actions, on every push to main, with no approval step |
staging platform, apps |
GitHub Actions, on every push to main, with no approval step |
The dev and staging stages deploy themselves after their initial bootstrap.
.github/workflows/check-and-deploy.yml runs make check on every pull request
and every push to main; a pull request then plans all four roots, and a push
applies and smoke-tests both stages. The OpenTofu version is pinned in the
workflow rather than taken from an operator's machine.
apps reads platform's state through terraform_remote_state, so ordering
matters: apply-dev-apps waits on apply-dev-platform, and the corresponding
staging jobs have the same edge. An apps job therefore never plans against
outputs an in-flight platform apply is about to change. Every root is applied
on every push, even one that touched only one of them. An empty plan costs about
a minute, and it means there is no path-filter list to forget to update when a
module moves.
In a pull request all four plans run at once, and each apps plan is computed against its last applied platform state rather than against this pull request's platform plan. A change to a platform output that apps consumes therefore shows its real apps plan only after platform applies.
An apply reports success as soon as AWS accepted the change, which for an ECS
service means a task definition was registered rather than that a task is
serving traffic on it. Each apps apply therefore ends by running
scripts/wait-services-stable.sh, which waits for every service in the cluster
to reach steady state, and a task that never becomes healthy fails the job after
twenty minutes. Without that wait a smoke test can pass against the revision the
push replaced, because a rolling update keeps the old task answering. The script
names the services it is still waiting on as it polls, and every two minutes
prints their task counts, their deployments and their recent ECS events, so a
long wait says whether a rollout is slow or has stopped moving.
smoke-dev and smoke-staging then call make smoke for their stage and retry
for four minutes. Steady state covers the task; a newly created Route53 record
or listener rule in front of it can take a moment longer. The smoke checks need
no credentials because every request goes over public HTTPS. See
Smoke-testing a stage.
When either the wait or the smoke test fails, scripts/tail-logs.sh prints the
tail of every log group the stage owns into the run, so the diagnosis is where
the failure is. The groups are discovered from CloudWatch, so a service added to
either root is covered. It runs on the apply role, because reading log events
needs logs:FilterLogEvents and the plan role deliberately has none of it.
A failed run on main posts to #filone-alerts in Slack with the stage that
failed, the failed jobs, the commit subject, its author and a link to the run.
The stage comes from the job name: apply-<stage>-<root> and smoke-<stage>
name a stage, apply-grafana reports as grafana, and check as itself.
When the commit is an image bump, one more line names the service commit that
produced the image and who wrote it, looked up from the source commit link Bump
deployed image puts in the commit body. A lookup that fails drops the line and
still sends the alert. Any failed job triggers it, from make check through the
smoke test. Pull request failures
are not announced, because the author already sees the red check on the pull
request. The job reads one repository secret, SLACK_BOT_TOKEN, holding the bot
token of a Slack app with the chat:write scope; without the secret the
notification step fails and nothing else about the run changes.
AWS credentials are never stored. Each job assumes an IAM role in the target
account through GitHub's OIDC federation, and the credentials expire with the job.
There are two roles, and the split matters: GitHub runs the workflow file from a
pull request's own head, so the role a plan job uses can describe infrastructure
and read nothing, and the role that can change anything is reachable only from
refs/heads/main. See terraform/modules/github-actions-iam.
The workflow applies the prod platform root on every merge to main. The prod
apps root has no CI job yet.
See Planned work for the manual steps that remain.
Staging's capacity, durability, identity and promotion choices are recorded in the staging environment decision.
The bootstrap roots are always applied locally. They run rarely, they create the things everything else depends on, and so there is nothing for a pipeline to trigger on and no earlier apply to have created their state. They come in two kinds, and the split is what keeps a second region cheap:
bootstrap/<account>/account/holds the state bucket every other root in the account keeps its state in, and the two CI roles GitHub Actions assumes to plan and apply the stages. One per account: a bucket name is global and IAM is not regional, so a second region must not create these again.bootstrap/<account>/<region>/holds the ECR repository for the provision Lambda image and the telemetry egress to Grafana Cloud. In prod it also holds the KMS keys of the Aurora cluster and of OpenBao's seal, which must outlive the platform root. One per account and region, described in Setting up an AWS region.
Both accounts this project uses already have an account root, and both have been applied.
Copy a bootstrap/<account>/ directory, both the account/ root and the
regional one beside it. In the copies, point each provider at the account id it
belongs to, set the bucket name in account/main.tf and in both versions.tofu
backend blocks, and add that id to terraform/modules/shared/constants if it is
not there yet.
Every root reads its account id from that module, so an apply run with credentials for the wrong account fails at plan time rather than building a second working copy of the stage somewhere unexpected.
One thing has to exist before the account root can be applied, and nothing here
creates it: the GitHub OIDC provider,
https://token.actions.githubusercontent.com. It is one per account and shared
with every other repository that deploys into that account, so
modules/github-actions-iam reads it as a data source rather than owning it.
Creating it here would fail for the second repository to try, and a destroy
would lock the first one out of its own CI.
Both accounts this project uses already have it, so this matters only for an account nobody has deployed to from GitHub Actions before. Check:
aws iam list-open-id-connect-providers \
--query "OpenIDConnectProviderList[?contains(Arn, 'token.actions.githubusercontent.com')]"If that comes back empty, create it before applying the account root, or the
apply that makes the CI roles fails on the lookup with NoSuchEntity:
aws iam create-open-id-connect-provider \
--url https://token.actions.githubusercontent.com \
--client-id-list sts.amazonaws.comsts.amazonaws.com is the audience aws-actions/configure-aws-credentials
requests when the workflow does not override it, and it is what both trust
policies require in their aud condition. Omit it from the client id list and
every sts:AssumeRoleWithWebIdentity call is rejected. No --thumbprint-list:
AWS no longer validates one for this provider.
The account root is the awkward one: its own backend points at the bucket it
creates, so the first apply in a fresh account cannot use that backend. Run it
against a local backend once, then move its state into the bucket it just made.
.gitignore already ignores *_override.tf, so the override cannot be committed
by accident:
cd terraform/envs/bootstrap/<account>/account
printf 'terraform {\n backend "local" {}\n}\n' > backend_override.tf
tofu init
tofu apply -target=module.tfstate # the bucket, and nothing else yet
rm backend_override.tf
tofu init -migrate-state # local state moves into the bucket
rm -f terraform.tfstate terraform.tfstate.backup
tofu apply # the CI rolesNote the two role ARNs it prints. .github/workflows/check-and-deploy.yml names
them literally, so if they differ from what is there, the workflow needs
updating.
Every root after this one is ordinary, because its backend block points at a bucket that now exists. The regional root beside it comes next.
The regional root, bootstrap/<account>/<region>/, holds two things:
forge-central/provision, the ECR repository for the provision Lambda image. Lambda pulls an image only from ECR in the same region as the function. Stages sharing an account and region share the repository and pin different digests. Its repository policy is what lets Lambda pull the image; without it, creating a stage's provision Lambda fails withAccessDeniedException.- The telemetry egress to Grafana Cloud: one log Firehose per stage, from
the stage list in
terraform/modules/shared/constants, and one CloudWatch metric stream, which covers one account in one region.
The first region of an account already has this directory next to the account root. For a further region, copy it:
cp -r terraform/envs/bootstrap/nonprod/us-east-2 terraform/envs/bootstrap/nonprod/us-west-2Change two things in the copy: the region in the provider block, and the key
in the backend "s3" block in versions.tofu (bootstrap/us-west-2.tfstate).
Leave the backend's region at us-east-2. It names the region the state
bucket is in, and the bucket is one per account, created by the account root.
Pointing it at us-west-2 makes tofu init fail against a bucket that is
sitting right there. Nothing else needs changing and nothing needs deleting: the
account-scoped resources are not in this directory to begin with, and the
telemetry module's backup bucket and IAM roles, which share the account's
namespace, take the region from the provider and so get their own names.
Applying the regional root needs three values from the Grafana Cloud stack,
passed as environment variables and committed nowhere. All three are in the
Forge Central item of the Fil One vault in 1Password, under the GRAFANA
section. With the 1Password CLI
signed in:
export TF_VAR_grafana_logs_user="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_LOGS_USER')"
export TF_VAR_grafana_metrics_user="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_METRICS_USER')"
export TF_VAR_grafana_push_token="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_CLOUD_PUSH_TOKEN')"op item get --vault "Fil One" "Forge Central" lists the section's fields
without revealing the token. The two *_USER values are the Loki and Prometheus
instance ids of the stack. The item's two *_URL fields are the Firehose
delivery endpoints, which are what the module defaults to, so nothing needs
setting for them. They are not the plain Loki and Prometheus push URLs Alloy
uses on the appliances: Firehose has its own delivery format and Grafana
receives it on aws-logs-* and aws-metric-streams-* hosts.
The token is a Grafana Cloud access policy token with the logs:write and
metrics:write scopes, created under Security → Access Policies in the
Grafana Cloud portal, the same kind infra-nodes' runbook describes for the
appliances. Rotating it is a new token in the 1Password item and a tofu apply
of this root. The token ends up in this root's state, which is why the
Firehoses live in this root rather than in a stage root; see
the telemetry decision.
The bucket already exists, so there is no bootstrap dance here:
cd terraform/envs/bootstrap/nonprod/us-east-2
tofu init
tofu apply # the image registry and the telemetry egressEvery image this project publishes to ECR lives under the forge-central/
prefix, one repository per image. Per-image repositories are what make per-image
push permissions, lifecycle policies, and tag immutability possible.
Then fill the repository. The provision image is built and pushed by hand from
a developer machine; nothing builds it automatically. make publish needs
Docker with buildx and AWS credentials for the target account, and it creates a
docker-container builder on first use, because Docker Desktop's default
builder cannot push by digest and cannot cross-build for arm64.
make publish STAGE=dev # AWS_REGION defaults to us-east-2
make publish STAGE=<stage> AWS_REGION=us-west-2 # a further regionIt pushes by digest and writes no tag, so the digest a stage pins is the only reference to the image. That is why the repository rejects tags and carries no expiry rule: an untagged image is indistinguishable from one a stage is running, and Lambda does not survive having its image deleted. Prune by hand when the image count starts to bother you.
A stage needs nothing copied from the bootstrap output. It builds the image URL
from its own account and region, which is the only registry its Lambda can pull
from anyway; the account ids and the repository name live in
terraform/modules/shared/constants. The digest is derived from the image
rather than from where it is stored, so a stage in a new region can pin the same
digest an existing stage already runs.
cp -r terraform/envs/dev terraform/envs/bajtosThen, in the copy:
- Set the
keyin bothversions.tofufiles tobajtos/platform.tfstateandbajtos/apps.tfstate, and thekeyin the apps root'sterraform_remote_stateblock to match the platform one. The bucket is already right: it is per account, and personal stages share the non-prod account. - Change
stage = "dev"to"bajtos"inplatform/main.tf, and theStagedefault tag in both roots. - In
platform/terraform.tfvars, sethostname_suffixto<STAGE>.dev.fil-forge.comandingot_hostname_suffixto<STAGE>.dev.filonecontent.com. Leavezone_nameasdev.fil-forge.com: the zone is already delegated and shared by every dev stage, so the DNS project needs no change. Point thechainblock at the network this stage transacts against. - Add the stage to
.github/workflows/check-and-deploy.yml: two more entries in theplanmatrix, namedbajtos-platformandbajtos-apps, two more apply jobs copied from dev's, withapply-bajtos-appsneedingapply-bajtos-platform, a smoke job for the new stage, and a diagnose job copied from dev's, which names the stage whose logs it tails. Add both apply jobs and the smoke job tonotify-failure'sneeds; a failure in a job it does not name announces nothing. - Add
"bajtos"tononprod_stagesinterraform/modules/shared/constants/outputs.tfand apply both bootstrap roots for the account. The account root grants the CI roles state access by stage prefix and the regional root creates the stage's log Firehose, so a stage missing from the list either cannot read its own state or fails its first platform apply creating its subscription filters. Why there is one list is in the telemetry decision. - Apply the new stage's platform root locally once. The apps root reads the platform's remote state, so its first CI plan cannot run until that state exists. Do not apply the apps root yet; the workflow will do that after the change merges.
- Update the branch protection rule on
main. The new stage adds two required checks,plan-bajtos-platformandplan-bajtos-apps, and a rule that does not name them will merge a pull request whose stage plan failed. - Merge. The stage's platform job reconciles the VPC, RDS, OpenBao and secrets, its apps job applies the six services, and its smoke job tests the public endpoints. Nothing in the apps root needs starting by hand.
The first platform apply is slow: it waits for the OpenBao task's cold start
before it can initialise it, inside a synchronous Lambda call that Lambda caps at
15 minutes. If it times out there, re-run the job. The seed phase regenerates
nothing that already exists, which is what protects funded wallets.
Prod differs from dev inside main.tf rather than by being a different
shape: an Aurora cluster with a writer and a reader in its own subnets,
deletion protection on, KMS keys for the cluster and for OpenBao's seal from
the regional bootstrap, a larger OpenBao connection budget, a zone per public
hostname, and a provision digest pinned in terraform.tfvars. It lives in its
own account and deploys on every merge, like staging. Its choices are recorded
in the first prod stack decision,
and staging's in the staging environment
decision.
Stage names are not limited to dev and prod. Copy envs/dev to envs/<you>, give
it its own state key in the same bucket, and apply it from your machine: no
commit, no merge, no workflow run to wait for, which is the fastest loop for
iterating on the provision Lambda. Leave it out of check-and-deploy.yml; that is what
makes it yours.
Set enable_log_forwarding = false in the copied platform/main.tf. The
regional bootstrap root creates a log Firehose only for the stages in
nonprod_stages, and a platform apply that forwards to a Firehose that does
not exist fails creating its first subscription filter. The stage's logs stay
in CloudWatch, where scripts/tail-logs.sh reads them. A sandbox that needs
its logs in Grafana follows step 5 of adding a stage
instead, at the cost of a commit.
What it costs is everything the dev stage gets from the workflow: a plan on every
pull request, applies that cannot disagree with main, and an OpenTofu and
provider version that is the same for everyone. Use it to iterate, not to host
anything anyone depends on.
A stage seals the appliances named in its appliance_regions, and
make mint-appliance-token issues a node's unseal credential. The ordering across
both repositories, how the credential is delivered to a node operator, and what
retiring a region destroys are in
docs/appliance-onboarding.md.
The seed phase mints two secp256k1 wallets and reports their addresses. Both start empty, and nothing works until they hold funds:
tofu -chdir=terraform/envs/dev/platform output wallet_addressesGas, for both wallets. The delegator's transactor signs provider approvals and the payer signs PDP operations, so each needs tFIL on Calibration or FIL on mainnet. Faucet: https://faucet.calibnet.chainsafe-fil.io/
USDFC, for the payer only. Faucet: https://forest-explorer.chainsafe.dev/faucet/calibnet_usdfc — capped at 10 USDFC per day, which is why the amounts below stay small.
Depositing into FilecoinPay. USDFC sitting in the payer's wallet is not
enough. Creating a proof set locks up around 0.9 USDFC, and lockup can only draw
on funds deposited into the FilecoinPay contract, so a freshly faucet-funded
wallet still fails with InsufficientLockupFunds(..., Available=0).
make fund-payer STAGE=devThat runs three transactions, the same ones as smelt's
scripts/staging-fund-payer.sh:
USDFC.approve(FilecoinPay, amount)— let FilecoinPay pull the tokensFilecoinPay.deposit(USDFC, payer, amount)— credit the payer's accountFilecoinPay.setOperatorApproval(USDFC, FWSS, true, rate, lockup, period)— let warm storage lock it up
The signing happens inside the provision Lambda, in AWS.
The script invokes the Lambda twice. The first call reads the chain and prints what it would do, signing nothing. You then confirm, and the second call broadcasts.
Amounts default to smelt's, which stay under the faucet's daily cap. Override them per run:
make fund-payer STAGE=dev DEPOSIT=5
make fund-payer FUND_ARGS="--rate-allowance 0.2 --lockup-allowance 5"
scripts/fund-payer.sh --stage dev --deposit 3 --force-depositTerraform never invokes this phase. An apply must not move money, so funding is always an explicit operator action.
Two preconditions are checked before anything is signed: the RPC must report the chain id the stage expects, and the payer wallet must already hold at least the deposit amount. Neither is recoverable by this tooling, so both fail loudly.
To read the balances without invoking anything:
PAYER=$(aws ssm get-parameter --name /forge-central/dev/signing-service/payer-key.address \
--query Parameter.Value --output text)
cast call "$USDFC_TOKEN_ADDRESS" "balanceOf(address)(uint256)" "$PAYER" \
--rpc-url https://api.calibration.node.glif.io/rpc/v1
cast call "$FILECOIN_PAY_ADDRESS" \
"accounts(address,address)(uint256,uint256,uint256,uint256)" \
"$USDFC_TOKEN_ADDRESS" "$PAYER" \
--rpc-url https://api.calibration.node.glif.io/rpc/v1make publish STAGE=devThat writes the new digest into the stage's image.auto.tfvars, so there is no
line to edit by hand. Commit that file and merge it. The stage is planned by a
workflow, which sees only what is in version control, so a digest left on your
machine is applied nowhere.
Dev and staging share an ECR repository. Promote the Lambda to staging by
copying dev's digest into
terraform/envs/staging/platform/image.auto.tfvars. The image is already in
ECR, so the promotion needs no build or push.
Prod is a separate account with its own ECR repository, so a dev digest means nothing there. Promote the Lambda to prod by publishing into the prod account:
make publish STAGE=prodFor prod the command writes nothing. It prints the provision_image_digest
line to paste into terraform/envs/prod/platform/terraform.tfvars, where prod
pins its digest. Commit that file and merge it.
Change its digest in the stage's image_digests and merge. Every stage pins
digests, dev included: dev is applied on every push to main, and a rolling tag
would make what dev runs depend on when a task last restarted rather than on what
was merged.
For dev, the services do this themselves. A publish workflow dispatches a
bump-deployed-image event carrying the digest it just pushed, and
bump-deployed-image.yml opens a
pull request that changes the one line, with auto-merge enabled so the deploy
lands as soon as the required checks pass. Those pull requests come from the
fil-forge-bot GitHub App on the branch bot/bump-<service>-image-dev, one
branch per service, so a second publish updates the open request instead of
stacking a stale one beside it.
Two things keep those pull requests mergeable while several are open at once.
The pins are spaced a blank line apart, because git conflicts on changes to
adjacent lines and each bump rewrites one line. And
refresh-bump-prs.yml rebuilds every
open bump branch on top of main, because the ruleset will not merge a branch
that is behind. It runs when main moves, when a bump branch is pushed, and
when a bump pull request is opened: a bump commit is built from the main its
run checked out, which can be behind by the time the push lands, and on a
service's first bump the branch is pushed before the pull request exists to be
found. A rebuild keeps the original commit message, so the link to the pull
request that published the image survives. A branch whose digest main already
pins is closed instead, and one whose service someone else moved meanwhile is
left alone: which digest dev should run is then a question rather than an edit,
and the pull request shows the conflict it has.
The same workflow bumps any of the six services on demand:
gh workflow run bump-deployed-image.yml -R fil-forge/infra-central \
-f service=sprue -f digest="$(crane digest ghcr.io/fil-forge/sprue:main)"Locally, scripts/set-dev-pin.sh makes the same edit.
It is what both workflows run, so a pin written by hand comes out identical to
one written by the bot:
scripts/set-dev-pin.sh sprue "$(crane digest ghcr.io/fil-forge/sprue:main)"A service is wired up with a dispatch step in its publish workflow plus the
fil-forge-bot credentials in that repository. The receiver only accepts
dispatches made as that app, and the source_repo in the payload has to be the
repository the service is published from, so the required client_payload is
service, digest and source_repo; commit, pr_url and run_url are
provenance links the commit message uses when present.
Staging and prod stay manual: a promotion copies dev's reviewed digest in a deliberate pull request.
The most important check after any apply. Read created_parameters from the
stage's platform apply job, such as apply-dev-platform or
apply-staging-platform, or from a shell:
tofu -chdir=terraform/envs/dev/platform output created_parametersEmpty means every key already existed and was reused. A non-empty list after the first apply of a stage means something was minted; find out what before assuming a wallet is intact.
make smoke STAGE=devEvery public service is checked over public HTTPS, needing no AWS credentials. A 200 from the health path covers the whole ingress route in one request: the Route53 record, the wildcard certificate, the listener rule, the target group and a task passing its container health check.
The second check is the one health cannot make. sprue, hilt and swarf mint an
ephemeral identity when no key is supplied and report themselves healthy either
way, so /.well-known/did.json is read and its id compared against
did:web:<hostname>. A mismatch means the service is running an identity
nothing has registered against.
OpenBao is checked too, at ssm.<suffix> rather than at its own name. The
request omits the uninitcode=200 its ALB health check passes: ECS has to keep a
fresh task alive long enough for the provision Lambda to initialise it, but a
stage that has finished deploying and is still uninitialised is a failure.
The same command runs in CI for dev and staging after every push to main. See
How each part is deployed.
The script reads hostname_suffix from the stage's
platform/terraform.tfvars, so it needs no Terraform state and no TFE token.
Services are probed concurrently: a task that accepts the connection and never
replies waits out the whole timeout, and several of those in sequence is a
minute of nothing.
plc and piri-signing-service get the health check alone, and the output says so
rather than passing over it. plc publishes no identity. piri-signing-service
takes a did:web and serves no document at it, so nothing resolves it today: it
is the only service no other service addresses by DID.
Run by hand it also says nothing about which revision answered. No service reports its build, so a stage mid-rollout can pass on the old task. In CI the stage's apps apply job closes that gap by waiting for steady state first.
Delete the parameter:
aws ssm delete-parameter --name /forge-central/dev/swarf/identityThen force the seed phase to run again. A plain run will not do it: the phase is
an aws_lambda_invocation, which re-invokes only when its input changes, and
deleting a parameter changes nothing Terraform can see. See Forcing a provision
phase to re-run.
The new DID appears in service_dids, and the rotation also refreshes
/forge-central/<stage>/<service>/identity.did, which holds the same value for
anyone reading it without decryption rights. Anything that had registered the
old DID has to be told about the new one, which is why this is a deliberate act
rather than something an apply does on its own.
Rotating an identity that signs a proof — sprue, indexer or etracker — re-issues that proof automatically in the same apply, because the old delegation would no longer verify against a key that does not exist.
A UCAN delegation is public but not reproducible: ucantone mints a random 16-byte nonce per delegation, so signing the same request twice produces different bytes and a different CID. Rewriting on every apply would churn the parameter and invalidate anything holding the previous delegation.
So proofs are written once and then left alone, and re-issued only when their issuer key was freshly minted. smelt tracks the same dependency, skipping a committed proof unless one of the keys behind it was regenerated that run.
The practical consequence: you cannot verify a proof by regenerating it and
diffing. Only the framing is reproducible, and that is what the tests pin: a
textual container is stored as ucantool writes it, trailing newline included,
while a bare DAG-CBOR delegation is stored base64-encoded.
Every proof reaches its consumer as an environment variable that the task's
entrypoint writes to a file, and an environment variable cannot carry a NUL
byte. A bare DAG-CBOR delegation is binary and contains them, so the container
never starts and runc reports only the variable's name. The delegator's two
proofs are therefore stored base64-encoded and decoded on the way to the file,
which is what secret_files_base64 on modules/shared/ecs-service is for.
hilt's proof is a base64+gzip container, already text, and travels as it is.
aws ssm delete-parameter --name /forge-central/dev/hilt/vault-secret-idThen force the vault phase to run again. There is no way to do that from a remote run yet — see Forcing a provision phase to re-run.
The vault phase also self-heals: if a stored secret_id no longer
authenticates, because OpenBao's storage was rebuilt underneath it, the next run
replaces it rather than leaving hilt unable to start.
make check # gofmt, go vet, go test, tofu fmt, shellcheck
make test # go test alone, for the inner loopBoth run offline against no deployed stage, which is why make smoke is
separate: it needs a stage to be up. See Smoke-testing a
stage.
Dependabot opens the pull requests, .github/dependabot.yml says which and how
often, and
auto-merge-dependabot.yml
merges the ones that are minor or patch bumps. The merge is squashed and armed
through fil-forge-bot, so main moves only after make check and all required
plans have passed. The push that lands applies both shared non-prod stages the
same way any other merge to main does.
A major bump stays open for someone to read. So does a group whose highest change is a major, and so does any Dependabot branch that carries a commit Dependabot did not write.
Dependabot rebases its pull requests in place, and the workflow runs again on each new head, so a branch that stops qualifying also loses the auto-merge it was given earlier. Each auto-merge is bound to the head the workflow inspected, which is what stops a head that arrives mid-run from merging on the previous head's decision.
See the following Linear tickets:
- FIL-1147 Stand up the first, disposable Forge Central prod stack
- FIL-1394 Hold prod applies for a human approval
- FIL-1396 Reset production after the test run
- FIL-1156 Narrow the apply role's IAM policy
- FIL-1090 Tooling for removing an appliance node from the network
- FIL-1091 Hilt: API to remove Ingot node
- FIL-1154 Hilt authenticates to OpenBao with AWS IAM auth
- FIL-1148 Publish the provision Lambda image from CI
- FIL-1150 Enable the OpenBao audit log on Forge Central
- FIL-1157 Decide the OpenBao availability target for Forge Central
- FIL-1149 Write and rehearse the RDS restore procedure
- FIL-1151 Grafana alarms for Forge Central services
- FIL-1152 Expose metrics from swarf, delegator and piri-signing-service
- FIL-1153 Write down the running cost of a Forge Central stage
- FIL-1160 Replace static Postgres passwords with RDS IAM authentication
- FIL-1155 Own the RDS parameter group and pin
rds.force_ssl - FIL-1161 Verify the RDS server certificate in every Forge service
- FIL-1162 Zero-downtime upgrades of Forge Central services
- FIL-1158 Automate promotion of Forge Central from dev to staging
aws_lambda_invocation re-invokes only when its input changes, which is what
seed_trigger and vault_trigger in modules/platform exist for. Rotating an
identity or hilt's OpenBao credential needs one of them bumped.
Neither is reachable today. The stage roots do not expose them, and the workflow
passes no -var flags, so the only way to bump one is to edit
envs/<stage>/platform/main.tf and merge.
That is arguably the right answer rather than a limitation. A committed, reviewed
diff is what re-running the thing that mints wallets should cost, and it stays
visible afterwards. The alternative — a workflow input anyone with write access
could set — puts that one text field away. A workflow_dispatch input would be the
middle ground if the merge ever proves too slow.
- smelt — the single-VM Docker Compose deployment this replaces, and the source of the key generation code.
- smelt#11 — its appliance registration scripts, which specified the writes the onboard phase performs, including the region mismatch hilt's own error cannot distinguish.
- fil-forge/infra-nodes — the appliance side: the node this repository seals, registers and can revoke.
- fil-one/RFC#21 — regional security and key management, which makes this OpenBao the root of trust for appliances.