Skip to content

About

Deployment configuration for Forge central services

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

180 Commits

Folders and files

Repository files navigation

infra-central

Deployment configuration for Forge central services on AWS ECS/Fargate.

Six services plus their dependencies, across as many stages as you need: sprue, hilt, swarf, piri-signing-service, delegator and plc, each on its own public hostname. All of them are backed by a shared RDS Postgres instance and an OpenBao that also serves as the root of trust for regional appliances.

This replaces the single-VM Docker Compose deployment in smelt, and carries over its secret and key generation code with one substantial change: keys are minted inside AWS by a Lambda rather than on an operator's laptop.

Contents

How it fits together

                              ALB  (*.latest.dev.fil-forge.com in dev)
                                  │
   ┌───────┬────────────┬─────────┼─────────┬───────────────┬───────────┐
 upload   auth       revoke      plc    delegator       signer        ssm
   │       │            │         │         │                       (OpenBao)
   │       └── AppRole ─┼─────────┼─────────┼───────────────┼───────────┘
   │       │            │         │         │
   ├──────── RDS Postgres (one database per service) ───────┤
   │                                        │
   S3                                   DynamoDB

Regional appliances reach OpenBao at ssm.<hostname_suffix> to unseal at boot, and the Ingot on an appliance reaches plc at plc.<hostname_suffix>. In dev these end in latest.dev.fil-forge.com. sprue, hilt and swarf reach plc over private DNS instead, which keeps the call inside the VPC.

piri-signing-service is spelled signing-service in AWS resource names and SSM parameter paths, but uses the stable public label signer. Similarly, sprue, hilt and swarf keep their implementation names internally while serving at upload, auth and revoke, as specified by the Forge service identity RFC.

Architecture decisions

Decision Why
Secrets minted by a Lambda in the VPC Private keys are generated where they will be used. Nothing is written to a laptop, and no key enters Terraform state: the function returns only DIDs, addresses and names.
SSM Parameter Store, per-service prefixes Each task execution role reads only /forge-central/<stage>/<service>/*. A compromised sprue task cannot read hilt's AppRole secret or the delegator's transactor key. smelt's 1Password item is all-or-nothing by comparison.
OpenBao stores in Postgres, not on a volume Fargate has no durable local disk. Reusing the RDS instance avoids an EFS filesystem and keeps a replaced task's data intact.
OpenBao seals with KMS There is no unseal key to store, share or leak, and no sidecar polling to apply one. A restarted task comes back ready with no operator step. smelt runs 1-of-1 Shamir with the key in 1Password.
hilt authenticates with AppRole, not root Its policy reaches only forge-central/hilt/data/tenant/*. smelt hands hilt the Vault root token and tracks that as debt.
Directory per stage over shared modules Each stage's root says what differs and nothing else; the shared modules stop the stages drifting apart.
Two roots per stage platform holds the VPC, RDS, OpenBao and ingress; apps holds the services. A routine image bump plans in seconds and never touches the database.
State in S3, one bucket per account The state lives in the account it describes, locked by S3 conditional writes rather than a DynamoDB table. Nothing outside AWS has to be reachable for a deploy to work.
OpenTofu, and Terraform refused outright A versions.tofu / versions.tf pair per root. Terraform stamps a version marker OpenTofu then reads as being from the future, so one stray terraform apply would lock OpenTofu out of that state.
Images pinned by digest A git SHA names the last commit, not the code you just built, so it collides with itself against a dirty tree. A manifest digest is content-derived and cannot move underneath a deploy.
S3 and DynamoDB through task roles No static access keys anywhere. This is what replaces MinIO's root user and password.

How a regional appliance is admitted to a stage is decided in docs/decisions/2026-08-region-onboarding.md: the transit key it seals against, the wrapped token that delivers its unseal credential, and the registration writes central performs on its behalf.

What each service needs

Service Port Health Postgres Other
sprue 8080 /health yes 3 S3 buckets, plc
hilt 8080 /health yes OpenBao, plc, calls sprue
swarf 8080 /health yes plc, serves an SSE stream
delegator 8080 /healthcheck no 2 DynamoDB tables, chain RPC
piri-signing-service 7446 /healthcheck no chain RPC
plc 3000 /_health yes nothing else

Two more names appear in the parameter store: indexer and etracker. They are not deployed, and they get identities anyway, because the delegator validates two UCAN proofs at startup that must be signed by their keys, exactly as in smelt. Both are expected to become real services, so their keys are kept rather than discarded, which is also what lets the proofs be re-signed after a rotation.

Sharp edges

These each cost an afternoon to rediscover.

  • hilt and swarf bind 127.0.0.1 by default. Without an explicit HILT_SERVER_HOST / SWARF_SERVER_HOST the health check can never pass.
  • sprue, hilt and swarf generate an ephemeral identity key when none is supplied, silently changing their DID on every restart. Supplying the key is mandatory, not an optimisation.
  • hilt and swarf accept the identity key only as a file path, and the delegator's UCAN proofs are file-only too (the inline variant panics). ECS injects secrets as environment variables, so the ecs-service module wraps the entrypoint to write them out before exec'ing the process.
  • Health paths disagree: /health, /healthcheck and /_health all appear.
  • Migrations run in-process via goose for sprue, hilt, swarf and plc. Concurrent starts race on the goose lock, so services run at desired_count = 1 until someone sets the relevant *_SKIP_MIGRATIONS.
  • No service exposes Prometheus metrics. Observability is JSON logs on stdout, collected by CloudWatch into /forge-central/<stage>/<service> and the provision Lambda into /aws/lambda/fc-<stage>-provision, both kept for 30 days, and forwarded from there to Grafana Cloud together with the AWS metrics for ECS, the ALB, RDS and the NAT gateway. See docs/observability.md for the labels and queries.
  • swarf's /revocations/:since is a long-lived SSE stream, so the ALB idle timeout is raised well above its 60-second default.
  • did:web resolution goes over the public internet. hilt resolves sprue at https://upload.<hostname_suffix>/.well-known/did.json, so a task in a private subnet reaches the public ALB back out through the NAT gateway.
  • Every plan warns that failure_threshold is deprecated. Expected, and the alternatives are worse: AWS fixed the Cloud Map custom health check wait at one 30-second interval and deprecated the parameter, but leaving it out makes the provider create the service with no custom health config at all, after which every plan schedules a replacement that lands in the same state. The comment in terraform/modules/shared/ecs-service/routing.tf has the full story and tracks the upstream issues that would end the warning.

Repository layout

# Go binary executed in AWS to provision DB & secrets
cmd/provision/                the Lambda: phase dispatch, seeding, OpenBao, funding, appliances
internal/keygen/              Ed25519 identities, secp256k1 wallets, UCAN proofs
internal/dbinit/              idempotent role and database creation
internal/vaultinit/           OpenBao init, mounts, hilt's AppRole
internal/ssmstore/            the never-overwrite parameter store
internal/fund/                the three FilecoinPay transactions
internal/onboard/             the appliance registration writes and their verification
build/                        Lambda container image
scripts/fund-payer.sh         invokes the fund phase, with a confirmation prompt
scripts/mint-appliance-token.sh  issues an appliance's unseal credential, wrapped
scripts/onboard-appliance.sh  registers an appliance and returns its S3 proof
scripts/retire-region.sh      removes a region's Ingot identity from hilt and SSM
scripts/refresh-bump-prs.sh   rebuilds every open bump branch on top of main
scripts/set-dev-pin.sh        pins one dev service at one image digest
scripts/smoke-test.sh         checks a deployed stage over public HTTPS
scripts/tail-logs.sh          prints the tail of every log group a stage owns
scripts/wait-services-stable.sh  waits for a cluster's ECS services to reach steady state

# Documentation beyond this file
docs/appliance-onboarding.md  the runbook for admitting a regional appliance
docs/observability.md         what a stage ships to Grafana and the queries that find it
docs/decisions/               why a thing is the way it is, one file per subject

# Infra configuration
terraform/
  modules/                                         the wiring; see Stages below
  envs/                                            one directory per root module
    bootstrap/<account>/account/                   state bucket, CI roles
    bootstrap/<account>/<region>/                  the image registry, telemetry egress to Grafana
    dev/platform/      dev/apps/                   applied on every push to main
    staging/platform/  staging/apps/               applied on every push to main
    prod/platform/     prod/apps/                  platform applied on every push to main; apps not yet

# Deployment
.github/workflows/check-and-deploy.yml    check, then plan on a PR or apply and smoke-test on main

Stages

A stage is a directory pair under terraform/envs/, backed by two state files in the account's bucket. Everything is namespaced by the stage name, so stages coexist in one AWS account without colliding:

  • fc-<stage>-* resources
  • /forge-central/<stage>/* parameters
  • RFC service identities under the stage's hostname_suffix (for example, upload.latest.dev.fil-forge.com in dev)

fc is short for forge-central, this repository's own deployment (as opposed to deployments of regional nodes). It is kept short because a target group name is capped at 32 characters, and <prefix>-<stage>-signing-service has to fit inside it. Only AWS resource names abbreviate; paths and namespaces spell forge-central out, since nothing there is close to a length limit.

Files:

terraform/envs/<stage>/platform/   VPC, RDS, S3, DynamoDB, ALB, OpenBao, provision Lambda
  main.tf                  module "platform" plus what this stage overrides
  terraform.tfvars         committed, non-secret: DNS, chain, contracts
  outputs.tf               re-exported for the apps root
  image.auto.tfvars        committed Lambda digest; published in dev, copied on promotion
  versions.tofu            OpenTofu version, S3 backend, providers
  versions.tf              refuses Terraform; OpenTofu never reads it

terraform/envs/<stage>/apps/       the six ECS services
  main.tf                  reads platform outputs via terraform_remote_state
  terraform.tfvars         committed: image digests
  versions.tofu            as above, with this root's own state key
  versions.tf              as above

Both roots stay short because modules/platform and modules/apps hold the wiring. That is the point of the split: a stage's root says what differs, and nothing else can drift between stages.

The module tree mirrors that split, so each root can name the directories it depends on:

terraform/modules/
  platform/                everything the platform root builds
    main.tf                the wiring, calling the eight below
    network/ kms/ database/ storage/ ingress/ provision/ openbao/
    aurora/                prod's database, in place of database/
    log-forwarding/        the role CloudWatch Logs ships a stage's groups to Grafana with
  apps/                    the six ECS services
  shared/                  used by more than one root
    ecs-service/           apps, and openbao inside platform
    constants/             every root, bootstrap included
  ecr/                     regional bootstrap only: the image registry
  telemetry/               regional bootstrap only: the Firehoses and metric stream to Grafana
  tfstate/                 account bootstrap only: the state bucket
  github-actions-iam/      account bootstrap only: the two CI roles

A module used by exactly one root lives under that root's composite module. shared/ holds the two that genuinely cross the boundary. Adding a module therefore never means editing a trigger pattern anywhere, because the workflow applies both roots on every push rather than choosing between them by path.

Chain configuration lives in the platform root and the apps root reads it from there, so a stage has one set of contract addresses rather than two copies to keep in step. That mirrors smelt's shared smart-contracts.env.

DNS

The public Forge domains delegate the zones used by this deployment to Route53. Dev stages share dev.fil-forge.com, staging uses staging.fil-forge.com, and production uses fil-forge.com directly.

fil-forge.com DNS
  ├── NS dev  ──►  Route53 zone dev.fil-forge.com  (non-prod account)
  │                 ├── upload.latest.dev.fil-forge.com
  │                 ├── ssm.latest.dev.fil-forge.com
  │                 └── upload.<STAGE>.dev.fil-forge.com
  ├── NS staging ─► Route53 zone staging.fil-forge.com  (non-prod account)
  │                 ├── upload.staging.fil-forge.com
  │                 └── ssm.staging.fil-forge.com
  ├── NS upload ──► Route53 zone upload.fil-forge.com  (production account)
  ├── NS ssm ─────► Route53 zone ssm.fil-forge.com     (production account)
  └── …            one zone per public service name

Adding a personal stage beneath dev.fil-forge.com requires no change to the DNS project. Shared domain roots such as staging are delegated once before a stage uses them.

Production carries no stage label: upload.fil-forge.com. Ephemeral and personal stages use <STAGE>.dev.fil-forge.com; this repository's dev stage is the long-lived latest stage, so Sprue is upload.latest.dev.fil-forge.com. The shared staging deployment is the RFC's separate <service>.staging.fil-forge.com namespace; it is not a stage label beneath dev.fil-forge.com.

Public labels are stable identities rather than implementation names: Sprue is upload, Hilt is auth, Swarf is revoke, piri-signing-service is signer, and Delegator and Indexer use delegator and indexer.

Two per-stage settings follow, and this is where they diverge:

  • zone_name is the delegated Route53 zone records are written into. Dev stages share dev.fil-forge.com; staging uses staging.fil-forge.com.
  • hostname_suffix is what that stage's hostnames end with, which for non-prod includes the stage label.
  • ingot_hostname_suffix is the corresponding suffix in the filonecontent.com namespace. Ingot identities are did:web:s3.<REGION>.<ingot_hostname_suffix>.

The delegation itself lives in fil-one/infrastructure and is added once per dev/staging domain root: an aws_route53_zone for the delegated name, plus a Cloudflare NS record carrying that zone's four name servers.

Production has no domain root to delegate, because its service names sit directly beneath fil-forge.com. Each public service name is a Route53 zone of its own instead, created by terraform/envs/bootstrap/prod/account from the constants module's public_hostname_labels, so the zones survive a rebuild of the prod stage. fil-one/infrastructure carries one Cloudflare NS record per zone, copied from that root's service_zone_name_servers output. The prod platform root sets no zone_name, which tells the ingress module to write each record into its hostname's own zone.

Those records are created with proxied = false, which matters: these hostnames serve did:web documents and terminate their own TLS at the ALB, so Cloudflare must not sit in front of them.

Certificates belong here, not in the fil-one/infrastructure project.

The ingress module issues *.<hostname_suffix>, writes the DNS validation records into the delegated zone, and waits for validation. In production a wildcard would validate through a record in the Cloudflare apex, so the certificate names every public hostname and validates each one in its own zone. Two reasons it cannot be one central certificate:

  • An ALB needs its certificate in the ALB's own region. A us-east-1 certificate, which is what CloudFront requires, cannot be attached.
  • A wildcard covers exactly one label, so *.dev.fil-forge.com does not match upload.latest.dev.fil-forge.com. Each stage needs its own.

What survives a destroy

terraform destroy deletes no parameter this project generates. The provision Lambda creates them, so Terraform has no record of them and never removes them. An accidental destroy therefore cannot burn a funded wallet or invalidate a DID that storage providers have already registered against.

They also stay readable, which takes deliberate arrangement. SecureStrings are encrypted under the account's AWS-managed SSM key rather than the stage's own customer-managed key. The stage's key is destroyed with the stage, and a key in PendingDeletion stops serving decryption at once, so tying the parameters to it would leave every secret unreadable the moment the stage came down and would fail the next apply that tried to rebuild it. The stage's key seals OpenBao and nothing else, and what it protects is meant to die with the stage: OpenBao's storage sits in the stage's database and goes at the same time.

So a destroyed and recreated stage silently comes back with its previous identities and wallets. An appliance's stored delegation is one of them: a rebuilt stage still holds the proof hilt signed for whatever Ingot DID it was addressed to, and onboarding returns that copy rather than issuing a new one. That is usually what you want, and it is occasionally a surprise, so check before assuming a rebuilt stage is fresh:

aws ssm get-parameters-by-path --path /forge-central/dev --recursive \
  --query 'Parameters[].Name' --output text

To retire a stage, delete the parameters after the destroy, having first confirmed the wallets hold no funds:

# Check the balances first. This is not reversible.
aws ssm get-parameter --name /forge-central/dev/signing-service/payer-key.address
aws ssm get-parameter --name /forge-central/dev/delegator/transactor-key.address

aws ssm get-parameters-by-path --path /forge-central/dev --recursive \
  --query 'Parameters[].Name' --output text \
  | xargs -n 10 aws ssm delete-parameters --names

Runbook

Prerequisites

  • AWS CLI, with credentials for the target account.
  • OpenTofu 1.12 or newer, for the bootstrap roots and for reading a stage's outputs. Stage applies themselves run in GitHub Actions. Terraform is not an alternative here and every root refuses it outright: it stamps a version marker into state that OpenTofu reads as being from the future, so one terraform apply would lock OpenTofu out of that state.
  • Docker with buildx, for make publish.
  • Go and make, for make check and make test.
  • ShellCheck, for the shell half of make check. CI pins 0.11.0, so that build is the one that decides a merge; an older local one can pass a script CI rejects.
  • Foundry's cast, only to read chain balances by hand. Nothing in the deploy path needs it.

How each part is deployed

Part How it is deployed
bootstrap roots tofu apply run locally, always
provision image make publish run locally, pushed to ECR by hand
dev platform, apps GitHub Actions, on every push to main, with no approval step
staging platform, apps GitHub Actions, on every push to main, with no approval step

The dev and staging stages deploy themselves after their initial bootstrap. .github/workflows/check-and-deploy.yml runs make check on every pull request and every push to main; a pull request then plans all four roots, and a push applies and smoke-tests both stages. The OpenTofu version is pinned in the workflow rather than taken from an operator's machine.

apps reads platform's state through terraform_remote_state, so ordering matters: apply-dev-apps waits on apply-dev-platform, and the corresponding staging jobs have the same edge. An apps job therefore never plans against outputs an in-flight platform apply is about to change. Every root is applied on every push, even one that touched only one of them. An empty plan costs about a minute, and it means there is no path-filter list to forget to update when a module moves.

In a pull request all four plans run at once, and each apps plan is computed against its last applied platform state rather than against this pull request's platform plan. A change to a platform output that apps consumes therefore shows its real apps plan only after platform applies.

An apply reports success as soon as AWS accepted the change, which for an ECS service means a task definition was registered rather than that a task is serving traffic on it. Each apps apply therefore ends by running scripts/wait-services-stable.sh, which waits for every service in the cluster to reach steady state, and a task that never becomes healthy fails the job after twenty minutes. Without that wait a smoke test can pass against the revision the push replaced, because a rolling update keeps the old task answering. The script names the services it is still waiting on as it polls, and every two minutes prints their task counts, their deployments and their recent ECS events, so a long wait says whether a rollout is slow or has stopped moving.

smoke-dev and smoke-staging then call make smoke for their stage and retry for four minutes. Steady state covers the task; a newly created Route53 record or listener rule in front of it can take a moment longer. The smoke checks need no credentials because every request goes over public HTTPS. See Smoke-testing a stage.

When either the wait or the smoke test fails, scripts/tail-logs.sh prints the tail of every log group the stage owns into the run, so the diagnosis is where the failure is. The groups are discovered from CloudWatch, so a service added to either root is covered. It runs on the apply role, because reading log events needs logs:FilterLogEvents and the plan role deliberately has none of it.

A failed run on main posts to #filone-alerts in Slack with the stage that failed, the failed jobs, the commit subject, its author and a link to the run. The stage comes from the job name: apply-<stage>-<root> and smoke-<stage> name a stage, apply-grafana reports as grafana, and check as itself. When the commit is an image bump, one more line names the service commit that produced the image and who wrote it, looked up from the source commit link Bump deployed image puts in the commit body. A lookup that fails drops the line and still sends the alert. Any failed job triggers it, from make check through the smoke test. Pull request failures are not announced, because the author already sees the red check on the pull request. The job reads one repository secret, SLACK_BOT_TOKEN, holding the bot token of a Slack app with the chat:write scope; without the secret the notification step fails and nothing else about the run changes.

AWS credentials are never stored. Each job assumes an IAM role in the target account through GitHub's OIDC federation, and the credentials expire with the job. There are two roles, and the split matters: GitHub runs the workflow file from a pull request's own head, so the role a plan job uses can describe infrastructure and read nothing, and the role that can change anything is reachable only from refs/heads/main. See terraform/modules/github-actions-iam.

The workflow applies the prod platform root on every merge to main. The prod apps root has no CI job yet.

See Planned work for the manual steps that remain.

Staging's capacity, durability, identity and promotion choices are recorded in the staging environment decision.

Setting up an AWS account

The bootstrap roots are always applied locally. They run rarely, they create the things everything else depends on, and so there is nothing for a pipeline to trigger on and no earlier apply to have created their state. They come in two kinds, and the split is what keeps a second region cheap:

  • bootstrap/<account>/account/ holds the state bucket every other root in the account keeps its state in, and the two CI roles GitHub Actions assumes to plan and apply the stages. One per account: a bucket name is global and IAM is not regional, so a second region must not create these again.
  • bootstrap/<account>/<region>/ holds the ECR repository for the provision Lambda image and the telemetry egress to Grafana Cloud. In prod it also holds the KMS keys of the Aurora cluster and of OpenBao's seal, which must outlive the platform root. One per account and region, described in Setting up an AWS region.

Both accounts this project uses already have an account root, and both have been applied.

Copying the root for a new account

Copy a bootstrap/<account>/ directory, both the account/ root and the regional one beside it. In the copies, point each provider at the account id it belongs to, set the bucket name in account/main.tf and in both versions.tofu backend blocks, and add that id to terraform/modules/shared/constants if it is not there yet.

Every root reads its account id from that module, so an apply run with credentials for the wrong account fails at plan time rather than building a second working copy of the stage somewhere unexpected.

The GitHub OIDC provider

One thing has to exist before the account root can be applied, and nothing here creates it: the GitHub OIDC provider, https://token.actions.githubusercontent.com. It is one per account and shared with every other repository that deploys into that account, so modules/github-actions-iam reads it as a data source rather than owning it. Creating it here would fail for the second repository to try, and a destroy would lock the first one out of its own CI.

Both accounts this project uses already have it, so this matters only for an account nobody has deployed to from GitHub Actions before. Check:

aws iam list-open-id-connect-providers \
  --query "OpenIDConnectProviderList[?contains(Arn, 'token.actions.githubusercontent.com')]"

If that comes back empty, create it before applying the account root, or the apply that makes the CI roles fails on the lookup with NoSuchEntity:

aws iam create-open-id-connect-provider \
  --url https://token.actions.githubusercontent.com \
  --client-id-list sts.amazonaws.com

sts.amazonaws.com is the audience aws-actions/configure-aws-credentials requests when the workflow does not override it, and it is what both trust policies require in their aud condition. Omit it from the client id list and every sts:AssumeRoleWithWebIdentity call is rejected. No --thumbprint-list: AWS no longer validates one for this provider.

First apply of the account root

The account root is the awkward one: its own backend points at the bucket it creates, so the first apply in a fresh account cannot use that backend. Run it against a local backend once, then move its state into the bucket it just made. .gitignore already ignores *_override.tf, so the override cannot be committed by accident:

cd terraform/envs/bootstrap/<account>/account

printf 'terraform {\n  backend "local" {}\n}\n' > backend_override.tf
tofu init
tofu apply -target=module.tfstate    # the bucket, and nothing else yet

rm backend_override.tf
tofu init -migrate-state             # local state moves into the bucket
rm -f terraform.tfstate terraform.tfstate.backup

tofu apply                           # the CI roles

Note the two role ARNs it prints. .github/workflows/check-and-deploy.yml names them literally, so if they differ from what is there, the workflow needs updating.

Every root after this one is ordinary, because its backend block points at a bucket that now exists. The regional root beside it comes next.

Setting up an AWS region

The regional root, bootstrap/<account>/<region>/, holds two things:

  • forge-central/provision, the ECR repository for the provision Lambda image. Lambda pulls an image only from ECR in the same region as the function. Stages sharing an account and region share the repository and pin different digests. Its repository policy is what lets Lambda pull the image; without it, creating a stage's provision Lambda fails with AccessDeniedException.
  • The telemetry egress to Grafana Cloud: one log Firehose per stage, from the stage list in terraform/modules/shared/constants, and one CloudWatch metric stream, which covers one account in one region.

The first region of an account already has this directory next to the account root. For a further region, copy it:

cp -r terraform/envs/bootstrap/nonprod/us-east-2 terraform/envs/bootstrap/nonprod/us-west-2

Change two things in the copy: the region in the provider block, and the key in the backend "s3" block in versions.tofu (bootstrap/us-west-2.tfstate). Leave the backend's region at us-east-2. It names the region the state bucket is in, and the bucket is one per account, created by the account root. Pointing it at us-west-2 makes tofu init fail against a bucket that is sitting right there. Nothing else needs changing and nothing needs deleting: the account-scoped resources are not in this directory to begin with, and the telemetry module's backup bucket and IAM roles, which share the account's namespace, take the region from the provider and so get their own names.

Grafana values

Applying the regional root needs three values from the Grafana Cloud stack, passed as environment variables and committed nowhere. All three are in the Forge Central item of the Fil One vault in 1Password, under the GRAFANA section. With the 1Password CLI signed in:

export TF_VAR_grafana_logs_user="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_LOGS_USER')"
export TF_VAR_grafana_metrics_user="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_METRICS_USER')"
export TF_VAR_grafana_push_token="$(op read 'op://Fil One/Forge Central/GRAFANA/GRAFANA_CLOUD_PUSH_TOKEN')"

op item get --vault "Fil One" "Forge Central" lists the section's fields without revealing the token. The two *_USER values are the Loki and Prometheus instance ids of the stack. The item's two *_URL fields are the Firehose delivery endpoints, which are what the module defaults to, so nothing needs setting for them. They are not the plain Loki and Prometheus push URLs Alloy uses on the appliances: Firehose has its own delivery format and Grafana receives it on aws-logs-* and aws-metric-streams-* hosts.

The token is a Grafana Cloud access policy token with the logs:write and metrics:write scopes, created under Security → Access Policies in the Grafana Cloud portal, the same kind infra-nodes' runbook describes for the appliances. Rotating it is a new token in the 1Password item and a tofu apply of this root. The token ends up in this root's state, which is why the Firehoses live in this root rather than in a stage root; see the telemetry decision.

Applying the root and filling the repository

The bucket already exists, so there is no bootstrap dance here:

cd terraform/envs/bootstrap/nonprod/us-east-2
tofu init
tofu apply                           # the image registry and the telemetry egress

Every image this project publishes to ECR lives under the forge-central/ prefix, one repository per image. Per-image repositories are what make per-image push permissions, lifecycle policies, and tag immutability possible.

Then fill the repository. The provision image is built and pushed by hand from a developer machine; nothing builds it automatically. make publish needs Docker with buildx and AWS credentials for the target account, and it creates a docker-container builder on first use, because Docker Desktop's default builder cannot push by digest and cannot cross-build for arm64.

make publish STAGE=dev                            # AWS_REGION defaults to us-east-2
make publish STAGE=<stage> AWS_REGION=us-west-2   # a further region

It pushes by digest and writes no tag, so the digest a stage pins is the only reference to the image. That is why the repository rejects tags and carries no expiry rule: an untagged image is indistinguishable from one a stage is running, and Lambda does not survive having its image deleted. Prune by hand when the image count starts to bother you.

A stage needs nothing copied from the bootstrap output. It builds the image URL from its own account and region, which is the only registry its Lambda can pull from anyway; the account ids and the repository name live in terraform/modules/shared/constants. The digest is derived from the image rather than from where it is stored, so a stage in a new region can pin the same digest an existing stage already runs.

Adding a stage

cp -r terraform/envs/dev terraform/envs/bajtos

Then, in the copy:

  1. Set the key in both versions.tofu files to bajtos/platform.tfstate and bajtos/apps.tfstate, and the key in the apps root's terraform_remote_state block to match the platform one. The bucket is already right: it is per account, and personal stages share the non-prod account.
  2. Change stage = "dev" to "bajtos" in platform/main.tf, and the Stage default tag in both roots.
  3. In platform/terraform.tfvars, set hostname_suffix to <STAGE>.dev.fil-forge.com and ingot_hostname_suffix to <STAGE>.dev.filonecontent.com. Leave zone_name as dev.fil-forge.com: the zone is already delegated and shared by every dev stage, so the DNS project needs no change. Point the chain block at the network this stage transacts against.
  4. Add the stage to .github/workflows/check-and-deploy.yml: two more entries in the plan matrix, named bajtos-platform and bajtos-apps, two more apply jobs copied from dev's, with apply-bajtos-apps needing apply-bajtos-platform, a smoke job for the new stage, and a diagnose job copied from dev's, which names the stage whose logs it tails. Add both apply jobs and the smoke job to notify-failure's needs; a failure in a job it does not name announces nothing.
  5. Add "bajtos" to nonprod_stages in terraform/modules/shared/constants/outputs.tf and apply both bootstrap roots for the account. The account root grants the CI roles state access by stage prefix and the regional root creates the stage's log Firehose, so a stage missing from the list either cannot read its own state or fails its first platform apply creating its subscription filters. Why there is one list is in the telemetry decision.
  6. Apply the new stage's platform root locally once. The apps root reads the platform's remote state, so its first CI plan cannot run until that state exists. Do not apply the apps root yet; the workflow will do that after the change merges.
  7. Update the branch protection rule on main. The new stage adds two required checks, plan-bajtos-platform and plan-bajtos-apps, and a rule that does not name them will merge a pull request whose stage plan failed.
  8. Merge. The stage's platform job reconciles the VPC, RDS, OpenBao and secrets, its apps job applies the six services, and its smoke job tests the public endpoints. Nothing in the apps root needs starting by hand.

The first platform apply is slow: it waits for the OpenBao task's cold start before it can initialise it, inside a synchronous Lambda call that Lambda caps at 15 minutes. If it times out there, re-run the job. The seed phase regenerates nothing that already exists, which is what protects funded wallets.

Prod differs from dev inside main.tf rather than by being a different shape: an Aurora cluster with a writer and a reader in its own subnets, deletion protection on, KMS keys for the cluster and for OpenBao's seal from the regional bootstrap, a larger OpenBao connection budget, a zone per public hostname, and a provision digest pinned in terraform.tfvars. It lives in its own account and deploys on every merge, like staging. Its choices are recorded in the first prod stack decision, and staging's in the staging environment decision.

A personal sandbox stage

Stage names are not limited to dev and prod. Copy envs/dev to envs/<you>, give it its own state key in the same bucket, and apply it from your machine: no commit, no merge, no workflow run to wait for, which is the fastest loop for iterating on the provision Lambda. Leave it out of check-and-deploy.yml; that is what makes it yours.

Set enable_log_forwarding = false in the copied platform/main.tf. The regional bootstrap root creates a log Firehose only for the stages in nonprod_stages, and a platform apply that forwards to a Firehose that does not exist fails creating its first subscription filter. The stage's logs stay in CloudWatch, where scripts/tail-logs.sh reads them. A sandbox that needs its logs in Grafana follows step 5 of adding a stage instead, at the cost of a commit.

What it costs is everything the dev stage gets from the workflow: a plan on every pull request, applies that cannot disagree with main, and an OpenTofu and provider version that is the same for everyone. Use it to iterate, not to host anything anyone depends on.

Onboarding a regional appliance

A stage seals the appliances named in its appliance_regions, and make mint-appliance-token issues a node's unseal credential. The ordering across both repositories, how the credential is delivered to a node operator, and what retiring a region destroys are in docs/appliance-onboarding.md.

Funding the wallets

The seed phase mints two secp256k1 wallets and reports their addresses. Both start empty, and nothing works until they hold funds:

tofu -chdir=terraform/envs/dev/platform output wallet_addresses

Gas, for both wallets. The delegator's transactor signs provider approvals and the payer signs PDP operations, so each needs tFIL on Calibration or FIL on mainnet. Faucet: https://faucet.calibnet.chainsafe-fil.io/

USDFC, for the payer only. Faucet: https://forest-explorer.chainsafe.dev/faucet/calibnet_usdfc — capped at 10 USDFC per day, which is why the amounts below stay small.

Depositing into FilecoinPay. USDFC sitting in the payer's wallet is not enough. Creating a proof set locks up around 0.9 USDFC, and lockup can only draw on funds deposited into the FilecoinPay contract, so a freshly faucet-funded wallet still fails with InsufficientLockupFunds(..., Available=0).

make fund-payer STAGE=dev

That runs three transactions, the same ones as smelt's scripts/staging-fund-payer.sh:

  1. USDFC.approve(FilecoinPay, amount) — let FilecoinPay pull the tokens
  2. FilecoinPay.deposit(USDFC, payer, amount) — credit the payer's account
  3. FilecoinPay.setOperatorApproval(USDFC, FWSS, true, rate, lockup, period) — let warm storage lock it up

The signing happens inside the provision Lambda, in AWS.

The script invokes the Lambda twice. The first call reads the chain and prints what it would do, signing nothing. You then confirm, and the second call broadcasts.

Amounts default to smelt's, which stay under the faucet's daily cap. Override them per run:

make fund-payer STAGE=dev DEPOSIT=5
make fund-payer FUND_ARGS="--rate-allowance 0.2 --lockup-allowance 5"
scripts/fund-payer.sh --stage dev --deposit 3 --force-deposit

Terraform never invokes this phase. An apply must not move money, so funding is always an explicit operator action.

Two preconditions are checked before anything is signed: the RPC must report the chain id the stage expects, and the payer wallet must already hold at least the deposit amount. Neither is recoverable by this tooling, so both fail loudly.

To read the balances without invoking anything:

PAYER=$(aws ssm get-parameter --name /forge-central/dev/signing-service/payer-key.address \
  --query Parameter.Value --output text)

cast call "$USDFC_TOKEN_ADDRESS" "balanceOf(address)(uint256)" "$PAYER" \
  --rpc-url https://api.calibration.node.glif.io/rpc/v1

cast call "$FILECOIN_PAY_ADDRESS" \
  "accounts(address,address)(uint256,uint256,uint256,uint256)" \
  "$USDFC_TOKEN_ADDRESS" "$PAYER" \
  --rpc-url https://api.calibration.node.glif.io/rpc/v1

Iterating on the provision Lambda

make publish STAGE=dev

That writes the new digest into the stage's image.auto.tfvars, so there is no line to edit by hand. Commit that file and merge it. The stage is planned by a workflow, which sees only what is in version control, so a digest left on your machine is applied nowhere.

Dev and staging share an ECR repository. Promote the Lambda to staging by copying dev's digest into terraform/envs/staging/platform/image.auto.tfvars. The image is already in ECR, so the promotion needs no build or push.

Prod is a separate account with its own ECR repository, so a dev digest means nothing there. Promote the Lambda to prod by publishing into the prod account:

make publish STAGE=prod

For prod the command writes nothing. It prints the provision_image_digest line to paste into terraform/envs/prod/platform/terraform.tfvars, where prod pins its digest. Commit that file and merge it.

Deploying a service

Change its digest in the stage's image_digests and merge. Every stage pins digests, dev included: dev is applied on every push to main, and a rolling tag would make what dev runs depend on when a task last restarted rather than on what was merged.

For dev, the services do this themselves. A publish workflow dispatches a bump-deployed-image event carrying the digest it just pushed, and bump-deployed-image.yml opens a pull request that changes the one line, with auto-merge enabled so the deploy lands as soon as the required checks pass. Those pull requests come from the fil-forge-bot GitHub App on the branch bot/bump-<service>-image-dev, one branch per service, so a second publish updates the open request instead of stacking a stale one beside it.

Two things keep those pull requests mergeable while several are open at once. The pins are spaced a blank line apart, because git conflicts on changes to adjacent lines and each bump rewrites one line. And refresh-bump-prs.yml rebuilds every open bump branch on top of main, because the ruleset will not merge a branch that is behind. It runs when main moves, when a bump branch is pushed, and when a bump pull request is opened: a bump commit is built from the main its run checked out, which can be behind by the time the push lands, and on a service's first bump the branch is pushed before the pull request exists to be found. A rebuild keeps the original commit message, so the link to the pull request that published the image survives. A branch whose digest main already pins is closed instead, and one whose service someone else moved meanwhile is left alone: which digest dev should run is then a question rather than an edit, and the pull request shows the conflict it has.

The same workflow bumps any of the six services on demand:

gh workflow run bump-deployed-image.yml -R fil-forge/infra-central \
  -f service=sprue -f digest="$(crane digest ghcr.io/fil-forge/sprue:main)"

Locally, scripts/set-dev-pin.sh makes the same edit. It is what both workflows run, so a pin written by hand comes out identical to one written by the bot:

scripts/set-dev-pin.sh sprue "$(crane digest ghcr.io/fil-forge/sprue:main)"

A service is wired up with a dispatch step in its publish workflow plus the fil-forge-bot credentials in that repository. The receiver only accepts dispatches made as that app, and the source_repo in the payload has to be the repository the service is published from, so the required client_payload is service, digest and source_repo; commit, pr_url and run_url are provenance links the commit message uses when present.

Staging and prod stay manual: a promotion copies dev's reviewed digest in a deliberate pull request.

Confirming nothing was regenerated

The most important check after any apply. Read created_parameters from the stage's platform apply job, such as apply-dev-platform or apply-staging-platform, or from a shell:

tofu -chdir=terraform/envs/dev/platform output created_parameters

Empty means every key already existed and was reused. A non-empty list after the first apply of a stage means something was minted; find out what before assuming a wallet is intact.

Smoke-testing a stage

make smoke STAGE=dev

Every public service is checked over public HTTPS, needing no AWS credentials. A 200 from the health path covers the whole ingress route in one request: the Route53 record, the wildcard certificate, the listener rule, the target group and a task passing its container health check.

The second check is the one health cannot make. sprue, hilt and swarf mint an ephemeral identity when no key is supplied and report themselves healthy either way, so /.well-known/did.json is read and its id compared against did:web:<hostname>. A mismatch means the service is running an identity nothing has registered against.

OpenBao is checked too, at ssm.<suffix> rather than at its own name. The request omits the uninitcode=200 its ALB health check passes: ECS has to keep a fresh task alive long enough for the provision Lambda to initialise it, but a stage that has finished deploying and is still uninitialised is a failure.

The same command runs in CI for dev and staging after every push to main. See How each part is deployed.

The script reads hostname_suffix from the stage's platform/terraform.tfvars, so it needs no Terraform state and no TFE token. Services are probed concurrently: a task that accepts the connection and never replies waits out the whole timeout, and several of those in sequence is a minute of nothing.

plc and piri-signing-service get the health check alone, and the output says so rather than passing over it. plc publishes no identity. piri-signing-service takes a did:web and serves no document at it, so nothing resolves it today: it is the only service no other service addresses by DID.

Run by hand it also says nothing about which revision answered. No service reports its build, so a stage mid-rollout can pass on the old task. In CI the stage's apps apply job closes that gap by waiting for steady state first.

Rotating a service identity

Delete the parameter:

aws ssm delete-parameter --name /forge-central/dev/swarf/identity

Then force the seed phase to run again. A plain run will not do it: the phase is an aws_lambda_invocation, which re-invokes only when its input changes, and deleting a parameter changes nothing Terraform can see. See Forcing a provision phase to re-run.

The new DID appears in service_dids, and the rotation also refreshes /forge-central/<stage>/<service>/identity.did, which holds the same value for anyone reading it without decryption rights. Anything that had registered the old DID has to be told about the new one, which is why this is a deliberate act rather than something an apply does on its own.

Rotating an identity that signs a proof — sprue, indexer or etracker — re-issues that proof automatically in the same apply, because the old delegation would no longer verify against a key that does not exist.

Why proofs are not rewritten every apply

A UCAN delegation is public but not reproducible: ucantone mints a random 16-byte nonce per delegation, so signing the same request twice produces different bytes and a different CID. Rewriting on every apply would churn the parameter and invalidate anything holding the previous delegation.

So proofs are written once and then left alone, and re-issued only when their issuer key was freshly minted. smelt tracks the same dependency, skipping a committed proof unless one of the keys behind it was regenerated that run.

The practical consequence: you cannot verify a proof by regenerating it and diffing. Only the framing is reproducible, and that is what the tests pin: a textual container is stored as ucantool writes it, trailing newline included, while a bare DAG-CBOR delegation is stored base64-encoded.

Why the delegator's proofs are stored base64

Every proof reaches its consumer as an environment variable that the task's entrypoint writes to a file, and an environment variable cannot carry a NUL byte. A bare DAG-CBOR delegation is binary and contains them, so the container never starts and runc reports only the variable's name. The delegator's two proofs are therefore stored base64-encoded and decoded on the way to the file, which is what secret_files_base64 on modules/shared/ecs-service is for. hilt's proof is a base64+gzip container, already text, and travels as it is.

Rotating hilt's OpenBao credential

aws ssm delete-parameter --name /forge-central/dev/hilt/vault-secret-id

Then force the vault phase to run again. There is no way to do that from a remote run yet — see Forcing a provision phase to re-run.

The vault phase also self-heals: if a stored secret_id no longer authenticates, because OpenBao's storage was rebuilt underneath it, the next run replaces it rather than leaving hilt unable to start.

Development

make check   # gofmt, go vet, go test, tofu fmt, shellcheck
make test    # go test alone, for the inner loop

Both run offline against no deployed stage, which is why make smoke is separate: it needs a stage to be up. See Smoke-testing a stage.

Dependency updates

Dependabot opens the pull requests, .github/dependabot.yml says which and how often, and auto-merge-dependabot.yml merges the ones that are minor or patch bumps. The merge is squashed and armed through fil-forge-bot, so main moves only after make check and all required plans have passed. The push that lands applies both shared non-prod stages the same way any other merge to main does.

A major bump stays open for someone to read. So does a group whose highest change is a major, and so does any Dependabot branch that carries a commit Dependabot did not write.

Dependabot rebases its pull requests in place, and the workflow runs again on each new head, so a branch that stops qualifying also loses the auto-merge it was given earlier. Each auto-merge is bound to the head the workflow inspected, which is what stops a head that arrives mid-run from merging on the previous head's decision.

Planned work

See the following Linear tickets:

  • FIL-1147 Stand up the first, disposable Forge Central prod stack
  • FIL-1394 Hold prod applies for a human approval
  • FIL-1396 Reset production after the test run
  • FIL-1156 Narrow the apply role's IAM policy
  • FIL-1090 Tooling for removing an appliance node from the network
  • FIL-1091 Hilt: API to remove Ingot node
  • FIL-1154 Hilt authenticates to OpenBao with AWS IAM auth
  • FIL-1148 Publish the provision Lambda image from CI
  • FIL-1150 Enable the OpenBao audit log on Forge Central
  • FIL-1157 Decide the OpenBao availability target for Forge Central
  • FIL-1149 Write and rehearse the RDS restore procedure
  • FIL-1151 Grafana alarms for Forge Central services
  • FIL-1152 Expose metrics from swarf, delegator and piri-signing-service
  • FIL-1153 Write down the running cost of a Forge Central stage
  • FIL-1160 Replace static Postgres passwords with RDS IAM authentication
  • FIL-1155 Own the RDS parameter group and pin rds.force_ssl
  • FIL-1161 Verify the RDS server certificate in every Forge service
  • FIL-1162 Zero-downtime upgrades of Forge Central services
  • FIL-1158 Automate promotion of Forge Central from dev to staging

Forcing a provision phase to re-run

aws_lambda_invocation re-invokes only when its input changes, which is what seed_trigger and vault_trigger in modules/platform exist for. Rotating an identity or hilt's OpenBao credential needs one of them bumped.

Neither is reachable today. The stage roots do not expose them, and the workflow passes no -var flags, so the only way to bump one is to edit envs/<stage>/platform/main.tf and merge.

That is arguably the right answer rather than a limitation. A committed, reviewed diff is what re-running the thing that mints wallets should cost, and it stays visible afterwards. The alternative — a workflow input anyone with write access could set — puts that one text field away. A workflow_dispatch input would be the middle ground if the merge ever proves too slow.

Related

  • smelt — the single-VM Docker Compose deployment this replaces, and the source of the key generation code.
  • smelt#11 — its appliance registration scripts, which specified the writes the onboard phase performs, including the region mismatch hilt's own error cannot distinguish.
  • fil-forge/infra-nodes — the appliance side: the node this repository seals, registers and can revoke.
  • fil-one/RFC#21 — regional security and key management, which makes this OpenBao the root of trust for appliances.

About

Deployment configuration for Forge central services

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages