Control plane — thin bootstrap + SSH provisioning
The split
Provisioning a new machine has two phases with different trust and cost profiles. They are deliberately separate:
Phase 1 — thin bootstrap (human-in-the-loop, via the new machine's browser/agent). The ONLY steps that require the new machine's local agent:
- Download the kit (
site/index.htmlstep 2 — the curl tarball). bin/pairing.sh gen— show the human the 4-word code + public key.- Append the operator's public key (from the pre-filled prompt) to the
target user's ~/.ssh/authorized_keys. 4. Ensure sshd is installed and running (no-op on identical containers). 5. Set REMOTE_FWD_PORT / TERMINAL_FWD_PORT in bin/gcp-tunnel-up.sh (values from the pre-filled prompt) and dial the tunnel.
Then STOP. Everything else is cheaper, faster, and more reliable over SSH.
Phase 2 — SSH provisioning (operator side, deterministic). The operator uplink (super) runs bin/provision-remote.sh <name>, which finishes the job over SSH through the VM jump host. No browser, no chat, no ambiguity about whether the machine is onboarded — the script verifies each step and reports.
Trust model
- Identity is centralized: the port registry (
gcp/machine-registry/)
plus the VM's authorized_keys are the source of truth for which machines exist — and /srv/board/allowed_signers + /srv/verify/levels.json are the source of truth for which agent identities are registered and at what level (verified, dev). - Secrets are decentralized: every secret is generated ON the machine that uses it (provision-remote.sh triggers remote generation; the plaintext never traverses the wire). The operator device and VM never hold machine secrets. - Access: the operator's public key (not a shared secret) is installed on each machine for support access — the consented front door from ETHICS.md. One keypair per operator; the private key never leaves the operator's machine. - Pairing is the trust root: the 4-word code binds the human's approval to the exact key being authorized. No code, no key install, no tunnel.
The SSH path
super (operator uplink) --(ProxyJump)--> [email protected] --(127.0.0.1:SSH_PORT)--> target
provision-remote.sh builds this as:
ssh -J [email protected] -p <ssh-port> <target-user>@127.0.0.1
Host keys: container host keys regenerate on rebuild, so TOFU pinning would break automation after every rebuild. The script does not pin target host keys (StrictHostKeyChecking=no, no persistent known_hosts). The trust root is the pairing ceremony plus the VM jump — both operator-controlled — not the target's ephemeral host key. This tradeoff is explicit and documented here, not silent.
The machine contract
A machine is not "onboarded" until it proves all four, verified by provision-remote.sh:
- SSH reachable — the operator can open a non-interactive SSH
session through the VM jump. 2. Tunnel live — the supervisor runs on the target; both ports listen on the VM's 127.0.0.1; the terminal subdomain answers (401 = healthy). 3. Scheduler present — some cron-like scheduler exists for rotation and watchdog jobs. v1 only checks and reports (crontab present?); installing cron bodies is a follow-up. 4. Human-notify path — there is a way to reach the human. v1: the main tmux session exists (banner delivery). Recorded, not yet automated.
What the provisioner does (phase 2)
- Resolve
<name>→ ports from the registry mirror. - Preflight SSH.
- Push
bin/(tar-pipe, no intermediate files). - Set the per-machine ports in the remote
gcp-tunnel-up.sh. - Generate secrets ON the target (
.ttyd-pass,.ttyd-cookie-secret). - Install the operator's public key (idempotent append).
- Run
recover-after-rebuild.sh(idempotent full provisioning). - Check scheduler + notify path; record the contract.
- End-to-end verify (supervisor, VM ports, terminal URL).
- Report. Stay silent on nothing — this is an attended run, so a full
report is the point (unlike the unattended watchdog).
Non-goals
- The provisioner never handles plaintext secrets (it triggers their
creation remotely). - The provisioner never touches the pairing ceremony — that stays human-in-the-loop, always. - The provisioner does not replace update.sh — updates remain pull-based per machine; provisioning is push-based, once.