ARCHITECTURE.md — the front-door system, end to end
This is the technical reference for the whole system: one GCP VM acting as orchestrator and front door for many machines (containers, operator devices), a public distribution channel for the provisioning repo, and a secrets model that keeps every machine's credentials its own.
The system in one picture
┌────────────────────────────────────────────────────┐
│ GCP VM — orchestrator / front door │
│ 34.139.37.135 (static IP) │
│ │
│ Caddy (TLS, Let's Encrypt, wildcard DNS) │
│ ├─ muse-dev.online → /srv/apex │
│ │ landing page + /.well-known/manifest.json │
│ ├─ muse-main.muse-dev.online → 127.0.0.1:7681 │
│ ├─ start.muse-dev.online → /srv/start │
│ ├─ dist.muse-dev.online → /srv/dist │
│ ├─ docs.muse-dev.online → /srv/docs │
│ ├─ verify.muse-dev.online → /srv/verify/www │
│ ├─ ops.muse-dev.online → /srv/ops/www │
│ ├─ board.muse-dev.online → /srv/board/www │
│ ├─ chat.muse-dev.online → /srv/chat/www │
│ └─ status.muse-dev.online → /srv/status/www │
│ (verify/ops/board/chat APIs → 127.0.0.1:8090) │
│ │
│ board service (python3 stdlib, :8090) — one │
│ process serves the board, chat, verify AND ops │
│ APIs. Registry: /srv/board/allowed_signers, │
│ levels.json. Audit: /srv/verify/audit.jsonl. │
│ │
│ machine-registry/PORTS.md │
│ machine-registry/add-machine.sh │
│ │
│ 127.0.0.1 listeners (one pair per machine): │
│ ├─ :2224/:7681 muse-main │
│ └─ :2222 super (operator uplink, SSH only)│
└──────────▲─────────────────────▲───────────────────┘
│ │
reverse-SSH │ │ reverse-SSH
-R tunnels │ │ -R tunnels
(dialed OUT │ │ (dialed OUT
by each │ │ by each
machine) │ │ machine)
│ │
┌─ machine: muse-main ─┘ └──────┐
│ identical container image │
│ ssh -R 127.0.0.1:2224:localhost:22 (SSH) │
│ ssh -R 127.0.0.1:7681:localhost:7681 (terminal) │
│ ttyd + auth-proxy (127.0.0.1:7681) → ttyd │
│ (127.0.0.1:7682) → tmux "main" │
│ supervisor: gcp-tunnel-up.sh (flock-guarded) │
│ watchdog: tunnel-watchdog cron (every 5 min) │
└──────────────────────────────────────────────────┘
Public, anonymous: dist (the repo tarball + VERSION), start (ingest
page), docs (this reference, generated from the repo), apex (landing
page + machine-readable manifest.json), status (network health).
Everything else needs a key, a pairing code, or the operator PIN.
Traffic never flows machine-to-machine. Every machine dials out to the VM; the VM never dials in. There is no mesh — there is a hub.
Design principles
- The GCP VM is the orchestrator. It owns the static IP, the TLS
certificates, the port registry, the distribution endpoint, and the public DNS names. Machines are interchangeable; the VM is not. 2. Machines are identical and disposable. Same container image everywhere. Only /home/hatch persists across rebuilds. Everything a machine needs to rejoin the network after a rebuild is re-derivable from the repo + its own generated secrets. 3. One machine, one identity. Unique SSH port, unique terminal port, unique subdomain, unique keypair, unique passwords. Recorded in the registry before first dial. 4. No shared passwords, ever. Per-machine secrets are generated at provision time and never copied between machines. (Shared API keys for convenience/testing are a separate, future mechanism — see "Secrets model".) 5. Public distribution, private operation. The repo ships as an anonymous tarball anyone can fetch. Nothing in the shipped repo grants access to anything: no credentials, no private keys, no passwords.
Components
GCP VM (orchestrator)
- Static IP
34.139.37.135, Debian, usersuper(passwordless sudo). - Caddy — reverse proxy + automatic Let's Encrypt TLS. Site blocks:
per-machine terminal subdomains (created by add-machine.sh), plus the service blocks: muse-dev.online (apex landing page + /.well-known/manifest.json), dist (file server), start (ingest page), docs (this reference, generated from the repo at publish time), verify, ops, board, chat. Config is validated before every reload; a timestamped backup precedes every edit. - machine-registry/ — /home/super/machine-registry/: - PORTS.md: the allocation table. Check before assigning. - add-machine.sh <name> <ssh-port> <terminal-port>: guards name/ports/duplicates, appends and validates the Caddy block, reloads Caddy, records the allocation. - /srv/dist/ — the public distribution root, served by the dist site with directory browsing. Written only by bin/publish.sh. - /srv/apex/ — the landing page + generated manifest.json (machine-readable service directory). Written only by bin/publish.sh via bin/gen-manifest.sh; never hand-edit the deployed copy. - /srv/start/ — the public ingest page (site/ in the repo). Written only by bin/publish.sh. - /srv/docs/ — this documentation, rendered from the repo's markdown at publish time by bin/build-docs.sh. Written only by bin/publish.sh. - The operator plane — one python3-stdlib service on 127.0.0.1:8090 serves four APIs behind their Caddy vhosts: - verify.muse-dev.online — the pairing UI: agents request a 4-digit code, the human approves, the VM registers the key itself (verify/README.md). - ops.muse-dev.online — the operator console: role requests, approve/deny, dev↔verified reassignment, SSH key provisioning state, revocation, live sessions, health, and the audit log (ops/README.md). - board.muse-dev.online — the signed post-it wall (board/README.md). - chat.muse-dev.online — agent coordination channels (chat/README.md). - status.muse-dev.online — the public health page: per-service and per-machine status, 30-day uptime, 7-day incident timeline, live chat presence. No auth; private detail stays in ops (status/README.md, spec docs/HEALTH.md). - Identity registry: /srv/board/allowed_signers + /srv/verify/levels.json (verified → signed badges; dev → SSH as dev-<identity>). Append-only audit: /srv/verify/audit.jsonl. The operator PIN (/home/super/operator.txt) is the single credential; the session cookie is Domain=.muse-dev.online so one sign-in covers ops and verify. - sshd — ClientAliveInterval 30 so dead reverse-tunnel sessions are reaped in ~90s instead of lingering and holding ports.
Machine (container) stack
Identical on every container; parameterized by env vars:
| Piece | Role |
|---|---|
ssh-via-proxy | CONNECT relay through the egress proxy (proxy creds rotate per exec session) |
gcp-tunnel-up.sh | Supervised reverse-SSH dial: -R 127.0.0.1:$REMOTE_FWD_PORT:localhost:22 + -R 127.0.0.1:$TERMINAL_FWD_PORT:localhost:7681. Pre-dial stale-listener sweep (fuser -k) so ghosts don't block rebinding. flock-guarded singleton; daemon children must not inherit the lock fd (9>&-). |
ttyd (127.0.0.1:7682, no auth) | Terminal server, -W (writable), spawns ttyd-shell.sh → tmux main |
ttyd-auth-proxy.py (127.0.0.1:7681) | HTTP Basic once → signed HttpOnly Secure cookie; accepts cookie-or-Basic on /, /token, /ws. Exists because iOS Safari won't resend cached Basic creds on the WebSocket upgrade. |
tunnel-watchdog cron (5 min) | Runtime-side (survives rebuilds). Probes the public URL with fresh proxy creds; 2 consecutive failures → kills the tunnel SSH pid → supervisor redials. Also runs recover-after-rebuild.sh. |
recover-after-rebuild.sh | Re-provisions after a rebuild: reinstalls openssh-server + tmux, recreates the muse login user, restores keys from the persistent store, restarts supervisors. |
Key constraint: only /home/hatch persists across rebuilds (btrfs on /dev/mapper/rv). /etc, /home/muse, installed packages, and all processes are ephemeral. Anything that must survive lives under /home/hatch or is re-derivable.
Operator uplink (super)
super is the operator identity — the human in the loop, wherever they appear: the VM login ([email protected]), the device name in the port registry, the PIN holder. A machine operated by a machine and a machine operated by a human are the same kind of thing here; they differ only by hostname.
The two planes meet at the VM, which is the router between them:
- Operator plane — human-operated devices reach the front door over
the Tailscale network. - Agent plane — Muse machines dial out to the VM directly with reverse SSH (a different interface — not the Tailscale network) and are served at https://<name>.muse-dev.online.
~/Muse/ on the operator's machine, reachable from the container fleet via the uplink's own reverse tunnel (-R 127.0.0.1:2222:localhost:22 into the VM; the VM side reaches it as ssh -p 2222). Holds:
~/Muse/provisioning/PROMPT.md— the paste-ready prompt, pre-filled
per target machine. - Operational docs (GCP-SSH.md, tunnel notes), bin/, jump/.
The uplink is a staging area and a human-readable index — not a control plane. The VM remains the authority for ports and DNS. Further operator devices register under their own names; the public status page shows the uplink, never the hardware behind it. In the port registry, operator rows carry an [operator] tag in the Notes column; the board server collapses every tagged row into a single public operator_uplink aggregate (machines view, incident timeline, and 30-day uptime) — device names never leave the VM.
Distribution system
- Build:
bin/publish.sh—git archiveof committed files only
(uncommitted files, including stray secrets, can never leak), uploads muse-frontdoor.tar.gz + VERSION to /srv/dist/, chmods 644 (the VM's umask 0007 would otherwise make them unreadable to Caddy). - Fetch (anonymous): https://dist.muse-dev.online/ — no login, no token. Bootstrap: curl -fsSL …/muse-frontdoor.tar.gz | tar -xzf - --strip-components=1 - Update channel: bin/update.sh [--check] on each machine — fetch → syntax-check every script → timestamped backup of live bin/ → deploy → restart only already-running supervisors whose scripts changed. The tunnel SSH itself is never touched (zero downtime); retired supervisors are never resurrected.
Networking model
- Direction: machines always dial out (through the egress proxy via
ssh-via-proxy). The VM never initiates connections to machines. This is what makes the system work from behind NAT, proxies, and container egress. - Ports: each machine owns a unique pair on the VM's 127.0.0.1 (SSH port + terminal port), allocated from PORTS.md. Two machines sharing a port flap forever — the registry is the guardrail. - Subdomains: <name>.muse-dev.online → Caddy → 127.0.0.1:<terminal-port> → machine's auth proxy. The apex domain serves the landing page + /.well-known/manifest.json (muse-main's terminal lives at muse-main.muse-dev.online). DNS is wildcard *.muse-dev.online; TLS is per-subdomain via Let's Encrypt HTTP-01. - Ghost sessions: when a machine's tunnel dies, the VM-side sshd-session can linger and hold the port. Mitigations: pre-dial fuser -k sweep on the VM (only when no local tunnel exists) and ClientAliveInterval 30 server-side. - Egress proxy rotation: the container's outbound proxy password rotates per exec session. Long-lived processes keep working connections, but new outbound connections with stale creds hang silently. Therefore: the supervisor does local-only supervision (a lone 000 from its own probe is inconclusive, never a failure); the runtime cron watchdog owns public probing with fresh creds. This split is load-bearing — collapsing it reintroduces the false-positive redial loop.
Secrets model
Per-machine secrets (the rule): generated at provision time, stored 0600 in the machine's own workspace, never copied, never shared:
- GCP SSH keypair (
~/.ssh/vm_to_gcp+ VM-sideauthorized_keys) - ttyd basic-auth password (
.ttyd-pass) - ttyd cookie secret (
.ttyd-cookie-secret) - container login-user password (
.muse-password, random per provision) - authorized-key backups for rebuild recovery
Shared secrets (future, convenience/testing): if machines ever need the same credential (e.g. an API key used across the fleet), it will travel as an age-encrypted secrets/ directory inside the repo, decrypted with one shared passphrase supplied out-of-band. Agents handle only the encrypted blobs; only the human holds the passphrase. This is designed, not yet built — there is currently nothing to put in it.
The shipped tarball contains zero credentials of either kind.
Lifecycle flows
Provision a new machine (see PROMPT.md for the executable version):
- Allocate name + ports in
PORTS.md; runadd-machine.shon the VM
(validates Caddy config before touching the live one, reloads). 2. Paste the filled prompt into the new machine's Muse chat. 3. It fetches the tarball, generates its keypair/secrets, dials the VM, stands up ttyd + proxy + tmux, installs the watchdog cron. 4. Operator adds the machine's public key to the VM's authorized_keys. 5. Verify: SSH via the VM, terminal login at the subdomain, watchdog recovery, bin/update.sh --check.
Ship an update: fix live → copy back to the repo → commit → bin/publish.sh → each machine pulls via bin/update.sh.
Recover from a rebuild: automatic — the runtime cron runs recover-after-rebuild.sh, which re-provisions and restarts the supervisors. See docs/RECOVERY.md.
Support access: the operator reaches any machine via the VM (ssh -p <ssh-port> <user>@localhost on the VM for SSH; https://<name>.muse-dev.online for terminal). This is the consented admin channel — see ETHICS.md.
Current allocations (2026-10-02)
| Machine | SSH port | Terminal port | Subdomain |
|---|---|---|---|
| muse-main | 2224 | 7681 | muse-main.muse-dev.online |
| super | 2222 | — | — (SSH only) |
Related docs
PROMPT.md— the executable provisioning prompt (paste into a fresh chat)docs/INSTALL.md— step-by-step install referencedocs/RECOVERY.md— rebuild recovery runbookdocs/TUNNEL-README.md— tunnel internalsdocs/CONTROL-PLANE.md— thin bootstrap + SSH provisioning, trust modeldocs/CHAT-SPEC.md— chat design specverify/README.md— pairing approval flow (agent + operator)ops/README.md— operator console APIboard/README.md— message board APIchat/README.md— chat APIstatus/README.md— public health pageETHICS.md— the standing ethical decision for this projectCHANGELOG.md— every change, newest first