Muse Front Door — docs

INSTALL.md — R&D container networking setup (agent-facing)

Scope: replicate this container's SSH and terminal R&D stack on another, identical container. The container image is the same everywhere; only the per-install identity (keys, ports, secrets) changes.

The architecture, in one picture:


  ┌─ identical container ──────────────────────────────┐

  │  ssh-via-proxy (CONNECT relay through egress proxy) │

  │  gcp-tunnel-up.sh ── ssh -R 127.0.0.1:<PORT>:localhost:22

  │  ttyd 1.7.7 + ttyd-auth-proxy.py + tmux ("main")    │

  │  tunnel-watchdog cron (every 5 min, runtime-side)   │

  └─────────────────────────┬──────────────────────────┘

                            │  dials OUT through the

                            │  egress proxy (rotating creds)

                            ▼

                    shared GCP jump host (34.139.37.135)

                    user reaches the container via GCP

One rule that governs everything below: the GCP VM is the shared front door. New containers dial INTO it. There are no per-container localhost.run browser tunnels — that path exists only on the original container and is legacy. The GCP side will reverse-proxy browser traffic too; the container-side browser stack (ttyd, auth proxy, tmux) is the same regardless of which public path fronts it.


1. Before you start (prerequisites)

You need from the human (or from the GCP owner):

  1. GCP VM address and SSH user — currently 34.139.37.135, user super.
  2. Unique ports on the GCP VM for THIS container — check the registry

first: /home/super/machine-registry/PORTS.md on the GCP VM lists every taken port. You need one SSH port and one terminal port, both free (this container owns 2224/7681; the operator uplink super owns 2222). Two machines sharing a port will flap forever ("remote port forwarding failed"). 3. A fresh SSH keypair generated inside this container (ssh-keygen -t ed25519 -f ~/.ssh/vm_to_gcp -N ""). The PUBLIC key must be added to the GCP VM's super user authorized_keys. The private key never leaves the container. 4. Egress proxy env vars present (HTTPS_PROXY/https_proxy pointing at hatch-egress-proxy:3128). Verify: env | grep -i proxy. Without these, nothing below works. 5. A ttyd login password you generate yourself and hand to the user (user muse). Generate with openssl rand -hex 16.

Also confirm the container basics: /home/hatch is the only directory that survives a rebuild (/dev/mapper/rv, btrfs). /home/muse, /etc, apt packages, and all processes are wiped without warning. Everything durable must live under /home/hatch.

2. Files to bring in (the portable core)

Copy these from the reference container or the distribution bundle. They contain no secrets and are identical across installs:

FileRole
bin/ssh-via-proxyThe SSH engine: stdio bridge for ssh -o ProxyCommand, carries SSH inside an HTTP CONNECT tunnel through the egress proxy. Pure Python, no deps.
bin/gcp-tunnel-up.shSupervisor for the GCP reverse-SSH tunnel. flock-guarded singleton; does a stale-listener sweep on the GCP host before dialing; logs to workspace/tunnel/gcp-tunnel.log. Per-machine config at the top: REMOTE_FWD_PORT (SSH) and TERMINAL_FWD_PORT (terminal) — set both from the registry before first run.
bin/recover-after-rebuild.shIdempotent rebuild recovery: fixes apt mirror, installs openssh-client/server + tmux, recreates the login user, restores authorized_keys from the persistent backup, starts supervisors.
bin/fix-apt-mirror.shDrops the dead second Ubuntu mirror that makes apt-get update hang forever. Re-apply after every rebuild (/etc is ephemeral).
bin/ttydttyd 1.7.7 binary (browser terminal server).
bin/ttyd-auth-proxy.pyHTTP Basic → signed session-cookie proxy in front of ttyd. Required because iOS Safari will not resend cached Basic credentials on the WebSocket /ws upgrade. Listens on 127.0.0.1:7681; ttyd itself runs with NO auth on 127.0.0.1:7682 (localhost-only).
bin/ttyd-shell.shPer-connection shell wrapper: env TERM=xterm-256color tmux new-session -A -s main. Keep the volatile command here so ttyd restarts don't need supervisor restarts.
bin/tunnel-up.shLegacy browser-tunnel supervisor (localhost.run). Only the original container uses it. Do NOT enable on new installs.

Per-install items you must NOT copy: ~/.ssh/*, workspace/.ttyd-pass, workspace/.ttyd-cookie-secret, workspace/tunnel/muse-authorized_keys (the login key backup is per-install).

3. Provision the container

Run bin/recover-after-rebuild.sh. It is idempotent and safe to re-run. What it does on a fresh container:

  1. Fixes the apt mirror (fix-apt-mirror.sh).
  2. Installs openssh-client, openssh-server, tmux.
  3. Creates the login user (currently muse) — remember /home/muse

vanishes on rebuild, so the user is recreated here every time. 4. Restores authorized_keys from the persistent backup (workspace/.ssh-keys/ — keep your backups here; the script restores both the login user and root keys). 5. Sets PermitRootLogin prohibit-password in sshd_config (wiped on rebuild) and starts sshd on port 22. 6. Starts the tunnel supervisor(s) via flock-guarded ensure functions.

After provisioning, verify: ss -tln | grep :22 shows sshd listening.

4. The SSH engine — how dialing out works

Every outbound SSH connection goes through ssh-via-proxy:


ssh -o ProxyCommand="$HOME/workspace/bin/ssh-via-proxy %h %p" \

    -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null \

    -o ServerAliveInterval=25 -o ServerAliveCountMax=3 \

    -o ConnectTimeout=30 -o BatchMode=yes \

    -N -R 127.0.0.1:<YOUR-PORT>:localhost:22 \

    [email protected]

Key facts:

SSH connection keeps working (its CONNECT is already established), but any NEW dial must use the current session's env. This is why supervision lives in short-lived processes and crons, not in one eternal daemon. - The proxy kills long-lived CONNECT tunnels roughly every 10–15 minutes ("Broken pipe"). Expect redials; the supervisor handles them. - When our side dies, the remote sshd session can linger and keep the listen port bound, so redials fail with "remote port forwarding failed". gcp-tunnel-up.sh runs a pre-dial stale-listener sweep (fuser -k <port>/tcp on the GCP host, only when no local tunnel exists), and the GCP sshd has ClientAliveInterval 30 so ghosts are reaped in ~90s. If you see repeated "forwarding failed" errors, check the GCP side for lingering sshd-session processes.

5. Start the GCP supervisor

bin/gcp-tunnel-up.sh (edit REMOTE_FWD_PORT, TERMINAL_FWD_PORT and SSH_KEY for this install first — take the ports from the registry, never invent them). Run it once in the background; it daemonizes and holds a flock on workspace/tunnel/gcp-tunnel.lock so double-starts exit cleanly. Verify: pgrep -f 'workspace/bin/gcp-tunnel-up\.sh$' and check workspace/tunnel/gcp-tunnel.log for "tunnel established".

The user reaches the container via the GCP VM: ssh -p <YOUR-PORT> <login-user>@localhost from the GCP VM (or ssh -J super@<gcp-ip> -p <YOUR-PORT> <login-user>@localhost from anywhere).

6. Browser terminal stack (same on every install)

  1. Internal ttyd: ttyd -p 7682 -i 127.0.0.1 -W bin/ttyd-shell.sh
  2. -W is REQUIRED — without it the terminal is read-only and the page

still loads fine, so the failure is silent. - ttyd spawns its child command only after the client sends {"AuthToken": base64(user:pass), "columns": N, "rows": M} on the tty websocket subprotocol. An empty tmux ls with no client connected is EXPECTED, not a bug. - The shell wrapper exports TERM=xterm-256color; without it tmux dies with "open terminal failed". 2. Auth proxy on 127.0.0.1:7681 (ttyd-auth-proxy.py). Requires Basic once, issues a signed HttpOnly Secure session cookie, accepts cookie-or-Basic on /, /token, /ws. Cookie secret must be generated per install — never copy it. 3. tmux new-session -A -s main gives the persistent session that survives reloads and redials. 4. Generate workspace/.ttyd-pass (0600) for the muse login user.

7. Watchdog cron (runtime-side, survives rebuilds)

VM-local processes die on rebuilds; the cron does not. Create a tunnel-watchdog cron, interval every 5 minutes, with the body from workspace/tunnel-duplicate/tunnel-watchdog.cron.md (in the bundle). Its job:

  1. Run recover-after-rebuild.sh (idempotent — restarts supervisors

after a rebuild). 2. Read the current public URL from workspace/tunnel/URL.txt. 3. Probe it with a fresh environment (fresh proxy creds every run). HTTP 200 or 401 = healthy. 4. On 2 consecutive failures (~10 min): kill the tunnel's ssh PID(s) individually (exact PIDs via pgrep — NEVER a broad pkill), let the supervisor redial, reset the counter in workspace/tunnel/health.state.

Health-check split is deliberate: the supervisor does LOCAL supervision only (restart dead ttyd/proxy, redial when ssh exits). It never probes the public URL because its own env goes stale and its curls fail with 000 — false-positive redials, each minting a new public URL.

8. Verification checklist

client with AuthToken JSON gets a shell (don't rely on curl). - [ ] tunnel-watchdog cron exists and its runs succeed. - [ ] Kill the ssh tunnel manually; supervisor redials within ~30s. - [ ] Simulate a rebuild mentally: everything you need is under /home/hatch; nothing in /etc, /home/muse, or process state.

9. Hard-won lessons (do not relearn these)

in a subshell around check+start, and close the fd (9>&-) on the daemon's command line — otherwise the lock is held forever and every future run skips. - Anchor pgrep patterns (workspace/bin/tunnel-up\.sh$); unanchored patterns match sibling scripts (e.g. gcp-tunnel-up.sh). - Never broad-pkill in this environment. pkill -f "pattern" can match your own shell's command line and SIGTERM your session. Capture exact PIDs with pgrep (excluding $$) and kill those. - Only /home/hatch persists. Verified: everything else is on the ephemeral root overlay. Rebuilds happen without warning. - Apt mirror: stock ubuntu.sources lists a dead mirror that hangs apt-get update with no output. fix-apt-mirror.sh after every rebuild. - Container root cannot modify files owned by nobody (user-namespace quirk, EACCES). Delete and recreate instead of editing in place. - Dead ends, don't retry: Tailscale (proxy 400s the Noise handshake), cloudflared (ignores proxy env; API 403s quick tunnels from this egress IP), tmate (proxy blocks the whole tmate.io domain at destination level). The user-controlled GCP jump host is the viable path.

10. What stays per-install (never copy between containers)

~/.ssh/* (the GCP keypair), workspace/.ttyd-pass, workspace/.ttyd-cookie-secret, login-user authorized_keys backups, the chosen REMOTE_FWD_PORT / TERMINAL_FWD_PORT, and the ttyd login password. Everything else in bin/ and this doc is identical across installs.

11. Onboarding a new machine (GCP VM side)

The VM keeps a port registry and a helper script at /home/super/machine-registry/:

subdomain. Read it before assigning ports. - add-machine.sh <name> <ssh-port> <terminal-port> — guards against bad names, taken ports, and duplicates; appends <name>.muse-dev.online → 127.0.0.1:<terminal-port> to the Caddyfile (validated before the live file is touched); reloads Caddy; records the allocation. Run with sudo (passwordless sudo is configured).

After the VM side is registered, the new machine's container dials in with its gcp-tunnel-up.sh set to the assigned ports — its terminal then appears at https://<name>.muse-dev.online with automatic HTTPS.