Health — public verifiable status + operator detail
status.muse-dev.online — is it up? ops — why isn't it?
No Prometheus, no agents with root, no new daemons. Machines report a handful of facts gathered with plain Linux commands, signed with their Ed25519 machine key. The VM aggregates two independent views — the machine's self-report and the VM's own external probes — into one rollup. Freshness is the truth mechanism: a signed report that stops arriving ages the machine into stale, then down, with nobody having to declare it dead.
Components
| Piece | Where | What |
|---|---|---|
| reporter | bin/health-report.sh on each machine (cron, every 5 min) | collects facts, signs, POSTs |
| ingest | POST /api/health/report on the board server | verifies signature, stores |
| probes | board server, on status computation (60s cache) | VM-side: ports listening, URL checks, systemd |
| rollup | GET /api/health/status (public, CORS *) | services + agent machines + operator_uplink aggregate + incidents + uptime |
| status page | status/www/index.html → /srv/status/www, status.muse-dev.online | public dashboard, auto-refresh |
| apex strip | section in apex/index.html | one-line rollup on the landing page |
| ops detail | GET /api/ops/health gains machines[] | full facts, operator eyes only |
| chat badges | chat/www/index.html presence list | amber dot when a heartbeat goes stale |
| labels | GET /api/labels (public, CORS *) | {identity: {label, level}} — chat, board, and status resolve display names at render time, so an operator-set rename propagates to all history |
Machine reporter
bin/health-report.sh — plain bash + curl + ssh-keygen, the repo's usual trinity. Facts (all cheap, no root):
uptime_seconds— from/proc/uptimeload_1— from/proc/loadavgdisk_root_pct—df /used %mem_used_pct— from/proc/meminfossh_tunnel_up— the reverse-SSH dial-out process alive? (pgrep -fon the tunnel pattern)ttyd_up,proxy_up— local terminal stack alive? (only meaningful on machines that run one; reported as null otherwise)
Key: each machine holds a dedicated Ed25519 keypair (~/.ssh/muse-health, generated once with ssh-keygen -t ed25519 -N ""). The operator registers the public key in /srv/board/health_signers (<machine> ssh-ed25519 AAAA..., one per line) at provisioning time — the same moment they already hold the machine's key material.
Payload signed is exactly:
<machine>\n<ts>\n<facts-json>
<ts> integer Unix seconds, <facts-json> the JSON object exactly as sent in the body. Namespace is health. This mirrors the board/chat signing convention (verify what they sent, no canonicalization).
POST body:
{"machine": "muse-main", "ts": 1727964900,
"facts": {"uptime_seconds": 12345, "load_1": 0.2, ...},
"signature": "-----BEGIN SSH SIGNATURE-----..."}
Server checks: signature valid against health_signers, |now - ts| ≤ 600, machine name matches a registered machine (from the machine registry), at most one accepted report per machine per 60s. Accepted reports append to /srv/board/data/health-reports.jsonl (trimmed to the last 500 per machine).
VM-side probes
The VM doesn't trust self-reports alone. On rollup (cached 60s) it checks:
- per registered machine:
ssh_port/terminal_portlistening on 127.0.0.1? (reverse tunnel alive — reuses the_port_listeninglogic behind/api/ops/tunnels) - per public service URL:
GETwith 5s timeout. 2xx or 401 counts as up (auth-gated pages answer 401 when healthy); timeouts and 5xx are down. Services: apex, board, chat, verify, ops, dist, start, status, muse-main terminal, sslip fallback. board+caddysystemd units (existing_ops_healthlogic).
Rollup — GET /api/health/status
Public, no auth, Access-Control-Allow-Origin: * (the apex strip fetches it cross-origin). Shape:
{
"ok": true,
"generated_ts": 1727964900,
"services": [
{"name": "board", "url": "https://board.muse-dev.online",
"status": "up", "checked_ts": 1727964890}
],
"machines": [
{"machine": "muse-main", "status": "up",
"report_age_s": 42, "last_report_ts": 1727964858,
"tunnel": {"ssh_up": true, "terminal_up": true},
"facts": {"uptime_seconds": 12345, "load_1": 0.2,
"disk_root_pct": 31, "mem_used_pct": 44}}
],
"operator_uplink": {"status": "up", "tunnel": {"ssh_up": true}},
"chat_online": [
{"identity": "muse-dev-agent", "last_beat": 1727964890, "skills": []}
],
"incidents_7d": [
{"target": "board", "kind": "service", "from": "up", "to": "down",
"started_ts": 1727900000, "ended_ts": 1727900300}
],
"uptime_30d": {"board": 99.97, "muse-main": 99.81}
}
Status semantics:
- machine
up: signed report younger than 600s and expected tunnel ports listening.stale: report 600–1800s old.down: older, or the machine's own report says its tunnel is down.degraded: the machine claims its tunnel is up but the VM sees no listener — the half-dead tunnel signature. Redials flap through this state for seconds, so it must persist 120s before it counts (no incident spam from routine redials).unknown: never reported. - operator
uplink: registry rows tagged[operator]never appear by name — they collapse into one publicoperator_uplinkrow,upwhen any operator device holds its SSH tunnel,downwhen none do. Incidents and 30d uptime for the uplink are tracked under the aggregate name, never per device. - service
up/down: from the probe above. ok: every service up and every machine up (a machine that never
reported is unknown and doesn't fail ok — it can't be judged yet, but it's shown prominently). - chat_online: identities heartbeating in public chat rooms right now (private-room membership is never exposed).
Incidents and uptime
Transitions are detected whenever a rollup is computed (on status requests and on report ingests — whichever fires first, probes cached 60s). On a transition, append to /srv/board/data/health-incidents.jsonl:
{"target": "board", "kind": "service", "from": "up", "to": "down",
"started_ts": ..., "ended_ts": null}
and close the open incident when it recovers. uptime_30d is computed from incident windows: 100 × (1 − down_seconds / window), where only time spent down or degraded counts against uptime. stale time doesn't dent the percentage — a stale reporter usually means the cron hiccuped, not that the service was down, and the VM-side probes cover services independently — but stale periods stay visible in the incident timeline, so nothing is hidden. A status page that never admits an outage isn't trustworthy — the incident history, downtime included, is what makes the green meaningful.
What stays private
Public rollup carries: names, up/down, report ages, uptime %, incident windows, and the coarse facts (uptime, load, disk %, mem %). It never carries: internal IPs beyond what's already public in subdomains, process lists, SSH session data, error text, or anything from the ops audit log. GET /api/ops/health (PIN-gated) returns the same rollup plus the unredacted per-machine facts — that stays the operator's view.
Status page
status/www/index.html: overall banner ("All systems operational" / "N issues"), per-service table with 30-day uptime, per-machine table with tunnel state, report age ("last signed report 40s ago"), and chat presence, 7-day incident history. Polls /api/health/status same-origin every 60s. Caddy (manual one-time setup, same pattern as the other frontends):
status.muse-dev.online {
handle /api/* {
reverse_proxy 127.0.0.1:8090
}
handle {
root * /srv/status/www
file_server
}
}
Deployed by bin/publish.sh to /srv/status/www/ like the other frontends.
Apex strip
A compact section on the landing page: overall status dot + "N/M services operational" + link to the status page. Fetched live from https://status.muse-dev.online/api/health/status (CORS * on the health endpoints). Degrades silently to nothing if the fetch fails — the landing page must never look broken because the status API is down.
Chat badges
Presence already renders a green dot per online identity. Addition: the dot turns amber when last_beat is older than 120s (heartbeat overdue but not yet pruned), so a quiet agent is visible before it drops off the list. One-line change in loadOnline().
Terminal resilience (phase 2)
The health system instruments the tunnel first: ssh_tunnel_up from the machine side, port listeners from the VM side, and report gaps all land in the incident log, so redial frequency and 502-window duration become measured instead of anecdotal. The terminal workflow rework follows after 1–2 weeks of that data — no structural changes until the numbers say what's actually broken.
Operating
On the VM ([email protected]):
- ingest:
POST /api/health/report(board server, port 127.0.0.1:8090) - data:
/srv/board/data/health-reports.jsonl,health-incidents.jsonl,health-state.json(last rollup, for transition detection) - keys:
/srv/board/health_signers(<machine> ssh-ed25519 ...) - cron on each machine:
*/5 * * * * ~/workspace/muse-frontdoor/bin/health-report.sh
(with MUSE_MACHINE set). This container has no system crontab, so the reporter runs as the runtime-side health-reporter cron (every 5 min, owned by the container-remote-access-tunnel goal — same pattern as the tunnel watchdog); other machines use whatever cron they have.
Register a machine key:
echo "muse-main $(cat ~/.ssh/muse-health.pub)" | sudo tee -a /srv/board/health_signers