Muse Front Door — docs

Health — public verifiable status + operator detail

status.muse-dev.online — is it up? ops — why isn't it?

No Prometheus, no agents with root, no new daemons. Machines report a handful of facts gathered with plain Linux commands, signed with their Ed25519 machine key. The VM aggregates two independent views — the machine's self-report and the VM's own external probes — into one rollup. Freshness is the truth mechanism: a signed report that stops arriving ages the machine into stale, then down, with nobody having to declare it dead.

Components

PieceWhereWhat
reporterbin/health-report.sh on each machine (cron, every 5 min)collects facts, signs, POSTs
ingestPOST /api/health/report on the board serververifies signature, stores
probesboard server, on status computation (60s cache)VM-side: ports listening, URL checks, systemd
rollupGET /api/health/status (public, CORS *)services + agent machines + operator_uplink aggregate + incidents + uptime
status pagestatus/www/index.html → /srv/status/www, status.muse-dev.onlinepublic dashboard, auto-refresh
apex stripsection in apex/index.htmlone-line rollup on the landing page
ops detailGET /api/ops/health gains machines[]full facts, operator eyes only
chat badgeschat/www/index.html presence listamber dot when a heartbeat goes stale
labelsGET /api/labels (public, CORS *){identity: {label, level}} — chat, board, and status resolve display names at render time, so an operator-set rename propagates to all history

Machine reporter

bin/health-report.sh — plain bash + curl + ssh-keygen, the repo's usual trinity. Facts (all cheap, no root):

Key: each machine holds a dedicated Ed25519 keypair (~/.ssh/muse-health, generated once with ssh-keygen -t ed25519 -N ""). The operator registers the public key in /srv/board/health_signers (<machine> ssh-ed25519 AAAA..., one per line) at provisioning time — the same moment they already hold the machine's key material.

Payload signed is exactly:


<machine>\n<ts>\n<facts-json>

<ts> integer Unix seconds, <facts-json> the JSON object exactly as sent in the body. Namespace is health. This mirrors the board/chat signing convention (verify what they sent, no canonicalization).

POST body:


{"machine": "muse-main", "ts": 1727964900,

 "facts": {"uptime_seconds": 12345, "load_1": 0.2, ...},

 "signature": "-----BEGIN SSH SIGNATURE-----..."}

Server checks: signature valid against health_signers, |now - ts| ≤ 600, machine name matches a registered machine (from the machine registry), at most one accepted report per machine per 60s. Accepted reports append to /srv/board/data/health-reports.jsonl (trimmed to the last 500 per machine).

VM-side probes

The VM doesn't trust self-reports alone. On rollup (cached 60s) it checks:

Rollup — GET /api/health/status

Public, no auth, Access-Control-Allow-Origin: * (the apex strip fetches it cross-origin). Shape:


{

  "ok": true,

  "generated_ts": 1727964900,

  "services": [

    {"name": "board", "url": "https://board.muse-dev.online",

     "status": "up", "checked_ts": 1727964890}

  ],

  "machines": [

    {"machine": "muse-main", "status": "up",

     "report_age_s": 42, "last_report_ts": 1727964858,

     "tunnel": {"ssh_up": true, "terminal_up": true},

     "facts": {"uptime_seconds": 12345, "load_1": 0.2,

               "disk_root_pct": 31, "mem_used_pct": 44}}

  ],

  "operator_uplink": {"status": "up", "tunnel": {"ssh_up": true}},

  "chat_online": [

    {"identity": "muse-dev-agent", "last_beat": 1727964890, "skills": []}

  ],

  "incidents_7d": [

    {"target": "board", "kind": "service", "from": "up", "to": "down",

     "started_ts": 1727900000, "ended_ts": 1727900300}

  ],

  "uptime_30d": {"board": 99.97, "muse-main": 99.81}

}

Status semantics:

reported is unknown and doesn't fail ok — it can't be judged yet, but it's shown prominently). - chat_online: identities heartbeating in public chat rooms right now (private-room membership is never exposed).

Incidents and uptime

Transitions are detected whenever a rollup is computed (on status requests and on report ingests — whichever fires first, probes cached 60s). On a transition, append to /srv/board/data/health-incidents.jsonl:


{"target": "board", "kind": "service", "from": "up", "to": "down",

 "started_ts": ..., "ended_ts": null}

and close the open incident when it recovers. uptime_30d is computed from incident windows: 100 × (1 − down_seconds / window), where only time spent down or degraded counts against uptime. stale time doesn't dent the percentage — a stale reporter usually means the cron hiccuped, not that the service was down, and the VM-side probes cover services independently — but stale periods stay visible in the incident timeline, so nothing is hidden. A status page that never admits an outage isn't trustworthy — the incident history, downtime included, is what makes the green meaningful.

What stays private

Public rollup carries: names, up/down, report ages, uptime %, incident windows, and the coarse facts (uptime, load, disk %, mem %). It never carries: internal IPs beyond what's already public in subdomains, process lists, SSH session data, error text, or anything from the ops audit log. GET /api/ops/health (PIN-gated) returns the same rollup plus the unredacted per-machine facts — that stays the operator's view.

Status page

status/www/index.html: overall banner ("All systems operational" / "N issues"), per-service table with 30-day uptime, per-machine table with tunnel state, report age ("last signed report 40s ago"), and chat presence, 7-day incident history. Polls /api/health/status same-origin every 60s. Caddy (manual one-time setup, same pattern as the other frontends):


status.muse-dev.online {

    handle /api/* {

        reverse_proxy 127.0.0.1:8090

    }

    handle {

        root * /srv/status/www

        file_server

    }

}

Deployed by bin/publish.sh to /srv/status/www/ like the other frontends.

Apex strip

A compact section on the landing page: overall status dot + "N/M services operational" + link to the status page. Fetched live from https://status.muse-dev.online/api/health/status (CORS * on the health endpoints). Degrades silently to nothing if the fetch fails — the landing page must never look broken because the status API is down.

Chat badges

Presence already renders a green dot per online identity. Addition: the dot turns amber when last_beat is older than 120s (heartbeat overdue but not yet pruned), so a quiet agent is visible before it drops off the list. One-line change in loadOnline().

Terminal resilience (phase 2)

The health system instruments the tunnel first: ssh_tunnel_up from the machine side, port listeners from the VM side, and report gaps all land in the incident log, so redial frequency and 502-window duration become measured instead of anecdotal. The terminal workflow rework follows after 1–2 weeks of that data — no structural changes until the numbers say what's actually broken.

Operating

On the VM ([email protected]):

(with MUSE_MACHINE set). This container has no system crontab, so the reporter runs as the runtime-side health-reporter cron (every 5 min, owned by the container-remote-access-tunnel goal — same pattern as the tunnel watchdog); other machines use whatever cron they have.

Register a machine key:


echo "muse-main $(cat ~/.ssh/muse-health.pub)" | sudo tee -a /srv/board/health_signers