Skip to main content

Control Plane (MON)

flo-control is the fleet daemon that holds all fleet SSH keys and answers the authenticated control-plane API. It runs on Mon, out-of-band from the tenant fleet, in a dedicated deployment bundle (cli/deploy/control/) that is separate from tenant images. Lease-bound tenants depend on MON for renewed write authority, so a MON outage can fence them — this is why the daemon is intentionally small and hardened.

Where it runs​

  • Dedicated bundle on Mon, mounted from ~/.flo-control, created by flo control bootstrap. It contains only safe VPS profiles and one control-plane key per target — never the operator's master.key, full config store, tenant secrets, or Cloudflare credentials.
  • It reuses the host's existing Traefik (flo-seo-traefik) over a shared Docker network; it publishes no host port. Every byte arrives through Traefik.
  • Container hardening: read-only rootfs, all Linux capabilities dropped, no-new-privileges, non-root user matching the bundle owner, CPU/memory limits, and restart: unless-stopped.

Two doors, one daemon​

DoorHostPath
Browser (Google gate)control.useflo.netTraefik → oauth2-proxy → daemon :8787
CLI / machineapi-control.useflo.netTraefik → daemon (opaque session token validated in-process)
  • The browser door forwards the Google OIDC ID token to the daemon; the daemon exchanges it at POST /auth/exchange for an opaque Flo session token. Google tokens are accepted only on that route.
  • stripSpoofableHeaders deletes inbound X-Forwarded-* / X-Auth-* so a bypass cannot forge identity; machine (lease) routes have their own rate limits.
  • Both doors enforce the operator allow-list (whitelist.txt): Google signature + aud + email_verified + whitelisted email. Workspace-only sign-in is an alternative mode (FLO_CONTROL_HOSTED_DOMAIN).

Sessions and tokens​

  • Operator sessions: minted by flo control login (Google sign-in in the browser) and stored locally at 0600. They expire 7 days after login, do not slide with activity, and logout revokes them immediately.
  • Read-only tokens: flo control login --read-only --token-file <path> mints a GET-only session for observing agents (chat bots, wall dashboards). The daemon refuses every non-GET route with 403 auth.read_only before reading a body; the only POST allowed is /auth/logout. Read-only tokens have a 30-day TTL with sliding renewal: while used, the expiry advances (a write at most once per half-life), so an active observer never expires, while an unused token still dies.
  • Useful read endpoints: GET /fleet/data.json (the dashboard snapshot, ~4 KB gzip), GET /api/history (24 h sparklines), GET /vps/:name/stats, GET /audit.

Dashboard and /fleet kiosk​

flo control dashboard start # local full-write dashboard on 127.0.0.1
flo control dashboard start --port 4173 --poll-seconds 10

The local dashboard proxies to the daemon with the CLI token held server-side.

https://control.useflo.net/fleet is a read-only kiosk view served by the daemon itself behind the Google gate: no write actions in the DOM, auto-reload when data goes stale or the session expires, dark token theme, PWA manifest, 30-day oauth2-proxy cookie for wall tablets. It shows:

  • per-host memory/disk sparklines from a SQLite ring buffer (metrics.db, sampled at snapshot build with a 60 s throttle);
  • DR posture per tenant (primary/standby, lease state, active/last incident) from dr.db;
  • backup freshness with a TTL-cached R2 probe;
  • version drift and correct "not running" states, with a persistent test-clone filter.

Telegram alerts​

flo control alerts mon # show channel + state
cat token.txt | flo control alerts mon --provision --token-stdin --chat <id> -y
flo control alerts mon --test --token-stdin < token.txt
flo control alerts mon --disable -y
  • The daemon evaluates production thresholds every minute: host unreachable, disk/memory, certificates, production containers down, instance not running, web, and backups — with stable issue keys for dedupe/cooldown.
  • One batch message per cycle (CRITICAL/WARNING), a recovery line when an issue clears, and a per-issue cooldown (default 6 h) so an active problem is not re-notified before it expires.
  • Multiple chats and multiple bots are supported (--chat CSV or --targets-stdin with a JSON array); one successful target is enough. State lives in metrics.db, so restarts do not lose it. Test clones are never alerted.
  • The bot token is read from stdin (never argv) and written to MON's .env.

Backup reader​

flo control backup-reader mon # show current state
flo control backup-reader mon --provision -y

Mints a bucket-scoped read-only R2 token, stores it, and injects FLO_S3_* into the MON bundle, then recreates the daemon. The dashboard probes each tenant's backup bucket once per 30 minutes and shows the real backup age. Without these credentials the Backups tab simply stays empty — never a false warning. Re-provision after creating a new bucket so the reader can see it (the remote backup schedule form itself runs without R2 write keys).

Audit (WORM on R2)​

Every write against the control plane produces a hash-chained, redacted audit line; the chain is mirrored off-box to R2 under a Bucket Lock rule (R2's object-lock equivalent), so history cannot be rewritten.

flo control audit --provision -y # create bucket + scoped token + lock
flo control audit --apply-lock -y # add the lock rule to an existing target
flo control audit # verify configuration

Defaults: bucket flo-control-audit, prefix audit, retention 3650 days (--retention-days). The scoped S3 credentials are encrypted in the operator's local flo config and are the only audit credentials sent to Mon. Operators can read their own records with GET /audit.

Remote command registry​

flo --vps <name> … never falls back to SSH: only registered forms are accepted, and they are executed by the daemon.

  • Read: vps stats|test, vps journal status, instance ls|test list, instance info|inspect|health|stats <id>, logs app|errors|db|days <id>, logs day <id> <date>, deploy history <id>, config env show/flags show, config ssl status, instance db tables <id>, backup status [<id>], watchdog status, failover status, secrets status <id>.
  • Write: deploy <id> --branch <branch> -y, deploy rollback <id> -y, instance start|stop|restart <id>, config env set <id> KEY=VALUE [--restart] -y, backup schedule <id> --bucket … --cron … --run-now -y, vps journal vacuum -y (fixed 500M/30-day policy), plus the finite DR lifecycle commands.
  • Every whitelisted account gets the same full write access to these actions. The registry does not expose arbitrary shell, SQL, file transfer, or local paths; use --local for an intentional local operation. Long DR jobs have a two-hour timeout.

Commands​

CommandDescriptionKey options
flo control loginBrowser Google sign-in; stores the session token--api, --app, --read-only, --token-file <path>
flo control logout / whoamiRevoke the local session / show identity and live sessions
flo control dashboard startLocal full-write dashboard proxying the daemon--port, --poll-seconds, --api, --no-open
flo control snapshotNon-secret fleet read-model snapshot (JSON)--out <file>, --check-backups
flo control buildRender the static read-only dashboard HTML from a snapshot--out, --from <file>, --check-backups
flo control serverRun the daemon (used inside the container)--port, --bind
flo control bootstrap [intermediary]Create the safe bundle + restricted per-target SSH keys--fleet <names> (required), --source, --bundle-home, --dry-run, -y
flo control deploy [vps]Cold-deploy MON with durable recovery--recover <id>, --repair-worker, --local-build, --no-build, --force-env, --force-whitelist, --enable/--disable-dr-automation, --maintenance, --shared-network, domains, -y
flo control doctor [vps]Read-only triage of the control host (disk, journal, transactions, daemon)
flo control auditConfigure/verify the WORM audit target--provision, --apply-lock, --retention-days, --replace, --bucket, --prefix, -y
flo control alerts [vps]Show/configure Telegram alerts--provision, --disable, --test, --chat, --token-stdin, --targets-stdin, --dir, -y
flo control backup-reader [vps]Show/provision the read-only R2 backup credentials--provision, --dry-run, --dir, -y
flo control agent-config <tenant>Prepare a private agent config; only the token hash reaches MON--node (required), --output (required), --init, --rotate-token, -y
flo control command status <job> / attach <job>Inspect or reattach to a remote command job still running on MON

Rollout and recovery​

flo vps test mon
flo control audit --provision -y
flo control bootstrap mon --fleet production,production-uk -y
flo control deploy mon --enable-dr-automation -y
  • Deploy gate: every deploy requires all MON tenants fully manual, no active incident, and no outstanding signed lease; the gate rejects unknown or unsafe state before stopping the daemon. --maintenance admits a live armed pair with idle DR state; --disable-dr-automation does not replace a full disarm.
  • --local-build compiles on the workstation for the target platform and transfers the verified image; --no-build requires the same engine version.
  • Interrupted deploy: keep the printed operation ID and run flo control deploy mon --recover <operation-id> -y. Recovery verifies the new container or restores the owned snapshot; never use docker compose down or delete operation files.
  • MON IP change: flo control bootstrap mon --fleet … --source <new-ip> -y then flo control deploy mon --no-build -y (see the bundle README in cli/deploy/control/README.md).
  • Before every remote VPS mutation run the read-only preflight: flo --vps <name> vps stats.

See Failover and Disaster Recovery for the DR lifecycle that MON coordinates, and Backups for the freshness data the dashboard reads.