Fleet Monitoring
Monitoring is layered: host metrics and container health from the CLI, long-term sampling for capacity trends, a beszel hub for at-a-glance history, and alerting through the control plane and failover watchers.
Beszel hub and agents
The fleet reports to a beszel hub with one agent per host. The hub and its four
agents are aligned to 0.20.0; Mon is registered as MON-FR-flo-primary with
disk and memory thresholds at 80%. The hub gives the fleet-wide time series
(CPU, memory, disk, network) that the CLI commands sample on demand.
flo vps stats — host and container view
flo vps stats # default VPS profile
flo vps stats production # named profile
flo --json vps stats # machine-readable
The command reports:
- Host: hostname, uptime, load average, CPU cores, memory (used/total/%), disk (used/total/%) with disk-pressure levels (warning ≥ 70%, critical ≥ 80%).
- Docker: running/total containers, down and unhealthy containers.
- Instances: containers grouped per tenant with live CPU/RAM caps read from the
daemon (
docker inspect), authoritative status/health, blue/green and infra containers. - Log volumes: per-tenant
/logssize, oldest roll, and rotation verdict (prod first, then test). This is the usual disk hog — a tenant with uncapped Serilog rolls is surfaced here before it fills the disk.
Container and instance health
flo instance health <id> # DB, API, disk, containers
flo instance stats <id> # CPU/memory/status (-w 5m, -i 15s; watch max 2h)
flo logs errors <id> --hours 6 # recent errors from blue, green, Traefik
flo logs days <id> # daily Serilog rolls
flo logs day <id> 2026-07-20 # one day's app log
Application-level health checks come from Docker Compose: the Flo app exposes
/health on internal port 10001 and PostgreSQL uses pg_isready. Nginx depends
on the app reaching healthy status but has no dedicated health check (see
Single-Tenant Deployment).
Long-term resource monitoring
flo instance monitor start <id> --interval 5m
flo instance monitor status
flo instance monitor report <id> --from 2026-05-10 --to 2026-05-16
flo instance monitor stop <id> --clean
instance monitor installs a cron job that samples container stats at a fixed
interval (60 s default), stores them on the host, and reports aggregates, hourly
trends, and threshold warnings over days or weeks. --all targets every instance
on the VPS.
flo doctor — system checks
flo doctor # local or --vps target
flo doctor --fix # provision missing local-dev toolchain
flo doctor --fix --dry-run
flo doctor cache-headers https://api.example.com
Checks run against the resolved target and exit non-zero when anything fails:
| Category | Checks |
|---|---|
| Docker | daemon running, Compose available, server version |
| Traefik | container running, traefik-public network exists |
| Security | master key present and 0600 (POSIX), SSL certificate expiry (warn < 30 days, fail if expired), OpenSSL available |
| Config | ~/.flo/config.json valid, data directory writable |
| System | disk space (fail < 1 GB, warn < 5 GB), memory, reboot pending |
| Tenants | registry valid, per-tenant active container healthy/starting/not running |
| Toolchain | local .NET/Node/Postgres image, PATH wiring for ~/.flo/toolchains (local only) |
--fix installs the missing local toolchain into ~/.flo/toolchains
(local-only; a VPS runs prebuilt containers). doctor cache-headers <base-url>
audits Cache-Control on flow-critical endpoints (admin configs, feature flags)
to catch stale-cache bugs after a deploy.
Journal and disk hygiene
flo vps journal status production
flo vps journal vacuum production --dry-run
flo vps journal vacuum production -y
flo vps cleanup production --dry-run
flo vps cleanup production --logs --older-than 7
flo strapi maintenance docker-cleanup -a
flo vps provisioninstalls a bounded systemd-journal policy:SystemMaxUse=500M,RuntimeMaxUse=200M,MaxRetentionSec=30dayvia a/etc/systemd/journald.conf.d/50-flo-journal.confdrop-in.journal vacuumrotates and vacuums to the fixed 500M/30-day policy (no arbitrary limits); the remote registry exposes the same fixed form.- The MON host additionally runs with a 200M journal cap and a weekly docker
prune;
flo control doctor <vps>reads the active journald lines (not commented defaults) for read-only triage. flo vps cleanupremoves unused Docker containers, images, volumes, and build cache, and with--logsprunes dated Serilog rolls from every tenant's log volume (never the live file).
Alerting
- Telegram (control plane): the daemon evaluates production thresholds every
minute and sends batched alerts with a per-issue cooldown (default 6 h) and
recovery messages. Configure with
flo control alerts mon --provision --token-stdin --chat <id> -y; verify with--test. See Control Plane. - Email (failover):
flo failover monitor --all --notify -yandflo failover watch --notifyemail on degraded replication or a suspected primary outage, using the notification config managed byflo config notify(SMTP or Cloudflare). - Backups:
flo backup statusgrades freshness and cron health; the control dashboard surfaces real backup age through read-only R2 credentials. - Auto-deploy watchdog:
flo watchdog statusandflo watchdog logsshow the CI-driven auto-deploy watchdog state.