Skip to main content

Fleet Monitoring

Monitoring is layered: host metrics and container health from the CLI, long-term sampling for capacity trends, a beszel hub for at-a-glance history, and alerting through the control plane and failover watchers.

Beszel hub and agents​

The fleet reports to a beszel hub with one agent per host. The hub and its four agents are aligned to 0.20.0; Mon is registered as MON-FR-flo-primary with disk and memory thresholds at 80%. The hub gives the fleet-wide time series (CPU, memory, disk, network) that the CLI commands sample on demand.

flo vps stats — host and container view​

flo vps stats # default VPS profile
flo vps stats production # named profile
flo --json vps stats # machine-readable

The command reports:

  • Host: hostname, uptime, load average, CPU cores, memory (used/total/%), disk (used/total/%) with disk-pressure levels (warning ≥ 70%, critical ≥ 80%).
  • Docker: running/total containers, down and unhealthy containers.
  • Instances: containers grouped per tenant with live CPU/RAM caps read from the daemon (docker inspect), authoritative status/health, blue/green and infra containers.
  • Log volumes: per-tenant /logs size, oldest roll, and rotation verdict (prod first, then test). This is the usual disk hog — a tenant with uncapped Serilog rolls is surfaced here before it fills the disk.

Container and instance health​

flo instance health <id> # DB, API, disk, containers
flo instance stats <id> # CPU/memory/status (-w 5m, -i 15s; watch max 2h)
flo logs errors <id> --hours 6 # recent errors from blue, green, Traefik
flo logs days <id> # daily Serilog rolls
flo logs day <id> 2026-07-20 # one day's app log

Application-level health checks come from Docker Compose: the Flo app exposes /health on internal port 10001 and PostgreSQL uses pg_isready. Nginx depends on the app reaching healthy status but has no dedicated health check (see Single-Tenant Deployment).

Long-term resource monitoring​

flo instance monitor start <id> --interval 5m
flo instance monitor status
flo instance monitor report <id> --from 2026-05-10 --to 2026-05-16
flo instance monitor stop <id> --clean

instance monitor installs a cron job that samples container stats at a fixed interval (60 s default), stores them on the host, and reports aggregates, hourly trends, and threshold warnings over days or weeks. --all targets every instance on the VPS.

flo doctor — system checks​

flo doctor # local or --vps target
flo doctor --fix # provision missing local-dev toolchain
flo doctor --fix --dry-run
flo doctor cache-headers https://api.example.com

Checks run against the resolved target and exit non-zero when anything fails:

CategoryChecks
Dockerdaemon running, Compose available, server version
Traefikcontainer running, traefik-public network exists
Securitymaster key present and 0600 (POSIX), SSL certificate expiry (warn < 30 days, fail if expired), OpenSSL available
Config~/.flo/config.json valid, data directory writable
Systemdisk space (fail < 1 GB, warn < 5 GB), memory, reboot pending
Tenantsregistry valid, per-tenant active container healthy/starting/not running
Toolchainlocal .NET/Node/Postgres image, PATH wiring for ~/.flo/toolchains (local only)

--fix installs the missing local toolchain into ~/.flo/toolchains (local-only; a VPS runs prebuilt containers). doctor cache-headers <base-url> audits Cache-Control on flow-critical endpoints (admin configs, feature flags) to catch stale-cache bugs after a deploy.

Journal and disk hygiene​

flo vps journal status production
flo vps journal vacuum production --dry-run
flo vps journal vacuum production -y
flo vps cleanup production --dry-run
flo vps cleanup production --logs --older-than 7
flo strapi maintenance docker-cleanup -a
  • flo vps provision installs a bounded systemd-journal policy: SystemMaxUse=500M, RuntimeMaxUse=200M, MaxRetentionSec=30day via a /etc/systemd/journald.conf.d/50-flo-journal.conf drop-in. journal vacuum rotates and vacuums to the fixed 500M/30-day policy (no arbitrary limits); the remote registry exposes the same fixed form.
  • The MON host additionally runs with a 200M journal cap and a weekly docker prune; flo control doctor <vps> reads the active journald lines (not commented defaults) for read-only triage.
  • flo vps cleanup removes unused Docker containers, images, volumes, and build cache, and with --logs prunes dated Serilog rolls from every tenant's log volume (never the live file).

Alerting​

  • Telegram (control plane): the daemon evaluates production thresholds every minute and sends batched alerts with a per-issue cooldown (default 6 h) and recovery messages. Configure with flo control alerts mon --provision --token-stdin --chat <id> -y; verify with --test. See Control Plane.
  • Email (failover): flo failover monitor --all --notify -y and flo failover watch --notify email on degraded replication or a suspected primary outage, using the notification config managed by flo config notify (SMTP or Cloudflare).
  • Backups: flo backup status grades freshness and cron health; the control dashboard surfaces real backup age through read-only R2 credentials.
  • Auto-deploy watchdog: flo watchdog status and flo watchdog logs show the CI-driven auto-deploy watchdog state.