Skip to main content

Failover and Disaster Recovery

Flo protects each tenant with a primary + standby VPS pair. PostgreSQL streams WAL from the primary to the standby over a private WireGuard mesh, and the MON control plane is the single authority that can grant a node write authority. Failover can be manual (operator command) or automatic (MON coordinator), but it is always fail-closed: a node without a valid lease stops serving.

The full operator runbook lives in docs/DisasterRecoveryDrill.md in the platform repository; this page is the technical reference.

Architecture​

MON (control plane, out-of-band)
│ durable DR store: epochs, leases, incidents, policy
│ signed short-lived leases
▼
┌──────────────────┐ WireGuard mesh (wg0, 10.88.0.0/24, UDP 51820) ┌──────────────────┐
│ primary VPS │◄─────────── PostgreSQL streaming replication ────►│ standby VPS (UK) │
│ app (active) │ │ app (stopped) │
│ postgres primary │ │ postgres replica │
│ failover agent │ │ failover agent │
└──────────────────┘ └──────────────────┘
▲ ▲
└────────────────── tenant traffic (DNS → Traefik) ──────────────────────┘
  • WireGuard mesh: flo vps link <primary> <standby> builds a two-node private network (wg0, RFC1918 10.88.0.0/24, default UDP port 51820). Only public keys and mesh IPs are recorded; private keys never leave the nodes. The PostgreSQL proxy exposes only the WireGuard address.
  • Streaming replication: flo failover replication setup <tenant> takes a base backup of the standby from the primary and starts streaming. Default mode is async; --mode sync is supported but must be an approved change.
  • Authority: MON assigns a monotonically increasing epoch per tenant. At most one node can hold a valid primary lease for an epoch. The signed lease carries tenant, node, role, epoch, issue/expiry time, incident ID, and signing key ID. Renewal interval is 10 s, lease duration 30 s, allowed clock skew 5 s (FLO_CONTROL_DR_LEASE_TTL_MS=30000, FLO_CONTROL_DR_LEASE_CLOCK_SKEW_MS=5000).
  • RPO barrier: on the controlled path MON revokes new writes, waits for active writers to drain, records a source WAL flush LSN barrier, and waits until the standby replays it before stopping the old source and authorizing the target. On the unreachable-source path MON waits for lease expiry plus clock skew and marks RPO as unknown.
  • Promotion order: epoch commit → target authority → start target application → verify PostgreSQL/application/direct-origin readiness → repoint DNS. Readiness checks are database-aware, and DNS moves only after they pass.
  • Manual failback: after an automatic failover MON keeps the standby active; failback is always an explicit operator command (flo failover back).

Anti-split-brain​

Anti-split-brain is enforced by MON's single-authority model: one epoch, one valid primary lease, fail-closed agents and backend guards. The earlier witness/quorum design (a third VPS voting on primary liveness, failover setup --witness, --arm/auto-promote configuration) was retired in the DR control security completion; those options are absent from the public and remote contracts, and the legacy witnessVps registry field is removed on load. flo failover watch is observe-only: without an external fencing provider it reports a suspected failure but cannot promote, change DNS, or dispatch commands.

Tenant-side protections​

When a tenant is armed, the Flo backend and the node agent both fail closed:

  • PrimaryLeaseGuard (IPrimaryLeaseGuard) is the one authority decision for backend execution paths. With an absent, expired, invalid, or stale lease it blocks startup migrations, database commands, background work, and new external side effects (media, email, jobs); CRUD controllers and health checks consult it too. Config editor writes and cleanup are also lease-guarded, so a standby is effectively read-only.
  • Hosted services wrapped by PrimaryLeaseHostedService<T> wait for a valid lease before starting, poll it every second, stop the inner service on lease loss, and stop the host if cleanup does not finish in time — a fresh process then waits for a new lease.
  • Node agent (flo failover agent ...) is the fail-closed local enforcer: on lease expiry or signature failure it stops the tenant's application containers (it never stops the VPS) and preserves the highest accepted epoch across restarts.
  • Startup gate: FAILOVER_ENABLED gates the whole lease subsystem. When it is absent or false, no lease middleware, EF interceptor, or lease health surface is active, and the app behaves exactly like a pre-DR build. An invalid value refuses startup (exit 78) instead of silently disarming an enrolled node.

Master switches​

SwitchSurfaceDefaultEffect
FAILOVER_ENABLEDTenant backendoffEnables the lease middleware, EF interceptor, lease health checks
FLO_CONTROL_DR_ENABLEDMONfalseGates the DR store, signing key, workers, and /internal/dr/* + /api/dr/* routes
FLO_CONTROL_DR_AUTOMATION_ENABLEDMONfalseAutomatic promotion; set only by flo control deploy mon --enable-dr-automation

Disabling a switch never disarms implicitly: --disable-dr-automation does not replace a full tenant disarm.

Operations​

All commands target the node that runs the child operation (--vps <node>); auto and incident read MON even when --vps points at the primary. Without --vps, failover commands resolve the tenant's location from the registry.

CommandDescriptionKey options
flo failover setup [tenant]Enrol a tenant pair (primary + standby)--standby <profile> (required), --primary <profile>, --all, --dry-run, -y
flo failover replication setup <tenant>Base-backup the standby and start streaming--mode async|sync, --dry-run, -y
flo failover replication status <tenant>Probe and record live replication state
flo failover statusDR posture for all enrolled instances (primary/standby, replication state)
flo failover monitorReplication health + image/secret drift; exits non-zero when degraded--tenant, --all, --notify, --max-lag <s>, -y
flo failover sync-image <tenant>Pull + pin the primary's current image on the standby-y
flo failover sync-secrets <tenant>Re-mirror OIDC certs + DB/audit secrets after rotation-y
flo failover sync-media <tenant>Mirror local-filesystem uploads primary→standby (no-op for R2)-y
flo failover runFail over to a standby (fence → promote → start → verify → DNS)--to <profile> (required), --tenant or --all, --dry-run, --resume-dns, -y
flo failover back <tenant>Fail back to the recovered original primary (re-clone → promote)--to <profile> (required), -y
flo failover watchObserve-only liveness watcher; records/emails suspected outages--tenant, --all, --interval, --threshold, --notify, -y
flo failover drill <tenant>Game-day test: disaster → failover → RTO/RPO verify → fail back--chaos, --max-rpo, --keep, --autonomous, --dry-run, --check, -y
flo failover auto register <tenant>Register the pair in the MON DR store-y
flo failover auto prepare <tenant>Enable incumbent leases after role separation; promotion stays off-y
flo failover auto status <tenant>Durable automatic-promotion lifecycle
flo failover auto arm <tenant>Allow MON automatic authority after full preflight-y
flo failover auto disarm <tenant>Two-step safe disarm; --finalize ends lease issue and disarms both agents--finalize, -y
flo failover auto restore <tenant>Idempotent re-arm: pair back to armed with a streaming standby-y
flo failover incident list|show|explain|export|reconcile|resume|abortDurable incident record: list active, inspect transitions, explain the decision, export JSON, reconcile after restart, resume/abort blocked incidentsshow/explain/export/reconcile take an incident ID; --out <path>; resume/abort need -y
flo failover agent install|upgradeInstall/upgrade the root-only agent + systemd unit on one VPS--config <path> (required), -y
flo failover agent enrollPrepared-to-enrolled transition for one tenant--tenant, --config, -y
flo failover agent status|daemon|uninstallAgent state, service process, removal (unarmed all-prepared only)uninstall needs --config, -y
flo failover teardown <tenant>Remove only the standby copy + replication (never touches Cloudflare)--dry-run, -y
flo failover stormBreak several test tenants at once, each with its own chaos, then re-arm--targets <tenant[:chaos],...> (required), --sequential, -y
flo failover soak <tenant>Repeat the autonomous drill N times with rotating chaos--runs, --chaos, --max-duration, --continue-on-failure, --report <path>, -y

Read forms (status, replication status, monitor, incident list/show, auto status) are part of the authenticated remote registry; write forms require the control plane and -y. The remote drill form validates --chaos, --max-rpo, --dry-run, --check and rejects --force and --keep.

Drills​

flo --vps production failover drill <test-id> --dry-run # print the plan
flo --vps production failover drill <test-id> --check # live prerequisites, no chaos
flo --vps production failover drill <test-id> --chaos kill-db --max-rpo 0 -y
flo --vps production failover drill <test-id> --autonomous --chaos stop-hard -y
  • Chaos modes: kill-db, kill-app, stop-hard, pause (or random).
  • Only tenants explicitly marked isTest: true may be drilled destructively.
  • --dry-run prints a plan (it does not prove success); --check verifies live prerequisites (streaming, mode, images, secrets, routing, authority) without canary, chaos, promotion, or DNS writes.
  • --autonomous waits for a matching automatic MON incident started after the fault and never creates a manual incident as a fallback. A timeout or blocked incident is not a success.
  • --max-rpo <writes> is an acceptance threshold, not an async-replication guarantee; the report measures committed writes actually lost.
  • storm and soak scale drills across tenants/runs; soak appends one JSON line per run to --report and re-arms between runs.

What happens to tenants during a failover​

  1. MON decides (or the operator runs failover run); one failed probe can never trigger a promotion, and MON first attempts one scoped local restart when policy allows.
  2. The source is fenced: new writes are revoked and drained (controlled path) or the lease is left to expire plus skew (unreachable path). The old primary's backend and agent refuse work without a valid lease.
  3. MON commits a new epoch and issues target authority; the standby PostgreSQL is promoted and the application starts on the target.
  4. Readiness is verified (PostgreSQL, application, direct origin) before DNS is repointed; the registry is only a projection of MON state.
  5. After a run, the pairs involved are UNPROTECTED (no hot standby) until the old primary recovers and is failed back / re-protected. flo failover status lists unprotected pairs with the exact re-protection command.

Incident checklist​

flo control whoami
flo --vps <primary> vps stats
flo --vps <primary> failover status
flo --vps <primary> failover monitor --all -y
flo --vps <primary> failover incident list
  • Read the durable record before acting: incident show <id>, then explain.
  • incident resume <id> -y only after a fresh live safety probe; incident abort <id> -y only after source and standby recovery are proven.
  • Never delete incidents, leases, epochs, or fence files to make a drill pass; do not hand-edit the MON database.
  • Complete a full DR disarm before tenant maintenance or a MON deploy; the deploy gate rejects outstanding leases, active incidents, or unknown authority.
  • Failback stays manual: flo --vps <standby> failover back <tenant> --to <primary> -y, then flo --vps <primary> failover replication setup <tenant> -y to recreate protection.
  • Keep docs/DisasterRecoveryDrill.md as the step-by-step runbook (authorization, role separation, agent enrollment, drill, disarm, MON deploy/recovery).