Skip to main content

Common Issues

A practical playbook for the issues that come up most often in production. Each entry follows the same shape: symptom, checks, fix.

Login lockout (5 failed attempts / 15 minutes)​

Symptom: a correct password is rejected and the response carries errors.auth.tooManyAttempts with the remaining minutes. This is the per-account lockout, not the per-IP rate limiter (a 429 with errors.general.rateLimitExceeded is the limiter instead).

Checks:

  • The lockout counter lives in the LoginAttempt table: one row per attempted email with FailedCount, LockedUntil, and LastAttemptAt. Because it is in the database, blue-green containers and restarts share the same view.
  • SuperAdmin: Settings > Accessi and errori (/dashboard/amministrazione/settings/accessi-errori, API /api/v1/auth-events) records masked auth events, including account_locked.
  • App logs show a masked warning (Login attempt for locked account) with the remaining minutes.

Fix:

  • Wait for the 15-minute window. The counter restarts after the window elapses, and a successful login deletes the row.
  • Do not delete or edit lockout rows by hand: the lockout is time-based and the next failed attempt would simply re-create the state.
  • If lockouts recur without the user retrying, treat it as credential stuffing: inspect auth events and the rate-limit logs, and rotate the password.

Email not delivered​

Symptom: a user does not receive an activation, OTP, password-recovery, or newsletter email.

Checks:

  • Transactional sending is gated by the enable_email_sender feature flag.
  • Provider configuration: EMAIL_SENDER_EMAIL_PROVIDER plus its credentials. In development, an empty provider means console mode (messages are logged, not sent). See Email Configuration.
  • Bulk/newsletter sending uses the provider selected in Communication settings (DB), with the env default as fallback. A quota-exhausted provider (429/402) parks the job until the next UTC midnight, and a non-empty test-email allow-list redirects every message to those addresses only.
  • Suppression list: newsletter_suppressions holds permanently excluded recipients (bounce, complaint, manual, unsubscribe). The send queue marks suppressed addresses as Suppressed, so sent + failed + suppressed reconciles to the job total.
  • Delivery log: Settings > Sent emails (SuperAdmin, /api/v1/email-delivery-logs) merges the local log (180-day retention) with live Cloudflare analytics when configured. Filter by recipient, status, provider, and date range.
  • Container logs for provider errors (flo logs app <tenant>, or the app log viewer in Settings and Audit).

Fix:

  • Correct the provider environment variables and restart the container: environment variables are read at startup.
  • For Gmail aliases, verify the "Send As" configuration; for custom SMTP, check host/port/TLS and credentials.
  • For newsletters, re-select the provider or set its credentials; retry after the quota reset if the provider was parked.
  • Review the suppression reason before un-suppressing: an unsubscribe is a legitimate preference, not a delivery failure.
  • Replay failed transactional messages with POST /api/v1/maintenance/email/replay-failed and the X-Maintenance-Key header (Maintenance__ApiKey). It supports dryRun and recipient/user filters, and answers 503 when no email provider is resolvable or refuses when enable_email_sender is off (errors.emailDeliveryLog.replayEmailDisabled).

Payment webhook issues​

Symptom: the provider reports a successful payment but the order stays pending, or the webhook returns 400 with errors.payment.invalidWebhook.

Checks:

  • Webhook URL configured at the provider: POST https://<domain>/api/v1/payments/{provider}/webhook, where {provider} is the configured gateway (stripe or paypal, matching PAYMENT__PROVIDER). The endpoint is anonymous, caps the body at 64 KB, and is rate-limited at 120/min per IP.
  • Signature verification: Stripe uses PAYMENT__STRIPE__WEBHOOKSECRET (the whsec_ value) against the Stripe-Signature header; PayPal uses PAYMENT__PAYPAL__WEBHOOKID for signature verification. A wrong or missing secret makes parsing fail, which surfaces as errors.payment.invalidWebhook.
  • Dedup and processing: PaymentWebhookEvent rows record the external event id, event type, received/processed timestamps, and any error message.
  • Order lifecycle: pending/held booking orders expire on schedule when no webhook confirms them.

Fix:

  • Copy the signing secret from the provider dashboard into the environment variable, restart the container, and re-send the event from the provider.
  • Commerce (studio plan) orders: POST /api/v1/commerce/orders/{id}/reconcile re-reads the gateway and updates the order.
  • Booking payment orders: resolve a needs-review order with POST /api/v1/payments/admin/orders/{id}/review; retry e-invoice emission with POST /api/v1/payments/orders/{id}/invoice/retry (available when the invoice job is NeedsReview or Failed).
  • Verify the gateway is configured: GET /api/v1/payments/admin/readiness.

Maintenance mode​

Symptom: guarded APIs answer 503 with errors.general.maintenanceMode.

Checks:

  • flo --vps production instance maintenance <tenant> status returns enabled, activeRequests, and leaseBoundMaintenanceBlocked for each backend.
  • The MAINTENANCE_MODE environment variable is the boot-time default; a live flip overrides it for the running process without a restart.

Fix:

  • Turn it off: flo --vps production instance maintenance <tenant> off. The command verifies the resulting state on every backend.
  • If leaseBoundMaintenanceBlocked is reported, MON controls the tenant: complete flo failover auto disarm before changing maintenance state.
  • If a flip fails and the state cannot be verified, the CLI stops the backends to block traffic. Treat that as an operational incident and inspect the cutover state and every container before starting anything.
  • To refresh the feature-flag cache while maintenance is on, use flo --vps production instance maintenance-cache-invalidate <tenant>.

Failover and lease warnings​

Symptom: a tenant on a standby stops serving, health checks fail, or logs report lease denials (lease-missing, lease-expired, lease-invalid-signature, lease-fenced, lease-stale-epoch, and similar reasons). An invalid FAILOVER_ENABLED value refuses startup instead of silently disarming a node.

Checks:

flo control whoami
flo --vps <node> vps stats
flo --vps <node> failover status
flo --vps <node> failover monitor --all -y
flo --vps <node> failover incident list
  • Read the durable record before acting: flo failover incident show <id>, then flo failover incident explain <id>.

Fix:

  • incident resume <id> -y only after a fresh live safety probe; incident abort <id> -y only after source and standby recovery are proven.
  • Never delete incidents, leases, epochs, or fence files to make a drill pass, and do not hand-edit the MON database.
  • Failback stays manual: flo --vps <standby> failover back <tenant> --to <primary> -y, then recreate protection with flo --vps <primary> failover replication setup <tenant> -y.
  • Deploys refuse fenced tenants and tenants without a published primary lease; fix the DR state instead of bypassing the guard.

Full operator runbook: Failover and Disaster Recovery.