Skip to content

High Availability and Fault-Tolerant Architecture

!!! warning "Status: Proposed / Design Review Required" High Availability / Fault-Tolerant Architecture is a design proposal. It is not implemented in the current default hxEASM deployment.

The default deployment remains the Docker Compose topology described in [setup.md](../administration/setup.md). This document preserves the proposed architecture for review before implementation.

Current Deployment Architecture

The current production-oriented deployment is a single Docker Compose stack. Only the frontend Nginx container publishes a host port; API, PostgreSQL, Redis, and MinIO are internal Docker-network services.

Service Role Stateful? Persistent data Dependencies Current replicas Healthcheck Failure impact
frontend Public Nginx entrypoint, React SPA, /api/ proxy, /easm-files/ proxy Mostly no generated Nginx config and optional htpasswd inside container API and MinIO for proxied routes 1 GET :3000/health UI/API entrypoint outage
api Go API, auth, RBAC, migrations, reports, settings, backup API Mostly stateless, with local report files/cache local ./reports; in-memory Update Center cache PostgreSQL, Redis, MinIO/S3 1 GET :8080/health/ready Application unavailable if single API dies
worker Redis scan consumer, plugin execution, scheduler, artifact upload Mostly stateless, with scanner cache nuclei/tool state under mounted worker home; temporary scan artifacts PostgreSQL, Redis, MinIO/S3 1 by default; scalable via WORKER_REPLICAS ./worker healthcheck Scans stop or may be left incomplete
postgres Primary relational data store Yes ~/.hxeasm/data/postgres host disk 1 pg_isready Critical outage; highest data-loss risk
redis Scan queue Transient state ~/.hxeasm/data/redis, no explicit AOF policy in Compose host disk 1 redis-cli ping Queued scans can be lost or delayed
minio S3-compatible object storage Yes ~/.hxeasm/data/minio host disk 1 MinIO /minio/health/ready files, screenshots, avatars, and evidence unavailable
mailpit Local/dev SMTP sink No none none 1 none Email OTP test delivery unavailable

Current logical topology:

Users
  |
  v
Frontend Nginx :3000
  |-- serves React SPA
  |-- /api/        -> api:8080
  |-- /easm-files/ -> minio:9000 for presigned object URLs
  |
API
  |-- PostgreSQL: users, orgs, assets, vulnerabilities, scans, audit, comments, notifications, reports metadata
  |-- Redis: scan queue easm:scan:queue
  |-- MinIO/S3: avatars, uploaded plugin artifacts, screenshots, evidence/files
  |
Worker(s)
  |-- BRPOP Redis scan queue
  |-- execute plugin chain
  |-- persist results to PostgreSQL
  |-- upload artifacts to MinIO/S3
  |-- run scheduled scan checker

Current Single Points of Failure

SPOF Failure scenario Service impact Data risk Current recovery behavior
Docker host Host, disk, or host network fails Full outage High if host data is lost Manual host/container/data restore
Frontend/Nginx Container dies or port 3000 unavailable UI and proxied API unavailable None directly Docker restart/manual restart
API API container dies API unavailable; UI shell may load but actions fail Low, except in-flight report generation Docker restart/manual restart
Worker Worker dies during scan New scans stop; running scan may stay incomplete Redis message may already be popped Restart and manual retry where possible
PostgreSQL Database unavailable or corrupt Full application outage Highest Restore database or host data
Redis Redis unavailable or data lost Scans cannot enqueue/dequeue; queued messages may be lost Medium for queued work, low for canonical business data Restart Redis; incomplete DB scan state may need manual cleanup/retry
MinIO Object storage unavailable or data lost File downloads/uploads, screenshots, avatars, and storage readiness fail High for objects Restore MinIO data or object backup
Local report directory API local ./reports lost or request routed to a different API in future HA Ready report metadata may point to missing file Medium Regenerate report if possible
Configuration/secrets Replicas use different JWT, 2FA, DB, Redis, S3, or Basic Auth secrets Login failures, invalid tokens, broken presigned URLs High operational risk Manual correction
Worker scheduler Multiple workers run scheduler simultaneously Duplicate scheduled scans are possible Low direct data risk; operational noise No current leader election or lock

High Availability vs Backup / DR

High Availability keeps service running through component failure. Examples include load-balanced frontend/API replicas, PostgreSQL failover, Redis Sentinel, and distributed object storage.

Backup / Disaster Recovery restores after corruption, deletion, data loss, or site loss. Examples include PostgreSQL base backups, WAL archive/PITR, object storage backup, and off-site secret/config backups.

Both are required. HA without backups does not protect against data corruption or operator error. Backups without HA do not prevent outages.

Proposed Operational Targets

These are proposed operating targets, not current SLA commitments.

Area Proposed target
Frontend/API replica failure RTO under 1-3 minutes
Worker failure RTO under about 5 minutes for new work; running scan recovery needs application changes
PostgreSQL primary failure RTO under 5-15 minutes depending on managed/self-hosted HA design
PostgreSQL RPO near-zero with synchronous HA; seconds/minutes with asynchronous replication
Object storage node failure near-zero outage in distributed or managed object storage
Site-level DR RTO about 1-4 hours initially
Availability objective practical enterprise target of 99.5-99.9%

Deployment Levels

A. Minimal Resilient Deployment

Suitable for smaller or on-prem customers that are not ready for a full orchestrator.

Possible improvements:

  • keep the current application tier;
  • add stronger PostgreSQL and object-storage backups;
  • use external or managed PostgreSQL where possible;
  • use external S3-compatible object storage where possible;
  • enable deliberate Redis persistence such as AOF if remaining self-hosted;
  • place an external load balancer or reverse proxy in front of the app;
  • monitor the existing healthchecks.

Docker Compose limitations remain:

  • no cross-host rescheduling;
  • no quorum;
  • no leader election;
  • no orchestrator-controlled rolling failover;
  • no native multi-host stateful HA.

The primary recommended architecture is a stateless application tier plus external HA stateful services:

  • orchestrated frontend/API/worker replicas;
  • external HA PostgreSQL;
  • managed Redis or Redis Sentinel;
  • managed S3-compatible storage or distributed MinIO;
  • external load balancer or ingress;
  • centralized secrets/configuration;
  • one active scheduler until leader election or DB locking exists.

C. Advanced HA / DR

Suitable for larger enterprise environments:

  • multi-node Kubernetes/RKE2/OpenShift, Nomad, or customer-standard orchestrator;
  • stronger database quorum/failover;
  • Redis HA;
  • distributed object storage;
  • blue/green or rolling deployments;
  • off-site DR environment;
  • scheduled failover and restore exercises.

None of these deployment levels are implemented by the default Compose stack today.

Recommended future architecture:

                    Users
                      |
              External LB / Ingress
                      |
          +-----------+-----------+
          |                       |
     Frontend 1               Frontend 2+
          |                       |
          +-----------+-----------+
                      |
                  API Service
            +---------+---------+
            |         |         |
          API-1     API-2     API-3
            |         |         |
            +---------+---------+
                      |
         +------------+-------------+
         |            |             |
    PostgreSQL HA   Redis HA     S3 / MinIO HA
         |
     WAL archive / backups

             Worker Pool
         Worker-1 Worker-2 ...

             Scheduler
           single active

This is proposed architecture. It is not currently generated by hxEASM.

Frontend / Ingress HA

The frontend is mostly stateless and can conceptually run active-active behind an external load balancer or ingress.

Requirements and caveats:

  • every frontend replica must receive the same Basic Auth htpasswd secret when Basic Auth is enabled;
  • the load balancer should health-check /health;
  • sticky sessions should not normally be required because application authentication uses JWTs/API keys;
  • the current Nginx upstream configuration assumes Docker service names such as api:8080 and minio:9000;
  • future HA deployment should make upstreams orchestrator-friendly.

API HA

The API is mostly stateless:

  • authentication tokens are JWT-based when all replicas share the same JWT secret;
  • refresh/session tokens are stateless JWTs with a non-sliding deadline; users, RBAC, 2FA challenges, comments, notifications, and audit data live in PostgreSQL;
  • file bytes live in S3-compatible storage;
  • Update Center cache is per-process and safe to lose.

Current blockers and caveats:

  1. API runs database migrations on startup.
  2. Reports are generated and stored in local ./reports.
  3. Database pool sizing must account for all API replicas.
  4. Every replica must use identical configuration and secrets.

Recommended future changes:

  • run migrations as a single pre-deploy job;
  • move generated report output to shared object storage;
  • add PgBouncer or connection planning if API replica count grows;
  • centralize all secrets/configuration.

Worker HA / Reliability

Current worker behavior:

  • API and worker push scan messages into Redis;
  • workers consume from Redis with BRPOP;
  • the queue message is removed before scan execution completes;
  • there is no ack-after-completion, processing list, or visibility timeout;
  • a worker crash can leave a scan or scan job incomplete;
  • every worker instance also starts the scheduler;
  • duplicate scheduled scans are possible if multiple workers process the same due schedule window.

Worker scaling currently improves throughput but does not yet provide strong reliable job execution semantics.

Recommended future changes:

  • separate the scheduler process, or add leader election / DB advisory locking;
  • add reliable job claiming through a processing list and visibility timeout, or move to DB-backed job claiming;
  • add stuck job recovery;
  • define worker drain behavior before upgrades.

PostgreSQL HA

PostgreSQL is the most critical stateful component.

Options:

Option Assessment
Managed PostgreSQL HA Preferred where available; lowest operational burden.
Patroni + etcd/Consul Preferred self-hosted/on-prem pattern.
Streaming replication + manual failover Simpler, but weaker RTO and higher operator risk.
PostgreSQL operator Acceptable when Kubernetes expertise exists.

Recommended design:

  • one writable primary;
  • one or more standby replicas;
  • quorum/leader management;
  • stable DB endpoint for API/worker;
  • WAL archiving;
  • PITR;
  • automated backups;
  • migrations executed once during deployment.

Never use two independent writable PostgreSQL primaries.

Redis HA

Redis primarily carries scan queue state, not canonical business records.

Redis loss can cause:

  • queued scans to be lost or delayed;
  • scans/jobs in PostgreSQL to require manual cleanup/retry;
  • temporary inability to enqueue or consume scans.

Recommended options:

  • Minimal: single Redis with AOF and monitoring.
  • Enterprise: managed Redis HA or Redis Sentinel.

Redis Cluster is not the default recommendation unless queue semantics and client support are reviewed. The current application config uses one Redis endpoint/client.

Object Storage HA

MinIO/S3 stores:

  • file artifacts;
  • screenshots;
  • evidence/raw outputs;
  • avatars;
  • uploaded/restored file objects.

A second standalone MinIO instance is not HA. It only creates another isolated storage endpoint.

Recommended options:

  • managed S3-compatible object storage;
  • distributed MinIO with quorum and erasure coding;
  • object replication/versioning where appropriate;
  • off-site object backups and restore verification.

Generated reports are currently local API files, not shared object-storage files.

Persistent Data Inventory

Data Stored in Criticality Replication needed Backup needed Acceptable loss
users/RBAC/organizations PostgreSQL Critical Yes Yes/PITR Near-zero
2FA state and recovery hashes PostgreSQL Critical Yes Yes/PITR Near-zero
API key hashes PostgreSQL Critical Yes Yes/PITR Near-zero
Assets and asset graph PostgreSQL Critical Yes Yes/PITR Near-zero
Vulnerabilities PostgreSQL Critical Yes Yes/PITR Near-zero
Scans and scan jobs PostgreSQL plus Redis queue High PostgreSQL yes; Redis recommended Yes Seconds/minutes for queue only if recoverable
Audit logs PostgreSQL High Yes Yes/PITR Near-zero
Asset History PostgreSQL High Yes Yes/PITR Near-zero
Vulnerability discussions/mentions PostgreSQL High Yes Yes/PITR Near-zero
User notifications PostgreSQL Medium Yes Yes Low but undesirable
Settings PostgreSQL/config High Yes Yes Near-zero
Redis jobs Redis Medium Recommended AOF optional Queued work can be recreated only with operational effort today
Files/screenshots/evidence S3/MinIO plus PostgreSQL metadata High Yes Yes Low
Reports PostgreSQL metadata plus API local ./reports Medium Shared storage needed Yes Often regenerable
Avatars S3/MinIO plus users metadata Medium Yes Yes Low
Configuration/secrets env/config/secret store Critical Yes Yes None for JWT/2FA/storage secrets

Failover Mechanisms

Proposed future failover:

  • Frontend/API: external load balancer removes unhealthy replicas.
  • PostgreSQL: managed or Patroni failover to a standby.
  • Redis: managed Redis or Sentinel failover.
  • Object storage: managed object storage or distributed MinIO quorum.
  • Site failure: DR DNS/VIP cutover to a secondary site.

These mechanisms are not active in the default deployment today.

Healthcheck Usage

Existing health surfaces:

Surface Future HA use
frontend /health Load-balancer readiness/liveness for Nginx
API /health/live Process liveness
API /health/ready Load-balancer readiness; checks PostgreSQL, Redis, and storage
worker ./worker healthcheck Container health; checks config, PostgreSQL, Redis, and storage
PostgreSQL pg_isready Container dependency health
Redis redis-cli ping Container dependency health
MinIO readiness endpoint Object storage readiness

Readiness should remove replicas from traffic. Liveness should avoid restart loops during temporary dependency outage.

Failure Scenarios

Scenario Detection Proposed failover User impact Data impact Proposed RTO/RPO
Frontend failure /health fails LB routes to another frontend brief UI/API proxy interruption none seconds/minutes, RPO 0
API failure /health/ready fails LB routes to another API in-flight requests fail report generation may fail seconds/minutes
Worker failure mid-scan worker health fails another worker handles new messages running scan may stick/fail popped Redis message may be lost RTO minutes; RPO queue-message risk
Redis failure ping/readiness fails Sentinel/managed failover scans cannot enqueue/dequeue during outage queued messages at risk without HA/AOF minutes; queue-dependent
PostgreSQL primary failure DB health/replication monitoring managed/Patroni failover app read/write outage during failover depends on sync/async replication 5-15 minutes; RPO 0 to seconds
MinIO node failure MinIO health/quorum monitoring distributed/managed storage masks node loss file functions remain available if quorum healthy no loss if quorum healthy near-zero
Application host failure host monitoring orchestrator reschedules or LB uses another host degraded if replicas exist local reports/tool cache lost minutes
Site failure site monitoring DR DNS/VIP to secondary site outage until DR promoted depends on PITR/object replication 1-4 hours initially
Network partition quorum monitoring leader systems fence minority degraded service possible split-brain if misconfigured depends on quorum design
Bad deployment readiness/errors rollback or blue/green switch degraded/outage DB migrations may limit rollback minutes if compatible

Split-Brain / Consistency

  • PostgreSQL requires leader/quorum management. Do not run two independent writable primaries.
  • Redis HA should use Sentinel/managed failover quorum when reliability is required.
  • MinIO HA should use distributed quorum/erasure coding or managed object storage.
  • Independent active-active MinIO nodes without replication/quorum are not safe HA.

Backup / DR Strategy

PostgreSQL:

  • periodic base backups;
  • continuous WAL archive;
  • PITR;
  • encrypted off-site retention;
  • restore verification.

Object storage:

  • replication and/or versioning;
  • off-site backup;
  • object recovery verification.

Configuration and secrets:

  • secret manager or reproducible deployment;
  • backup critical keys such as JWT secret, 2FA encryption key, storage credentials, SMTP credentials, and Basic Auth credentials.

Redis:

  • AOF is useful for queued work;
  • Redis backup is lower priority than PostgreSQL/S3 because canonical business data lives elsewhere.

The default deployment does not provide these HA/DR guarantees automatically.

Restore Testing

Untested backups are insufficient.

Recommended procedures:

  • regular PostgreSQL PITR restore test;
  • object storage restore test;
  • full environment rebuild exercise;
  • quarterly failover drill;
  • documented recovery runbooks;
  • measured RTO/RPO after each exercise.

Upgrade Strategy

Current blockers:

  • API runs migrations at startup;
  • reports are local to one API filesystem;
  • workers do not have drain/ack semantics;
  • scheduler runs in every worker.

Proposed HA upgrade process:

  1. Control the scheduler.
  2. Drain or stop workers.
  3. Run migrations once.
  4. Roll API replicas.
  5. Roll frontend replicas.
  6. Roll workers.
  7. Validate health, queue depth, scan jobs, and failed jobs.

Near-zero downtime requires backward-compatible database migrations.

Observability Requirements

Recommended monitoring:

  • frontend/API/worker health;
  • API dependency readiness;
  • PostgreSQL primary/replica health;
  • PostgreSQL replication lag;
  • WAL backup status;
  • Redis health, memory, and persistence status;
  • Redis queue depth;
  • MinIO quorum and disk capacity;
  • stuck scans and aging running jobs;
  • failed scan/job rate;
  • backup success;
  • restore-test age;
  • TLS certificate expiry;
  • configuration/secret drift between replicas.

Secrets / Configuration

All replicas must receive consistent:

  • JWT secret;
  • 2FA encryption key;
  • Update Center API key;
  • SMTP credentials;
  • storage credentials;
  • Basic Auth credentials;
  • DB/Redis credentials;
  • scanner/profile config.

Recommended secret stores:

  • Kubernetes Secrets;
  • Vault;
  • customer secret manager;
  • Docker secrets for VM deployments.

Do not hand-edit per-node secrets in HA deployments.

Minimum Node / Replica Guidance

Component Proposed minimum
Frontend 2+ replicas
API 2+ replicas, often 3
Worker 2+ consumers, but one active scheduler until leader election exists
PostgreSQL managed HA or quorum-based primary/standby design
Redis managed HA or Sentinel topology
Object storage managed S3 or distributed MinIO
Load balancer managed LB or redundant HAProxy/Keepalived pair

Exact counts depend on customer requirements, failure domains, and deployment technology.

Application Changes Required Before Full HA

All items below are not implemented.

  1. Separate database migrations from API startup.
  2. Move reports from API local disk to shared storage.
  3. Improve worker job reliability and acknowledgement.
  4. Separate scheduler or introduce leader election.
  5. Add HA Redis endpoint/Sentinel support if needed.
  6. Make frontend upstream configuration orchestrator-friendly.
  7. Ensure all replica secrets/config are centralized.

Implementation Roadmap

Phase 1 - Harden Current Deployment

  • backups;
  • monitoring;
  • deliberate Redis persistence;
  • external DB/S3 where possible;
  • secrets discipline.

Phase 2 - Stateless Application Replicas

  • external load balancer;
  • multiple frontend/API replicas;
  • centralized secrets;
  • one migration job;
  • shared report storage.

Phase 3 - Worker Reliability

  • singleton scheduler;
  • reliable job claim/ack;
  • stuck job recovery.

Phase 4 - Stateful HA

  • PostgreSQL HA;
  • Redis HA;
  • distributed or managed object storage.

Phase 5 - Orchestrated Enterprise Deployment

  • Kubernetes/RKE2/OpenShift/customer orchestrator;
  • rolling or blue/green upgrades;
  • monitoring and alerting;
  • DR exercises.

These phases describe future implementation work, not current product behavior.

Acceptance Criteria Mapping

Design requirement Covered
SPOFs identified Docker host, frontend, API, worker, PostgreSQL, Redis, MinIO, local reports, secrets, scheduler
HA configuration selected Stateless app tier plus external HA stateful services
Redundancy defined frontend/API replicas, worker pool, PostgreSQL HA, Redis HA, object storage HA
Failover defined LB health, Postgres failover, Redis failover, object storage quorum, DR DNS/VIP
Recovery requirements defined PITR, object backups, secret backups, restore testing
Final architecture prepared Recommended Enterprise Architecture section

This matrix covers the design task only. HA implementation is not complete.

Review Required Before Implementation

Suggested reviewers:

  • hxEASM architecture/backend;
  • DevOps/SRE;
  • database/platform engineer;
  • security engineering;
  • customer infrastructure/operations representative for on-prem scenarios.

Questions requiring review:

  1. Is Kubernetes/RKE2 the default enterprise target or only one supported profile?
  2. Do we support VM-based HA as a first-class deployment?
  3. What RTO/RPO do we actually commit to?
  4. What are the managed vs self-hosted PostgreSQL support expectations?
  5. Is Redis reliable queue redesign required before an HA release?
  6. Should reports move to S3 before multi-API support?
  7. What is the supported MinIO topology?
  8. Which secrets manager approaches are officially supported?
  9. What level of DR is part of the product vs the customer's infrastructure responsibility?